Dynamic optimization of large model inference performance and hardware-aware compression methods

By extracting input data features and monitoring hardware status in real time, generating dynamic control signals, and dynamically reconfiguring large model weights and activation values, the problems of low resource utilization and energy efficiency imbalance in large model reasoning are solved, and efficient computing is achieved under dynamic input and heterogeneous hardware environments.

CN120494006BActive Publication Date: 2025-09-12KARAMAY HONGYOU SOFTWARE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510983404.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-09-12
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Existing technologies have problems with large computational complexity, high memory requirements, and low energy efficiency in large-model inference, especially when deployed in edge devices and real-time scenarios. In addition, existing compression methods cannot adapt to dynamic inputs and heterogeneous hardware environments, resulting in latency fluctuations and energy efficiency imbalance.

Method used

By extracting the complexity characteristics of input data in real time and synchronously monitoring the hardware load status, joint control signals for quantization bit width, sparsification ratio and operator scheduling are generated, and model weights and activation values ​​are dynamically reconfigured to form a closed-loop optimization mechanism, achieving dynamic adaptation of computing accuracy and hardware resources.

Benefits of technology

Under dynamic input and heterogeneous hardware environments, resource utilization is improved, latency fluctuations and energy efficiency imbalances are reduced, and the inference performance of large models is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494006B_ABST
    Figure CN120494006B_ABST
Patent Text Reader

Abstract

The large-model inference efficiency dynamic optimization and hardware-aware compression method of the present invention includes five steps: S1: generating an input complexity signal that characterizes the computational complexity; S2: synchronously monitoring the hardware resource indicators of the operating platform to generate a hardware status signal that reflects the real-time load; S3: inputting the input complexity signal and the hardware status signal into the dynamic strategy selector, and generating a compression control signal through a pre-trained decision model; S4: performing a dynamic reconfiguration operation on the large model weights and activation values ​​of the current inference task according to the compression control signal; S5: using the reconfigured large model to perform inference calculations, and during the calculation process, feeding back the hardware resource indicators to S2 in real time to form a closed-loop optimization link. The large-model inference efficiency dynamic optimization and hardware-aware compression method of the present invention can solve the problems of low resource utilization, delay fluctuations and energy efficiency imbalance caused by static compression methods in dynamic input and heterogeneous hardware environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large model reasoning optimization and hardware-aware compression, and in particular to a method for dynamic optimization of large model reasoning efficiency and hardware-aware compression. Background Art

[0002] With the widespread adoption of large models with tens of billions of parameters, their enormous computational and memory requirements lead to high inference latency and low energy efficiency, severely restricting their deployment on edge devices and in real-time scenarios. Existing technologies primarily rely on static compression methods: knowledge distillation requires additional training and has a fixed compression rate, making it difficult to adapt to changes in input complexity. While structured pruning can reduce computational complexity, static sparsity patterns cannot match the acceleration characteristics of different hardware. While quantization techniques can reduce memory usage, static bit width settings can easily lead to accuracy collapse or resource waste when input complexity changes suddenly. Regarding hardware optimization, traditional solutions, such as precompiled operator libraries (such as TensorRT), only perform static optimizations at deployment time and cannot respond to hardware state fluctuations at runtime.

[0003] While some dynamic technologies (such as early exit mechanisms) have emerged, they only adjust the computation path without incorporating real-time hardware state, and lack the ability to jointly control compression parameters. Current research gaps lie in the lack of a dynamic mapping mechanism between input features, hardware load, and multi-dimensional compression strategies. This results in static solutions facing three real-world disconnects: a disconnect between the dynamic nature of input complexity and fixed compression strategies, a disconnect between hardware resource fluctuations and static optimization strategies, and a disconnect between multi-objective requirements (latency, accuracy, and energy efficiency) and single-dimensional optimization. Particularly in heterogeneous environments like edge computing, these shortcomings can result in up to 40% of potential computing power being idle and energy efficiency being lost. Summary of the Invention

[0004] In view of the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a dynamic optimization and hardware-aware compression method for large-model inference performance, which is used to solve the problems of low resource utilization, delay fluctuation and energy efficiency imbalance caused by static compression methods in dynamic input and heterogeneous hardware environments. The present invention extracts the complexity characteristics of the input data in real time and synchronously monitors the hardware load status, driving the dynamic strategy selector to generate a joint control signal for quantization bit width, sparsification ratio and operator scheduling; based on the signal, the model weights and activation values ​​are reconfigured at runtime to achieve dynamic adaptation of computing accuracy and hardware resources; the hardware indicator feedback of the inference process is used to form a closed-loop optimization, so that the large model automatically maintains optimal performance in changing scenarios.

[0005] The present invention provides a method for dynamic optimization of large model inference performance and hardware-aware compression, including:

[0006] S1: Extract the feature vector of input data in real time and generate an input complexity signal that represents the computational complexity;

[0007] S2: Synchronously monitors the hardware resource indicators of the operating platform and generates hardware status signals that reflect real-time load. Hardware resource indicators include memory usage, computing unit utilization, and power consumption data.

[0008] S3: Input complexity signals and hardware status signals are fed into the dynamic policy selector, which generates a compression control signal through a pre-trained decision model. This signal contains a combination of quantization bit width, sparsification ratio, and operator scheduling strategy instructions.

[0009] S4: Based on the compression control signal, dynamically reconfigure the weights and activation values ​​of the large model of the current inference task. This includes switching between floating-point and fixed-point computing modes based on quantization bit width instructions, activating the structured mask of the corresponding layer based on sparsification ratio instructions, and adapting the hardware acceleration kernel based on operator scheduling instructions.

[0010] S5: Use the reconfigured large model to perform inference calculations, and during the calculation process, feed back hardware resource indicators to S2 in real time, forming a closed-loop optimization link.

[0011] In one embodiment of the present invention, extracting the feature vector of the input data in S1 specifically includes analyzing the sequence length and attention distribution discreteness of the input text in real time through a lightweight convolutional network to generate a feature vector containing hierarchical semantic density information, wherein the input complexity signal is formed by fusing the sequence length feature and the attention entropy value feature through a gated recurrent unit, and the signal dynamically reflects the difference in theoretical computational amount caused by different input samples in each computational layer of the model.

[0012] In one embodiment of the present invention, the process of generating the hardware status signal in S2 includes establishing a direct data channel with the underlying hardware driver, periodically collecting the graphics processor memory bandwidth utilization, the tensor core idle cycle ratio and the on-chip cache miss rate, and mapping the heterogeneous hardware indicators into a load evaluation coefficient of a unified dimension through normalization processing, wherein the power consumption data is obtained by reading the real-time current-voltage product of the onboard power management chip.

[0013] In one embodiment of the present invention, the pre-trained decision model in S3 adopts a dual-channel graph neural network architecture. The first channel analyzes the spatiotemporal correlation features of the input complexity signal, and the second channel learns the fluctuation pattern of the hardware status signal in the time series. After fusing the dual-channel features through the cross-attention mechanism, the compressed control signal is output. The training process of the model uses a reinforcement learning framework with inference delay and energy consumption ratio as the joint reward function.

[0014] In one embodiment of the present invention, when the calculation mode is switched based on the quantization bit width instruction in S4, a pre-compiled integer calculation kernel or a mixed precision calculation kernel is dynamically loaded according to the bit width parameter specified by the control signal, and a real-time calibration node is inserted into the calculation graph to compensate for the quantization error, wherein the fixed-point calculation mode adopts a symmetric quantization strategy to map the floating-point weights to an integer representation with a scaling factor.

[0015] In one embodiment of the present invention, the activation process of the structured mask includes generating a block sparse pattern that meets the requirements of the hardware acceleration unit based on the sparsification ratio instruction, dynamically masking a specified ratio of low weight value areas in the weight matrix, and submitting a sparse matrix compression format identifier to the computing engine to trigger a dedicated computing pipeline.

[0016] In one embodiment of the present invention, the specific implementation of the operator scheduling instruction adaptation hardware acceleration kernel is to select the calculation graph segmentation strategy and memory allocation scheme according to the control signal, enable the asynchronous pipeline parallel mechanism for the graphics processor, start the data flow sharding calculation mode for the neural network processor, and bind the large page memory prefetch strategy for the central processing unit.

[0017] In one embodiment of the present invention, the decision model continuously receives hardware resource indicators fed back by S5 during operation, and updates the edge weight parameters of the graph neural network through an online incremental learning mechanism. The model update process uses a sliding window mechanism to retain historical state features, and ensures strategy stability through knowledge distillation constraints.

[0018] In one embodiment of the present invention, the feedback process of S5 includes an exception handling mechanism. When it is detected that the hardware resource indicator suddenly changes and exceeds the safety threshold, an interrupt signal is immediately sent to the dynamic strategy selector to trigger the emergency strategy, forcing the quantization bit width of the compression control signal to be increased to a safety level and disabling the sparsification operation.

[0019] In one embodiment of the present invention, the closed-loop optimization link is implemented in an edge computing scenario by collecting multi-node hardware status signals through a distributed monitoring agent, and generating a global compression control signal by a central policy selector, wherein the dynamic reconfiguration operation synchronously coordinates the model segmentation calculation and result fusion process among multiple devices.

[0020] The method for dynamic optimization of large-model inference performance and hardware-aware compression provided by the present invention extracts the complexity characteristics of input data in real time and synchronously monitors the hardware load status, driving the dynamic strategy selector to generate a joint control signal for quantization bit width, sparsification ratio and operator scheduling; based on the signal, the model weights and activation values ​​are reconfigured at runtime to achieve dynamic adaptation of computing accuracy and hardware resources; and closed-loop optimization is formed by using hardware indicator feedback of the inference process, so that the large model automatically maintains optimal performance in changing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 Flowchart of the method for dynamic optimization of large model inference performance and hardware-aware compression;

[0023] Figure 2 A schematic diagram showing the dynamic optimization of large model inference performance and hardware-aware compression methods. DETAILED DESCRIPTION

[0024] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0025] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0026] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0027] See Figure 1, which shows the dynamic optimization and hardware-aware compression method for large-scale model inference performance of the present invention. The dynamic optimization and hardware-aware compression method for large-scale model inference performance of the present invention includes five steps: S1: extracting the feature vector of the input data in real time to generate an input complexity signal representing the computational complexity; S2: synchronously monitoring the hardware resource indicators of the operating platform to generate a hardware status signal reflecting the real-time load. The hardware resource indicators include memory occupancy, computing unit utilization, and power consumption data; S3: inputting the input complexity signal and the hardware status signal into the dynamic strategy selector, and generating a compression control signal through a pre-trained decision model. The signal includes a combination of instructions for quantization bit width, sparsification ratio, and operator scheduling strategy; S4: based on the compression control signal, dynamically reconfiguring the weights and activation values ​​of the large model of the current inference task, including switching between floating-point and fixed-point calculation modes based on quantization bit width instructions, activating the structured mask of the corresponding layer based on sparsification ratio instructions, and adapting the hardware acceleration kernel based on operator scheduling instructions; S5: using the reconfigured large model to perform inference calculations, and feeding back the hardware resource indicators to S2 in real time during the calculation process, forming a closed-loop optimization link.

[0028] like Figure 1 As shown, the dynamic optimization and hardware-aware compression system constructed by the present invention achieves real-time improvement in the reasoning efficiency of large models through a five-level linkage mechanism. A lightweight feature extraction module is deployed in the input complexity perception stage. This module uses a convolution kernel to slide and scan the input text sequence to capture the local semantic density change characteristics, and at the same time models long-distance context dependencies through a recurrent neural network. According to the characteristics of the Transformer architecture, the discreteness of the output probability distribution of each attention layer is calculated. When it is detected that the attention is focused on a few tokens, a low-complexity tag is generated, and a uniform distribution triggers a high-complexity alarm. After the feature vector fuses the logarithmic transformation value of the sequence length and the attention entropy value, it is compressed by a three-layer perceptron to generate a hierarchically encoded input complexity signal, which carries the expected load intensity prediction of each computing layer. The hardware status monitoring system establishes a cross-level data collection channel: At the register level, it directly reads the memory controller's row buffer hit counter and dynamically calculates the memory bandwidth pressure index based on the ratio of misses to total accesses. At the instruction pipeline level, it monitors the status of the Tensor Core instruction queue, flagging underutilization of the compute unit when consecutive idle cycles exceed a threshold. The system also converts power consumption into watts in real time by multiplying the current and voltage of the onboard power management chip. All metrics are processed by device-specific normalization functions and then fed into a dynamically weighted fusion engine to generate scalarized hardware status signals. Edge devices prioritize power consumption, while data center-class GPUs prioritize compute utilization.

[0029] Specifically, the decision model of the dynamic policy selector adopts a two-channel graph neural network architecture. The first channel maps the input complexity signal into a spatiotemporal graph structure, where nodes correspond to model computation layers and edge weights represent the strength of computational dependencies between layers. Node features are propagated through a three-layer graph convolutional network to extract dynamic load characteristics of the model's internal computation paths. The second channel constructs a hardware topology graph, where nodes represent physical components such as processors, memory, and cache, and edge attributes reflect the hardware interconnection bandwidth. A graph attention mechanism is used to learn resource competition patterns between components. The output features of the two channels are deeply interacted through a cross-attention module: model computation features are used as query vectors and hardware topology features as key-value pairs to generate hardware-aware computational policy recommendations. This policy is decoded into three-dimensional control instructions by a fully connected layer: a quantized bit width vector specifies the numerical precision level of each layer's weight activation values, a sparsified scaling matrix controls the parameter pruning strength of each attention head, and an operator scheduling code selects the optimal kernel version for the current hardware. The runtime model reconfiguration engine achieves millisecond-level switching based on control instructions. The quantization control unit inserts real-time calibration nodes before weights are loaded into the compute unit. It dynamically calculates floating-point to fixed-point scaling factors based on the bit width directive. For extreme quantization scenarios below 8 bits, it uses piecewise linear functions to approximate nonlinear activations to minimize precision loss. The sparsification unit generates block-shaped mask matrices based on scaling directives and performs structured pruning on the weight tensor. It also submits a hardware-recognizable sparse format identifier to the compute engine, triggering the 2:4 sparse acceleration mode of the GPU tensor cores or the sparse multiply-add array of the NPU. After parsing the policy encoding, the operator scheduler dynamically loads a precompiled kernel component library, binding warp-level optimized asynchronous pipeline kernels to the GPU, configuring dataflow sharding computation routines for the NPU, and deploying a vectorized instruction set bound to large page memory for the CPU. A closed-loop optimization mechanism continuously collects actual hardware load data during inference execution: memory bandwidth utilization is obtained through memory controller performance counters, compute unit active cycles are counted by the instruction issue unit, and power consumption data is read from the power management chip via the I²C bus. These raw metrics are fed back to the hardware signal generation module with millisecond latency, driving the dynamic policy selector for incremental policy tuning. If sparsification causes a sharp drop in cache hit rate, the sparsity ratio for the next inference cycle is automatically reduced. If low-precision quantization causes activation value overflow, the bit width configuration of key layers is dynamically increased. This feedback loop ensures stable system performance despite sudden changes in input distribution or hardware resource competition.

[0030] Furthermore, a gated recurrent unit is employed to process the complexity evolution of continuous input sequences: the sequence length scale value and attention entropy value at the current time step are fed into the update gate to control the proportion of historical complexity features inherited; candidate states fuse text-structured convolutional features with attention distribution features; and the reset gate selects feature dimensions strongly correlated with the current computational load. The final hidden state is mapped to each Transformer layer via a layer-by-layer allocation matrix, generating a fine-grained complexity prediction signal. This mechanism accurately captures complexity jumps caused by the increasing number of objects in a video stream or computational pressure changes caused by increasing question complexity in conversational scenarios. A ring buffer data acquisition architecture is constructed: the bottom layer directly accesses the GPU memory controller's row buffer status registers via PCIe configuration space registers to calculate real-time bandwidth utilization; the driver layer intercepts CUDA runtime API call sequences and analyzes kernel launch intervals to derive unit utilization; the operating system layer collects last-level cache miss rates through performance monitoring events. The collected data is then classified and processed by the device feature encoder: a compute utilization-dominated weighting strategy is applied to GPUs, while a power constraint-prioritized fusion algorithm is used for mobile chips to generate a unified hardware state index. The spatiotemporal graph convolutional network performs three-order neighborhood feature aggregation on the model computation graph: the first order aggregates node features within a layer to extract single-layer computational characteristics; the second order fuses features from adjacent layers to capture short-range dependencies such as residual connections; and the third order propagates features across multiple attention heads to model long-range interaction patterns. The hardware topology graph attention network uses a multi-head mechanism to concurrently learn resource competition relationships across different dimensions: one head focuses on the data supply bottleneck between memory and compute units, while another head analyzes instruction supply delays between caches and cores. Dual-channel features are deeply coupled through multi-head cross-attention: each attention head optimizes the computational strategy under specific hardware constraints, ultimately synthesizing a multi-objective balanced compression control signal.

[0031] like Figure 1 As shown in the figure, a security protection system for policy execution is established: when hardware feedback indicates that the cache miss rate exceeds a threshold, sparse instructions are automatically frozen and the system falls back to dense computing mode. If low-bit quantization causes an abnormal output distribution, the system immediately switches to a backup high-precision kernel. The configuration version manager saves historical policy snapshots. When the accuracy monitoring module detects a continuous decrease in output confidence, it automatically rolls back to the most recent stable configuration and triggers the online learning process of the decision model to correct policy generation deviations.

[0032] like Figure 2As shown, upon receiving the quantization bit width instruction from the compression control signal, the system initiates a real-time calibration pipeline: First, it analyzes the weight distribution characteristics of the target layer, identifies the value range and discreteness characteristics, and dynamically calculates the scaling factor for the floating-point to fixed-point conversion based on this information. For specific activation functions in the Transformer architecture (such as GELU), a piecewise linear approximation strategy is automatically activated when the bit width is detected to be less than 8 bits, breaking complex operations into multiple linear computation segments to maintain accuracy. The kernel switching module precompiles kernel libraries based on the bit width instruction index. This loads low-bit integer matrix multiply-add kernels for hardware supporting integer tensor cores and enables mixed-precision analog computing for devices supporting only floating-point operations. This process incorporates residual calibration technology, adding error compensation channels to the output nodes of the computation graph. This dynamically corrects deviations by comparing the activation values ​​before and after quantization, ensuring reliable results in extreme compression scenarios. The reconfiguration operation is completed during the computation graph compilation phase. The just-in-time compiler embeds calibration nodes and kernel call instructions into the computation flow, enabling precision mode switching within five milliseconds.

[0033] Furthermore, a sparsification ratio instruction drives the mask generation engine to perform block-structured pruning: the weight matrix is ​​partitioned into sub-blocks that meet the requirements of the hardware accelerator. Low-weight regions are filtered based on a specified ratio threshold, retaining only high-energy blocks for computation. The resulting sparsity pattern is encoded in a compressed sparse row format and submitted to the compute engine with a hardware-recognizable format identifier. For NVIDIA GPUs, the compute flow attribute flag is set to activate the 2:4 sparse tensor cores, while for Ascend NPUs, a DMA engine descriptor is configured to enable the sparse matrix multiply-accumulate unit. During computation execution, the hardware accelerator automatically skips zero-valued blocks based on the format identifier, processing only non-zero blocks and accumulating partial sums. To address the potential loss of cache locality caused by dynamic sparsification, the system adds data prefetch instructions to the memory access layer, preloading adjacent data blocks into the cache based on the non-zero block distribution pattern. This solution reduces the number of matrix multiplications by 40% while reducing the additional control overhead to less than 1% through hardware co-design.

[0034] like Figure 2As shown, the operator scheduling instruction parsing engine maps the policy encoding into a three-stage optimization scheme. The first-stage hardware adaptation layer selects a computational graph partitioning strategy based on the device type. For GPU devices, it enables asynchronous pipeline parallelism, splitting the computational graph into multiple subgraph streams and executing them concurrently on different compute units. For NPU devices, it configures a data stream sharding mode, slicing tensor data based on the hardware array size and enabling multi-core parallel computing. The second-stage memory optimization layer implements differentiated storage management: GPU-bound texture memory accelerates constant data reads, CPUs deploy large page memory to reduce address translation overhead, and mobile chips enable memory compression to reduce data transfer. The third-stage instruction optimization layer invokes device-specific acceleration libraries: activating vectorized dot product kernels on CPUs supporting the AVX-512 instruction set and enabling OpenCL subgroup operations to accelerate reduction computations on Mali GPUs. The scheduler dynamically injects synchronization primitives at runtime to coordinate multi-level parallelism and automatically inserts memory barriers when data dependencies are detected, ensuring computational correctness while maximizing hardware utilization.

[0035] Specifically, the dynamic policy selector establishes an incremental learning loop during service operation: it receives hardware metric change data collected by the feedback loop in real time and converts it into policy evaluation samples for input into the online training module. This module adopts a two-stage update strategy: short-term incremental learning retains state features from the last two hundred inference cycles in a sliding window manner, and fine-tunes the graph neural network edge weight parameters through lightweight gradient descent to quickly adapt to hardware state drift; the long-term knowledge distillation phase is launched monthly, migrating the online model's policy generation capabilities to the backup model, and ensuring policy stability through output distribution alignment constraints. The model update process incorporates a secure isolation design: test replicas run in parallel in a shadow environment, and hot switching to the production model is triggered only when their policy achieves less than one percent accuracy loss and a latency optimization rate exceeding 15 percent on the validation set. This mechanism enables continuous system evolution after deployment. For example, when a new computing unit is added to the cluster, the model can autonomously learn the optimal scheduling policy within 24 hours.

[0036] The feedback loop incorporates a multi-layered fault-fault mechanism. When the hardware probe detects a transient increase in memory usage exceeding a safety threshold, it immediately sends a high-priority interrupt signal to the policy selector, triggering a three-level emergency response. The first level forces the quantization bit width of all layers to a safe level, disabling sparsification and reverting to dense computing mode. The second level initiates computational resource isolation, pausing low-priority inference tasks to free up hardware resources. The third level activates model segmented offloading, migrating some computational layers to collaborative devices. A real-time monitoring agent is deployed in the accuracy assurance layer to compare the output distribution before and after reconfiguration. If a key metric (such as classification confidence) continuously degrades beyond a tolerance threshold, it automatically rolls back to the last stable policy snapshot and generates an alert log. The system maintains a track of compression policy versions for each inference task, supports point-in-time restoration of historical configurations, and triggers a targeted retraining process for the decision model to correct parameter deviations. This design ensures service availability during sudden load spikes or hardware failures, limiting the impact of anomalies to milliseconds.

[0037] The large-model inference performance dynamic optimization and hardware-aware compression method of the present invention extracts the complexity characteristics of the input data in real time and synchronously monitors the hardware load status, driving the dynamic strategy selector to generate a joint control signal for quantization bit width, sparsification ratio and operator scheduling; based on the signal, the model weights and activation values ​​are reconfigured at runtime to achieve dynamic adaptation of computing accuracy and hardware resources; and the hardware indicator feedback of the inference process is used to form a closed-loop optimization, so that the large model automatically maintains optimal performance in changing scenarios.

[0038] Therefore, the dynamic optimization of large model inference efficiency and hardware-aware compression method of the present invention can solve the problems of low resource utilization, delay fluctuation and energy efficiency imbalance caused by static compression methods in dynamic input and heterogeneous hardware environments.

[0039] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.

Claims

1. Dynamic optimization of large model inference performance and hardware-aware compression method, characterized by: include: S1: Extracting feature vectors of input data in real time to generate an input complexity signal representing computational complexity. Extracting feature vectors of input data in S1 specifically includes analyzing the sequence length and attention distribution dispersion of the input text in real time through a lightweight convolutional network to generate a feature vector containing hierarchical semantic density information. The input complexity signal is formed by fusing sequence length features and attention entropy features through a gated recurrent unit. This signal dynamically reflects the difference in theoretical computational effort caused by different input samples at each computational layer of the model. S2: Synchronously monitor the hardware resource indicators of the operating platform and generate hardware status signals reflecting real-time load. The hardware resource indicators include memory occupancy, computing unit utilization, and power consumption data. S3: Input the input complexity signal and the hardware status signal into a dynamic strategy selector, and generate a compression control signal through a pre-trained decision model, where the signal includes a combination instruction of quantization bit width, sparsification ratio, and operator scheduling strategy; S4: According to the compression control signal, dynamic reconfiguration operations are performed on the large model weights and activation values ​​of the current inference task, including switching floating-point or fixed-point calculation modes based on quantization bit width instructions, activating the structured mask of the corresponding layer based on sparse ratio instructions, and adapting the hardware acceleration kernel based on operator scheduling instructions. When switching the calculation mode based on the quantization bit width instruction in S4, the pre-compiled integer calculation kernel or mixed precision calculation kernel is dynamically loaded according to the bit width parameter specified by the control signal, and a real-time calibration node is inserted in the calculation graph to compensate for the quantization error. The fixed-point calculation mode adopts a symmetric quantization strategy to map the floating-point weights to the scaling kernel. The integer representation of the factor; the activation process of the structured mask includes generating a block sparsity pattern that meets the requirements of the hardware acceleration unit according to the sparsification ratio instruction, dynamically masking the low-weight value area of ​​the specified ratio in the weight matrix, and submitting a sparse matrix compression format identifier to the computing engine to trigger a dedicated computing pipeline; the specific implementation of the operator scheduling instruction adaptation hardware acceleration kernel is to select the computing graph partitioning strategy and memory allocation scheme according to the control signal, enable the asynchronous pipeline parallel mechanism for the graphics processor, start the data flow sharding computing mode for the neural network processor, and bind the large page memory prefetch strategy for the central processing unit; S5: Use the reconfigured large model to perform inference calculations, and during the calculation process, feed back hardware resource indicators to S2 in real time, forming a closed-loop optimization link.

2. The large model inference efficiency dynamic optimization and hardware-aware compression method according to claim 1 is characterized in that: The generation process of the hardware status signal in S2 includes establishing a direct data channel with the underlying hardware driver, periodically collecting the graphics processor's video memory bandwidth utilization, the tensor core idle cycle ratio, and the on-chip cache miss rate, and mapping the heterogeneous hardware indicators into a load evaluation coefficient of a unified dimension through normalization processing. The power consumption data is obtained by reading the real-time current-voltage product of the onboard power management chip.

3. The large model inference efficiency dynamic optimization and hardware-aware compression method according to claim 1 is characterized in that: The pre-trained decision model in S3 adopts a dual-channel graph neural network architecture. The first channel analyzes the spatiotemporal correlation features of the input complexity signal, and the second channel learns the fluctuation pattern of the hardware status signal in the time series. After fusing the dual-channel features through the cross-attention mechanism, the compression control signal is output. The training process of this model uses a reinforcement learning framework with inference delay and energy consumption ratio as the joint reward function.

4. The large model inference efficiency dynamic optimization and hardware-aware compression method according to claim 1 is characterized in that: During runtime, the decision model continuously receives hardware resource indicators fed back by S5 and updates the edge weight parameters of the graph neural network through an online incremental learning mechanism. The model update process uses a sliding window mechanism to retain historical state features and ensures strategy stability through knowledge distillation constraints.

5. The large model inference efficiency dynamic optimization and hardware-aware compression method according to claim 1 is characterized in that: The feedback process of S5 includes an exception handling mechanism. When it is detected that the hardware resource indicator suddenly exceeds the safety threshold, an interrupt signal is immediately sent to the dynamic strategy selector to trigger the emergency strategy, forcing the quantization bit width of the compression control signal to be increased to the safety level and disabling the sparse operation.

6. The large model inference efficiency dynamic optimization and hardware-aware compression method according to claim 1 is characterized in that: The closed-loop optimization link is implemented in the edge computing scenario by collecting multi-node hardware status signals through a distributed monitoring agent, and generating a global compression control signal by a central policy selector, where dynamic reconfiguration operations synchronously coordinate the model segmentation calculation and result fusion process among multiple devices.

Citation Information

Patent Citations

  • Model quantitative reasoning acceleration method and device, equipment and medium

    CN120086355A

  • Multi-model collaborative operation method based on large model efficient training

    CN120295784A