Hardware accelerated dynamic sparsity optimization engine
Through the hardware-accelerated dynamic sparse optimization engine, real-time detection and prediction of sparse changes, dynamic adjustment of sparse storage and computing paths, the problem of insufficient resource utilization of sparse matrix in neural network models is solved, and efficient computing and storage optimization is achieved.
Patent Information
- Application Number
- CN202510328899.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the sparseness assumption of sparse matrices in neural network models is fixed and cannot be adjusted dynamically, resulting in insufficient resource utilization and inability to adapt to changes in different stages or task scenarios.
It adopts a hardware-accelerated dynamic sparse optimization engine, including a sparse dynamic detection module, a dynamic sparse storage optimization module, a sparse computing path optimization module and a hardware firmware collaboration module, and realizes sparse dynamic adjustment and optimization through hardware sensors, trend prediction units and firmware strategy engines.
Improve computing efficiency, improve computing performance by more than 30%, and reduce redundant memory access by 15%-20%, optimizing resource utilization.
Smart Images

Figure CN120256107A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural processing units, and specifically relates to a hardware-accelerated dynamic sparsity optimization engine. Background Art
[0002] In the fields of deep learning and artificial intelligence, sparse matrices play an important role in neural network models. Especially in weight matrices, sparsity (i.e., the proportion of zero values) is widespread, which can significantly reduce storage and computational requirements.
[0003] However, existing technologies usually assume fixed sparsity and do not support runtime dynamic adjustment. This static approach leads to insufficient resource utilization and cannot adapt to changes in sparsity in different stages or different task scenarios of the model.
[0004] In order to optimize hardware resource utilization and improve model performance, there is an urgent need for an innovative mechanism that can dynamically identify and adjust sparsity and is based on hardware acceleration to maximize the storage and computational efficiency of sparse matrices. Summary of the Invention
[0005] The purpose of the present invention is to provide a hardware-accelerated dynamic sparsity optimization engine to solve the above-mentioned problems.
[0006] The technical solution adopted by the present invention is as follows: a hardware-accelerated dynamic sparsity optimization engine, which includes: a sparsity dynamic detection module, a dynamic sparse storage optimization module, a sparse computing path optimization module, a hardware-firmware cooperation module, and a system bus and interconnection module;
[0007] The internal of the sparsity dynamic detection module is provided with a hardware sensor unit, a threshold adaptive management unit, and a trend prediction unit;
[0008] The internal of the dynamic sparse storage optimization module is provided with a multi-format storage controller, a storage compression unit, and a cache management unit;
[0009] The internal of the sparse computing path optimization module is provided with a zero-value skipping unit, a block parallel accelerator, and a dynamic computing scheduler;
[0010] The internal of the hardware-firmware cooperation module is provided with a firmware policy engine, a real-time monitoring and feedback unit, and a hardware abstraction layer;
[0011] The output end of the sparsity dynamic detection module is connected to the input end of the dynamic sparse storage optimization module through a low-latency data channel to drive dynamic adjustment of the storage format;
[0012] The output end of the dynamic sparse storage optimization module is connected to the input end of the sparse computing path optimization module through the data loading interface of the block parallel accelerator, providing optimized non-zero data blocks for the computing path;
[0013] The output end of the sparse computing path optimization module is fed back to the input end of the hardware-firmware cooperation module through the control interface of the hardware abstraction layer, completing the closed-loop optimization of computing efficiency and resource occupancy;
[0014] The output end of the hardware-firmware cooperation module is respectively connected to the threshold adaptive management unit of the sparsity dynamic detection module and the zero-value skipping unit of the sparse computing path optimization module through the control signal synchronization unit, realizing global policy synchronization;
[0015] The system bus and interconnection module internally integrate a dedicated data channel, directly connecting the output end of the storage compression unit to the input end of the block parallel accelerator, ensuring that high-sparse data is quickly transmitted to the computing unit in a compressed format. At the same time, it sends the storage load status to the cache management unit through the status signal interface of the hardware sensor unit, dynamically allocating cache resources.
[0016] In a preferred embodiment, the hardware sensor unit is composed of a dedicated sensor array and a high-speed sampling circuit. Through hardware-level parallel scanning technology, it can collect the distribution density, spatial continuity, and dynamic change characteristics of zero values in the weight matrix in real time. The sensor adopts a multi-channel design, and each channel independently monitors a specific block area of the matrix and generates a sparsity statistical report.
[0017] In a preferred embodiment, the threshold adaptive management unit has a built-in multi-level policy library, pre-setting threshold rules for training, inference, and edge devices. The unit dynamically calculates the optimal threshold by analyzing the real-time data reported by the hardware sensor and combining the task priorities issued by the firmware policy engine. This unit also supports developers to customize the threshold adjustment algorithm through the firmware interface to achieve scenario adaptation.
[0018] In a preferred embodiment, the trend prediction unit receives the historical sparsity data of the hardware sensor and predicts the areas where sparsity may mutate in the future computing cycle. The prediction result notifies the dynamic sparse storage optimization module in advance through a dedicated control signal, triggering the preparation for storage format switching and avoiding the impact of format conversion delay on computing performance.
[0019] In a preferred embodiment, the multi-format storage controller is the core of the dynamic sparse storage optimization module, supporting hardware-level switching of multiple formats. The controller selects the storage format through register configuration and integrates a dedicated format conversion accelerator to complete the conversion from a dense matrix to the target sparse format in an extremely short time. The controller also provides a format compatibility check function to ensure that the switched storage format matches the hardware architecture of the downstream computation path optimization module;
[0020] The storage compression unit is designed for high-sparse matrices and uses a hybrid compression algorithm to further reduce storage occupancy. First, the positions of non-zero blocks are marked by a hardware-level bitmask, and then differential encoding is applied to the data within the non-zero blocks. The compressed data is transmitted to the cache management unit through a dedicated bandwidth optimization channel;
[0021] The cache management unit analyzes the data access patterns of the block parallel accelerator, preferentially caches non-zero blocks with high-frequency access, and uses the LRU algorithm to eliminate low-frequency data. The unit is linked with the hardware sensor unit to obtain real-time sparsity change information and release invalid cache space in advance. Under a specific storage format, the unit supports the prefetch mechanism for non-zero blocks to improve the cache hit rate through the sparsity trend provided by the prediction unit.
[0022] In a preferred embodiment, the zero-value skipping unit is a basic component of the sparse computation path optimization module, and its core function is to quickly identify and filter invalid computation tasks through a hardware-level circuit. The unit consists of a multi-stage comparator array and a gated logic circuit, which can scan the zero-value distribution of the input matrix in real time before the data enters the computation unit. When it detects that a certain row or a certain block area is all zero values, the unit immediately sends a masking signal to the computation pipeline to directly skip the multiplication and accumulation operations in this area, significantly reducing redundant computations.
[0023] In a preferred embodiment, the block parallel accelerator is the core execution unit for sparse computation path optimization. By dividing the sparse matrix into independent computation blocks and enabling hardware-level parallel processing, it maximizes the utilization rate of computation resources. Multiple computation cluster clusters are deployed inside the accelerator, and each cluster contains a set of multiplication and accumulation units, a local register file, and a data prefetch controller. When a non-zero data block is loaded from the cache management unit, the accelerator only activates the computation clusters containing non-zero values according to the bitmask information of the block sparse coding, and the rest of the clusters enter the low-power sleep state.
[0024] In a preferred embodiment, the dynamic computation scheduler is responsible for the intelligent scheduling and resource allocation of global computation tasks, and its design goal is to achieve an optimal balance among the storage format, the hardware resource status, and the real-time computation requirements. The scheduler has a built-in multi-stage priority queue, which dynamically sorts the computation tasks according to the sparsity of the data blocks, the storage compression rate, and the urgency of the tasks.
[0025] In a preferred embodiment, the firmware policy engine is a policy management center with collaborative optimization of hardware and software, driving global optimization decisions through a preset rule library and an adaptive learning mechanism. The engine has multiple predefined policy modes built-in, including high-throughput mode, low-power mode, balanced mode, and extreme compression mode for edge devices;
[0026] The real-time monitoring and feedback unit collects runtime data in all dimensions through a distributed sensor network and hardware counters. The monitoring metrics cover all aspects of the computing link: including the response latency of the sparsity detection module, storage compression rate, cache hit rate, zero-value skip ratio, computing unit utilization, and system power consumption. After being aggregated and compressed, the data is transmitted to an external management platform through a standardized interface provided by the hardware abstraction layer, supporting visual analysis and historical trend backtracking. The unit has an intelligent alarm mechanism built-in. When an abnormal event is detected, it immediately triggers an interrupt signal and pushes an alarm code to the firmware policy engine to drive system degradation or recovery operations. In addition, the unit provides a debug mode, allowing developers to inject test data and trace the entire link processing process to accelerate fault location and performance tuning.
[0027] In a preferred embodiment, the hardware abstraction layer is the core interface layer that shields the underlying hardware differences and realizes cross-platform compatibility, providing a unified hardware operation view for the upper-layer firmware. This layer defines a standardized instruction set, covering key operations such as storage format switching, compression level setting, computing task dispatching, and status query. The "STOR_FORMAT_SWITCH" instruction is used to trigger the storage controller to switch formats, or the "TASK_PRIORITY_SET" instruction is used to adjust the task weights of the computing scheduler. The hardware abstraction layer abstracts physical resources into a virtual resource pool, enabling the firmware policy engine to manage resources without being aware of the specific hardware architecture. In terms of security, this layer implements hierarchical permission control, restricts the execution of illegal instructions, and prevents conflicts between computing tasks and storage formats through a data verification mechanism. In addition, the hardware abstraction layer supports dynamic expansion and can adapt to the access requirements of future new computing units or storage media.
[0028] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:
[0029] 1. In the present invention, the trend prediction unit predicts the zero-value distribution trend in the future calculation cycle by analyzing the historical data of sparsity changes in real time and combining with a lightweight time series model, enabling the system to anticipate sparsity mutations in advance and trigger optimization actions. This innovative mechanism completely changes the passive mode of traditional static sparsity processing, upgrading the optimization decision from "post-response" to "pre-preparation". In dynamic sparse training, when it is predicted that the sparsity of a certain weight matrix block will significantly increase in the next stage, the system switches to the block sparse storage format in advance and pre-allocates computing resources to avoid delays caused by format conversion or resource contention. At the same time, the firmware policy engine dynamically loads preset policies or developer-defined rules to adjust the parameters and optimization thresholds of the prediction model in real time, ensuring a high degree of adaptation of the prediction accuracy to the task requirements. This closed-loop coordination of prediction and policy enables the system to improve the computing efficiency by more than 30% and reduce redundant memory access by 15%-20% in resource-sensitive scenarios such as edge devices and high-concurrency inference.
[0030] 2. In the present invention, the hardware abstraction layer shields the differences of different hardware architectures through standardized interfaces and virtualized resource management, enabling the block parallel accelerator to seamlessly adapt to diverse computing units and storage media. Based on the unified instruction set provided by the hardware abstraction layer, the block parallel accelerator divides the sparse matrix into dynamically adjustable block granularities and activates only the computing clusters corresponding to non-zero blocks to achieve on-demand allocation of computing resources. In inference tasks with extremely high sparsity, the accelerator quickly calls block sparse encoded data through the hardware abstraction layer, executes non-zero block calculations in parallel in units of 8x8 blocks, and skips all-zero blocks at the same time, increasing the computing throughput to 4-5 times that of traditional dense computing. This combined innovation not only significantly reduces the cross-platform porting cost for developers, but also optimizes the system energy efficiency ratio to more than 1.8 times that of similar solutions through deep coordination of hardware-level parallelism and zero-value skipping, providing core support for efficient deployment in data centers and edge devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is the overall system block diagram of the present invention;
[0032] Figure 2 It is the system block diagram of the dynamic sparsity detection module of the present invention.
[0033] Figure 3 It is the system block diagram of the dynamic sparse storage optimization module of the present invention.
[0034] Figure 4 It is the system block diagram of the sparse computing path optimization module of the present invention.
[0035] Figure 5 It is the system block diagram of the hardware-firmware cooperation module of the present invention. Detailed Implementation Manner
[0036] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] Embodiment:
[0038] Refer to Figures 1-5 , a hardware-accelerated dynamic sparsity optimization engine, which includes: a sparsity dynamic detection module, a dynamic sparse storage optimization module, a sparse computing path optimization module, a hardware-firmware cooperation module, and a system bus and interconnection module;
[0039] Inside the sparsity dynamic detection module, there are a hardware sensor unit, a threshold adaptive management unit, and a trend prediction unit;
[0040] Inside the dynamic sparse storage optimization module, there are a multi-format storage controller, a storage compression unit, and a cache management unit;
[0041] Inside the sparse computing path optimization module, there are a zero-value skipping unit, a block parallel accelerator, and a dynamic computing scheduler;
[0042] Inside the hardware-firmware cooperation module, there are a firmware policy engine, a real-time monitoring and feedback unit, and a hardware abstraction layer;
[0043] The output end of the sparsity dynamic detection module is connected to the input end of the dynamic sparse storage optimization module through a low-latency data channel to drive the dynamic adjustment of the storage format;
[0044] The output end of the dynamic sparse storage optimization module is connected to the input end of the sparse computing path optimization module through the data loading interface of the block parallel accelerator to provide optimized non-zero data blocks for the computing path;
[0045] The output end of the sparse computing path optimization module is fed back to the input end of the hardware-firmware cooperation module through the control interface of the hardware abstraction layer to complete the closed-loop optimization of computing efficiency and resource occupancy;
[0046] The output end of the hardware-firmware cooperation module is respectively connected to the threshold adaptive management unit of the sparsity dynamic detection module and the zero-value skipping unit of the sparse computing path optimization module through a control signal synchronization unit to achieve global policy synchronization;
[0047] The system bus and the interconnection module internally integrate a dedicated data channel, which directly connects the output end of the storage compression unit and the input end of the block parallel accelerator, ensuring that high-sparse data is quickly transmitted to the computing unit in a compressed format. At the same time, the storage load status is sent to the cache management unit through the status signal interface of the hardware sensor unit to dynamically allocate cache resources.
[0048] The hardware sensor unit is the core component of the sparsity dynamic detection module and is composed of a dedicated sensor array and a high-speed sampling circuit. This unit uses hardware-level parallel scanning technology to collect the distribution density, spatial continuity, and dynamic change characteristics of zero values in the weight matrix in real time. The sensor adopts a multi-channel design, and each channel independently monitors a specific block area of the matrix and generates a sparsity statistical report. The data is transmitted to the threshold adaptive management unit through a low-latency bus to trigger storage and computing optimization decisions. The hardware sensor supports configurable sampling frequencies and can complete real-time monitoring without interrupting the computing task, making it suitable for dynamic sparse training and inference scenarios with high-frequency sparsity changes.
[0049] The threshold adaptive management unit is responsible for dynamically adjusting the sparsity determination threshold according to task requirements and resource status. This unit has a built-in multi-level policy library with threshold rules preset for training, inference, and edge devices. The unit dynamically calculates the optimal threshold by analyzing the real-time data reported by the hardware sensor and combining the task priorities issued by the firmware policy engine. This unit also supports developers to customize the threshold adjustment algorithm through the firmware interface to achieve scenario adaptation.
[0050] The trend prediction unit is based on a time series analysis model to make short-term predictions on sparsity changes. This unit receives the historical sparsity data of the hardware sensor and predicts the areas where sparsity may mutate in the future computing cycles. The prediction results are notified in advance to the dynamic sparse storage optimization module through dedicated control signals to trigger the preparation for storage format switching and avoid the impact of format conversion delay on computing performance. The prediction unit uses fixed-point arithmetic friendly to hardware, and the prediction accuracy error is controlled within a reasonable range, making it suitable for fast-changing scenarios in dynamic sparse training.
[0051] The multi-format storage controller is the core of the dynamic sparse storage optimization module and supports hardware-level switching of multiple formats. The controller selects the storage format through register configuration and integrates a dedicated format conversion accelerator, which can complete the conversion from a dense matrix to the target sparse format in an extremely short time. The controller also provides a format compatibility check function to ensure that the switched storage format matches the hardware architecture of the downstream computing path optimization module;
[0052] The storage compression unit is designed for highly sparse matrices and uses a hybrid compression algorithm to further reduce storage usage. The unit first marks the location of non-zero blocks through a hardware-level bit mask, and then uses differential encoding for the data in the non-zero blocks. The compressed data is transmitted to the cache management unit through a dedicated bandwidth optimization channel. The unit also supports dynamic compression level adjustment, enabling the highest compression rate mode when the edge device memory is tight;
[0053] The cache management unit is responsible for dynamically allocating on-chip cache resources based on computing requirements and storage formats. The unit analyzes the data access pattern of the block parallel accelerator, prioritizes caching of high-frequency accessed non-zero blocks, and uses the LRU algorithm to eliminate low-frequency data. The unit works in conjunction with the hardware sensor unit to obtain sparsity change information in real time and release invalid cache space in advance. Under specific storage formats, the unit supports a pre-fetch mechanism for non-zero blocks, and improves the cache hit rate by predicting the sparsity trend provided by the unit.
[0054] The zero-value skip unit is a basic component of the sparse computing path optimization module. Its core function is to quickly identify and filter invalid computing tasks through hardware-level circuits. The unit consists of a multi-level comparator array and a gated logic circuit, which can scan the zero-value distribution of the input matrix in real time before the data enters the computing unit. When it is detected that a row or a block area is full of zero values, the unit immediately sends a shielding signal to the computing pipeline to directly skip the multiplication and accumulation operations in the area, significantly reducing redundant calculations. The zero-value skip unit supports flexible skip granularity configuration, such as filtering by row, column or custom block size to adapt to different sparse distribution patterns. The unit adopts a pipeline design to ensure that zero-value detection is executed synchronously with the computing task, and the skip decision delay is controlled at the nanosecond level, which has almost no impact on the overall computing throughput. In high-sparseness scenarios, the unit can reduce more than 90% of invalid calculations, significantly reduce power consumption and improve energy efficiency.
[0055] The block parallel accelerator is the core execution unit for sparse computing path optimization. It maximizes computing resource utilization by dividing sparse matrices into independent computing blocks and enabling hardware-level parallel processing. Multiple computing clusters are deployed inside the accelerator, each of which contains a set of multiplication and accumulation units, a local register file, and a data prefetch controller. When a non-zero data block is loaded from the cache management unit, the accelerator only activates the computing cluster containing non-zero values according to the bit mask information of the block sparse coding, and the remaining clusters enter a low-power sleep state. The block strategy supports dynamic adjustment. For example, when the sparse distribution is uneven, the default 8x8 block is automatically switched to 4x4 or 2x2 block to match the clustering characteristics of non-zero values. The accelerator also integrates data dependency analysis functions, which can predict the correlation between computing blocks in advance, optimize the task scheduling order, and reduce memory access conflicts. In typical reasoning tasks, the accelerator can increase the sparse matrix computing throughput by 3-5 times.
[0056] The dynamic computing scheduler is responsible for the intelligent scheduling and resource allocation of global computing tasks. Its design goal is to achieve an optimal balance among the storage format, the status of hardware resources, and the real-time computing requirements. The scheduler has a built-in multi-level priority queue, which dynamically sorts computing tasks according to the sparsity of data blocks, the storage compression ratio, and the urgency of tasks. For example, for data blocks using the block sparse format with a compression ratio higher than 80%, high-bandwidth channels are preferentially allocated and parallel computing is started; for data blocks with low sparsity or uncompressed data, serial computing is used to save resources. The scheduler cooperates deeply with the zero-value skipping unit, receives zero-value distribution information in real time, and dynamically adjusts the startup order of the computing pipeline. In the scenario of multi-task concurrency on edge devices, the scheduler introduces a preemptive task management mechanism, allowing high-priority inference tasks to interrupt low-priority training tasks to ensure the quality of service for critical services. In addition, the scheduler provides a fault tolerance function. When it detects that the computing unit is overloaded or the data is abnormal, it automatically switches to the degraded mode and reports the error log.
[0057] The firmware policy engine is the policy management center for the collaborative optimization of hardware and software, driving global optimization decisions through a preset rule library and an adaptive learning mechanism. The engine has a variety of predefined policy modes built in, including high-throughput mode, low-power mode, balanced mode, and extreme compression mode for edge devices. Developers can customize policy parameters through standardized API interfaces, such as setting the sparsity detection frequency, the storage format switching threshold, or the computing task priority weight. The policy engine receives runtime metrics from the monitoring unit in real time and dynamically adjusts the policy according to the task type - for example, enabling a dynamic threshold learning algorithm in training tasks to automatically optimize the sparsity determination rule according to the gradient change trend; in inference tasks, the joint optimization of storage compression and zero-value skipping is forced. The engine also supports the policy hot-loading function, allowing seamless switching of policy configurations during system operation to avoid service interruption;
[0058] The real-time monitoring and feedback unit is the central system for the health status and performance analysis of the engine, collecting runtime data in all dimensions through a distributed sensor network and hardware counters. The monitoring metrics cover all aspects of the computing link: including the response latency of the sparsity detection module, the storage compression ratio, the cache hit rate, the zero-value skipping ratio, the utilization rate of the computing unit, and the system power consumption. After being aggregated and compressed, the data is transmitted to the external management platform through the standardized interface provided by the hardware abstraction layer to support visual analysis and historical trend backtracking. The unit has an intelligent alarm mechanism. When an abnormal event is detected (such as a sparsity mutation exceeding the threshold, frequent switching of the storage format, or continuous overload of the computing unit), an interrupt signal is immediately triggered and an alarm code is pushed to the firmware policy engine to drive system degradation or recovery operations. In addition, the unit provides a debug mode, allowing developers to inject test data and trace the whole-link processing process to accelerate fault location and performance tuning.
[0059] The hardware abstraction layer is the core interface layer that shields the underlying hardware differences and enables cross-platform compatibility, providing a unified hardware operation view for the upper-layer firmware. This layer defines a standardized instruction set covering key operations such as storage format switching, compression level setting, computing task dispatching, and status query. The storage controller switches the format by triggering the "STOR_FORMAT_SWITCH" instruction, or the task weight of the computing scheduler is adjusted by the "TASK_PRIORITY_SET" instruction. The hardware abstraction layer abstracts physical resources (such as computing clusters, cache capacity, sensor channels) into a virtual resource pool, enabling the firmware policy engine to manage resources without perceiving the specific hardware architecture. In terms of security, this layer implements hierarchical permission control, restricts the execution of illegal instructions, and prevents conflicts between computing tasks and storage formats through a data verification mechanism. In addition, the hardware abstraction layer supports dynamic expansion and can adapt to the access requirements of future new computing units or storage media.
[0060] As can be seen from the above: In the present invention, the trend prediction unit predicts the zero-value distribution trend in the future computing cycle by analyzing the historical data of the sparsity change in real time and combining with a lightweight time series model, enabling the system to anticipate the sparsity mutation in advance and trigger optimization actions. This innovative mechanism completely changes the passive mode of traditional static sparsity processing, upgrading the optimization decision from "post-response" to "preparation in advance". For example, in dynamic sparse training, when it is predicted that the sparsity of a certain weight matrix block will increase significantly in the next stage, the system switches to the block sparse storage format in advance and pre-allocates computing resources to avoid delays caused by format conversion or resource contention. At the same time, the firmware policy engine dynamically loads preset policies or developer-defined rules to adjust the parameters and optimization thresholds of the prediction model in real time, ensuring a high degree of adaptation between the prediction accuracy and the task requirements. This closed-loop collaboration between prediction and policy enables the system to improve the computing efficiency by more than 30% and reduce the redundant memory access by 15%-20% in resource-sensitive scenarios such as edge devices and high-concurrency inference.
[0061] In the present invention, the hardware abstraction layer manages virtualized resources through a standardized interface, shielding the differences in different hardware architectures, enabling the block parallel accelerator to seamlessly adapt to diverse computing units and storage media. Based on the unified instruction set provided by the hardware abstraction layer, the block parallel accelerator divides the sparse matrix into dynamically adjustable block granularities, and only activates the computing clusters corresponding to the non-zero blocks to achieve on-demand allocation of computing resources. For example, in inference tasks with extremely high sparsity, the accelerator quickly invokes block sparse encoded data through the hardware abstraction layer, executes non-zero block calculations in parallel in units of 8x8 blocks, and skips all-zero blocks at the same time, increasing the computing throughput to 4-5 times that of traditional dense computing. This combined innovation not only significantly reduces the cross-platform porting cost for developers, but also optimizes the system energy efficiency ratio to more than 1.8 times that of similar solutions through the deep coordination of hardware-level parallelism and zero-value skipping, providing core support for the efficient deployment of data centers and edge devices.
[0062] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article, or device comprising the element.
[0063] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hardware-accelerated dynamic sparsity optimization engine, characterized in that: The engine includes: a sparsity dynamic detection module, a dynamic sparse storage optimization module, a sparse computing path optimization module, a hardware-firmware cooperation module, and a system bus and interconnection module; Inside the sparsity dynamic detection module, there are a hardware sensor unit, a threshold adaptive management unit, and a trend prediction unit; Inside the dynamic sparse storage optimization module, there are a multi-format storage controller, a storage compression unit, and a cache management unit; Inside the sparse computing path optimization module, there are a zero-value skip unit, a block parallel accelerator, and a dynamic computing scheduler; Inside the hardware-firmware cooperation module, there are a firmware policy engine, a real-time monitoring and feedback unit, and a hardware abstraction layer; The output end of the sparsity dynamic detection module is connected to the input end of the dynamic sparse storage optimization module through a low-latency data channel, driving the dynamic adjustment of the storage format; The output end of the dynamic sparse storage optimization module is connected to the input end of the sparse computing path optimization module through the data loading interface of the block parallel accelerator, providing optimized non-zero data blocks for the computing path; The output end of the sparse computing path optimization module is fed back to the input end of the hardware-firmware cooperation module through the control interface of the hardware abstraction layer, completing the closed-loop optimization of computing efficiency and resource occupancy; The output end of the hardware-firmware cooperation module is respectively connected to the threshold adaptive management unit of the sparsity dynamic detection module and the zero-value skip unit of the sparse computing path optimization module through a control signal synchronization unit, realizing global policy synchronization; The system bus and interconnection module internally integrates a dedicated data channel, directly connecting the output end of the storage compression unit to the input end of the block parallel accelerator, ensuring that high-sparse data is quickly transmitted to the computing unit in a compressed format. At the same time, the storage load status is sent to the cache management unit through the status signal interface of the hardware sensor unit to dynamically allocate cache resources.
2. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, wherein: The hardware sensor unit is composed of a dedicated sensor array and a high-speed sampling circuit; through hardware-level parallel scanning technology, it can collect the distribution density, spatial continuity, and dynamic change characteristics of zero values in the weight matrix in real time; The sensor adopts a multi-channel design, and each channel independently monitors a specific block area of the matrix and generates a sparsity statistical report.
3. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, wherein: The threshold adaptive management unit has a built-in multi-level policy library, with threshold rules preset for training, inference, and edge devices; the unit dynamically calculates the optimal threshold by analyzing the real-time data reported by the hardware sensor and combining the task priorities issued by the firmware policy engine; the unit also supports developers to customize the threshold adjustment algorithm through the firmware interface to achieve scenario adaptation.
4. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, wherein: The trend prediction unit receives the historical sparsity data of the hardware sensor and predicts the areas where sparsity may mutate in the future computing cycle; the prediction result is notified to the dynamic sparse storage optimization module in advance through a dedicated control signal, triggering the preparation for storage format switching and avoiding the impact of format conversion delay on computing performance.
5. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, characterized in that: The multi-format storage controller is the core of the dynamic sparse storage optimization module and supports hardware-level switching of multiple formats. The controller selects the storage format through register configuration and integrates a dedicated format conversion accelerator to complete the conversion of dense matrix to target sparse format in a very short time. The controller also provides a format compatibility check function to ensure that the switched storage format matches the hardware architecture of the downstream computational path optimization module. The storage compression unit is designed for highly sparse matrices and uses a hybrid compression algorithm to further reduce storage occupancy; first, the position of the non-zero block is marked by a hardware-level bit mask, and then the data in the non-zero block is differentially encoded; The compressed data is transmitted to the cache management unit through a dedicated bandwidth-optimized channel; The cache management unit analyzes the data access mode of the block parallel accelerator, prioritizes caching high-frequency accessed non-zero blocks, and uses the LRU algorithm to eliminate low-frequency data; the unit is linked with the hardware sensor unit to obtain sparsity change information in real time and release invalid cache space in advance.
6. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, wherein: The zero-value skip unit is a basic component of the sparse computing path optimization module, and its core function is to quickly identify and filter invalid computing tasks through hardware-level circuits; the unit is composed of a multi-level comparator array and a gated logic circuit, and can scan the zero-value distribution of the input matrix in real time before the data enters the computing unit.
7. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, wherein: The block parallel accelerator is the core execution unit for sparse computing path optimization, which maximizes computing resource utilization by dividing sparse matrices into independent computing blocks and enabling hardware-level parallel processing; Multiple computing clusters are deployed inside the accelerator, each of which contains a set of multiplication and accumulation units, a local register file, and a data prefetch controller; when a non-zero data block is loaded from the cache management unit, the accelerator only activates the computing cluster containing non-zero values according to the bit mask information of the block sparse coding, and the remaining clusters enter a low-power sleep state.
8. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, characterized in that: The dynamic computing scheduler is responsible for the intelligent scheduling and resource allocation of global computing tasks, and its design goal is to achieve the optimal balance between storage format, hardware resource status and real-time computing requirements; The scheduler has a built-in multi-level priority queue that dynamically sorts computing tasks based on the sparsity of data blocks, storage compression rate, and task urgency.
9. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, wherein: The firmware policy engine is a policy management center for hardware and software collaborative optimization, driving global optimization decisions through a preset rule base and an adaptive learning mechanism; The engine has multiple predefined strategy modes built in, including high throughput mode, low power mode, balanced mode, and extreme compression mode for edge devices; The real-time monitoring and feedback unit collects runtime data in all dimensions through a distributed sensor network and hardware counters; the monitoring metrics cover all links of the computing chain, including the response latency of the sparsity detection module, storage compression rate, cache hit rate, zero-value skip ratio, computing unit utilization rate, and system power consumption; after the data is aggregated and compressed, it is transmitted to the external management platform through the standardized interface provided by the hardware abstraction layer, supporting visual analysis and historical trend backtracking; the unit is built with an intelligent alarm mechanism, which immediately triggers an interrupt signal and pushes an alarm code to the firmware policy engine when an abnormal event is detected, driving system degradation or recovery operations; in addition, the unit provides a debug mode that allows developers to inject test data and track the full-link processing process, accelerating fault location and performance tuning.
10. The hardware-accelerated dynamic sparsity optimization engine according to claim 1, wherein: The hardware abstraction layer is the core interface layer that shields the underlying hardware differences and realizes cross-platform compatibility, providing a unified hardware operation view for the upper-layer firmware; this layer defines a standardized instruction set covering key operations such as storage format switching, compression level setting, computing task dispatching, and status query.