AI voice chip ultra-low power consumption design method based on dynamic multi-domain power consumption regulation and control
By adopting a multi-level hierarchical clock gating architecture and dynamic power domain management on the AI voice chip, combined with voice data cache compression and minimum addressable unit grouping optimization, the problems of dynamic energy efficiency management being too coarse, real-time wake-up and leakage suppression in the existing technology are solved, and the two-way optimization of dynamic energy efficiency management and timing accuracy is achieved.
Patent Information
- Application Number
- CN202510319363.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
AI Technical Summary
The existing AI voice chips have too rough particle size in dynamic energy efficiency management, and the contradiction between real-time wake-up and leakage suppression is prominent, and the energy efficiency bottleneck of the memory subsystem is significant.
Using a design method based on dynamic multi-domain hierarchical power consumption regulation, the two-way optimization of dynamic energy efficiency management and timing accuracy is achieved through the multi-level hierarchical clock gate architecture control system, dynamic power domain division, voice data cache compression and minimum addressable unit grouping optimization.
It realizes two-way optimization of dynamic energy efficiency management and timing accuracy, reduces overall clock power consumption, balances wake-up delay and leakage current, and solves the problem of energy efficiency bottleneck of memory subsystems.
Smart Images

Figure CN120197574A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of chip technology, and particularly relates to a design method of an AI voice chip based on dynamic multi-domain hierarchical power consumption regulation. Background Art
[0002] The low-power design techniques adopted in existing chip designs mainly rely on the following several techniques: 1. Dynamic Voltage and Frequency Scaling (DVFS): Dynamically adjusts the supply voltage and clock frequency of the chip according to the chip's computing load. However, in voice chips, it faces the problems of voltage conversion delay and accuracy compensation.
[0003] 2. Clock Gating: Reduces dynamic power consumption by turning off the clocks of idle modules. However, it does not perform dynamic fine-grained control for the sparse matrix calculation characteristics of neural network forward inference, and it is difficult to match the burstiness and discontinuity characteristics of voice processing tasks.
[0004] 3. Power Gating: Cuts off the power supply of non-working modules to reduce the leakage current of the chip. However, frequent power-on and power-off will generate voltage oscillation noise, affecting the signal-to-noise ratio of the analog front end.
[0005] 4. Near-Threshold Computing (NTC): Runs the circuit in the near-threshold voltage domain to reduce power consumption. However, the conventional near-threshold computing design is limited by process variations, resulting in difficult-to-converge timing, and it is very difficult to be implemented in commercial products.
[0006] Although the above technologies have been widely used in general low-power chips, there are still the following key defects in the specific field of AI voice chips: 1. Coarse-grained dynamic energy efficiency management Traditional DVFS usually adjusts the voltage at the module level and cluster level (CLUSTER). The signal chain of voice processing includes multiple heterogeneous units such as analog front end (AFE), digital signal processing (DSP), and neural network accelerator (NPU). The voltage and frequency sensitivities of each unit are significantly different. Coarse-grained regulation results in relatively large energy efficiency losses. Existing solutions do not dynamically adjust the voltage margin in combination with the acoustic event probability density function output by voice activity detection (VAD), resulting in significant misjudgment power consumption in a low signal-to-noise ratio environment, leading to insufficient balance optimization between wake-up delay and power consumption.
[0007] 2. The contradiction between real-time wake-up and leakage current suppression is prominent The voice chip needs to meet the sub-millisecond wake-up latency to respond to sudden voice commands. However, due to the delays in the power supply network and clock network, the conventional power gating strategy is difficult to balance fast wake-up and low leakage. Keeping the power domain always on will cause a large standby leakage current, and turning off the power domain requires the power supply network to be re-powered, which often takes a long stabilization time and cannot meet the real-time requirements.
[0008] 3. The energy efficiency bottleneck of the memory subsystem is significant Existing low-power SRAM architectures are difficult to synchronously support the memory access patterns for neural network parameter caching and the bursty data processing characteristics of voice chips. There are significant spatio-temporal locality differences between voice feature data (such as MFCC coefficients) and neural network parameters. However, the existing low-power SRAM architectures do not optimize the Bank grouping strategy for the bursty access characteristics of voice data. The number of row activations far exceeds the actual requirements, and there is a lack of hardware co-design for parameter zero-value compression and feature map sparsification, resulting in a large amount of dynamic power consumption of the memory being wasted on invalid data transfer. Summary of the Invention
[0009] To overcome the technical defects of the existing technology, the present invention discloses a design method for an AI voice chip based on dynamic multi-domain hierarchical power consumption regulation.
[0010] The ultra-low power design method for an AI voice chip based on dynamic multi-domain power consumption regulation of the present invention includes the following steps: Step 1. Divide the chip power domain according to the application scenario of the system, design the system clock and module clocks, and obtain the chip power supply scheme and clock scheme; Step 2. Design a multi-level hierarchical clock gating architecture control system according to the chip power supply scheme and clock scheme determined in Step 1. Specifically, by means of the hierarchical gating layer design method, the clock control is divided into three levels: chip-level gating, module-level gating, and data-level gating; The chip-level gating is gated according to the current working state of the entire chip, the module-level gating refers to gating according to the current working state of each module, and the data-level gating refers to gating according to the current data receiving and sending state of the chip; Step 3. Divide the power supply network into a dynamic power domain and a always-on power domain. Based on the principle of minimum resource selection, the current that is always under detection needs to be designed as the always-on power domain, and the power supply to inactive modules and inactive clusters is turned off; Step 4. Conduct cache compression design for voice data. Specifically, perform dynamic bit-width reorganization on the data through a cache compression technology oriented to voice bursts, adopt a hierarchical compression technology according to the importance of the data stream features output by the feature extraction module and the data characteristics, and adopt different bit-width adjustment strategies according to different data attributes; Step 5. Optimize the grouping of the minimum addressable units of the memory to support selective shutdown of the minimum addressable units with different granularities; Step 6. Perform system modeling and simulation. After the simulation passes and meets the expected goals, the design is completed.
[0011] Preferably, in the above Step 1, the clock scheme is to adopt a segmented clock tree structure. Specifically, multiple sub-clocks are obtained by dividing the system clock once or multiple times to reduce the frequency, and a tree structure is formed. Different sub-clocks are applied to different modules or clusters.
[0012] Preferably, in the above Step 2, the specific implementation method of the module-level gating is to deploy a state prediction unit in different modules or clusters. The state prediction unit is used to detect whether the module or cluster is in an idle state, and to turn off the clocks and power supplies of the modules and clusters in the idle state.
[0013] Preferably, the different bit-width adjustment strategies in the above Step 4 include: For the data of the higher-order coefficients of the Mel-frequency cepstral coefficients, data compression is performed by retaining the sign bit and truncating the lower bits; for the fundamental frequency feature data, a differential coding method is used to obtain the data compression ratio; for the voice activity flag data, Huffman coding is used for compression.
[0014] Preferably, in the above Step 4, it further includes integrating a data dynamic reorganization unit at the output end of the feature extraction module, and implanting a zero-value statistical function, a sign-bit prediction logic function, and a configurable truncation shifter function in the data dynamic reorganization unit.
[0015] Preferably, the above Step 5 is specifically: based on the burst access characteristics of voice data and the access requirement characteristics of voice chips for data in the running scenario, optimize the grouping of the minimum addressable units of the memory, and perform selective power cut-off on the minimum addressable units according to the time-domain prediction of the voice silent segment.
[0016] The AI voice chip design method based on dynamic multi-domain hierarchical power consumption regulation in the present invention has the following advantages compared with the prior art: First, the multi-level hierarchical clock dynamic gating architecture control system realizes the two-way optimization of dynamic energy efficiency management and timing accuracy through the method of hierarchical gating layer design.
[0017] Second, based on the voice activity detection module, control the startup and shutdown of the overall clock domain of the chip to achieve chip-level gating, and reduce the overall clock power consumption.
[0018] Third, based on the differential dynamic regulation of the task load, realize module-level gating, formulate different frequency strategies for different functional modules or clusters (CLUSTER) to achieve dynamic frequency modulation, and turn off the clock drive of the non-critical timing paths in the module in advance.
[0019] IV. Adaptive regulation based on computational features realizes data-level gating and enables the system to adjust adaptively.
[0020] V. A reconfigurable clock network is realized through a segmented clock tree structure to achieve dynamic energy efficiency management and balance wake-up latency and power consumption.
[0021] VI. By dividing different functional areas in the dynamic area and the pseudo-static area and implementing an asymmetric power supply strategy, leakage current suppression is achieved.
[0022] VII. Through cache compression technology for voice bursts, dynamic bit-width reorganization of data is carried out to achieve dynamic compression and reorganization of data, reduce data memory occupancy, and improve data efficient operation.
[0023] VIII. Selective power cut-off of the memory bank is implemented based on time-domain prediction of voice silent segments to solve the problem of the energy efficiency bottleneck of the memory subsystem. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic flowchart of a specific implementation manner of the AI voice chip design method based on dynamic multi-domain hierarchical power consumption regulation described in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The design method of ultra-low power consumption of the AI voice chip based on dynamic multi-domain power consumption regulation described in the present invention, as Figure 1 shown, is implemented according to the following steps: Step 1. Determine the chip power supply scheme, clock scheme, and low-power scheme according to the chip architecture definition, cluster (CLUSTER) division, subsystem division, and module function division. A cluster is a combination of some modules divided inside the chip according to functions, and each divided cluster can implement a part of the functions of the chip.
[0026] Divide the chip power supply domain according to the application scenario of the system, design the system clock and module clock, and obtain the chip power supply scheme and clock scheme. The above schemes enable each module and cluster of the chip to perform automatic gating functions according to the function module and cluster load conditions.
[0027] A specific implementation manner of the clock scheme is to adopt a segmented clock tree structure. Multiple sub-clocks are obtained by single or multiple frequency divisions and frequency reductions of the system clock, forming a tree structure. Different sub-clocks are applied to different modules or clusters. The structure of the clock tree can make the corresponding module or cluster not work when a certain sub-clock is turned off, thereby realizing power consumption management through the clock.
[0028] Step 2. Design a multi-level hierarchical clock gating architecture control system according to the chip power supply scheme and clock scheme determined in Step 1. Specifically: Through the method of hierarchical gating layer design, the clock control is divided into three levels: chip-level gating, module-level gating, and data-level gating; Chip-level gating is performed according to the current working state of the entire chip. Module-level gating refers to gating according to the current working state of each module. Data-level gating refers to gating according to the current data reception and transmission state of the chip.
[0029] For example, when the system is in the standby mode, only the circuits related to VAD (voice activity detection) are retained through chip-level gating to maintain a low-speed clock, and the standby power consumption at this time is extremely low. When the VAD module detects valid voice features, the initial frequency of the main clock tree is started. At this time, differential dynamic regulation based on the task load realizes module-level gating, and different frequency strategies are formulated for different functional modules or clusters to achieve dynamic frequency modulation. A zero-value detection circuit is implanted in the memory array of the neural network accelerator (NPU). The specific function of the zero-value detection circuit is that when the sparsity of the input feature map exceeds the set threshold, the system automatically performs row-by-row gating on the inactive processing units in the memory array, that is, gates the memory array in units of rows, so that the inactive processing units do not perform operations such as refreshing, reading, and writing.
[0030] State prediction units are deployed in different modules or clusters. The state prediction units are used to detect whether the module or cluster is in the idle state. The hardware automatically judges whether to turn off the clocks and power supplies of the modules and clusters in the idle state. Through the segmented clock tree structure, differential clock frequency strategies can be implemented for sub-modules such as noise suppression, feature extraction, and neural network inference, realizing a reconfigurable clock network and achieving two-way optimization of dynamic energy efficiency management and timing accuracy.
[0031] Step 3. Design the asynchronous data path triggered by events, manifested as: The power supply network is divided into a dynamic power domain and an always-on (ALWAYSON) power domain according to the architecture characteristics, and modeling and simulation are carried out. Based on the principle of resource-minimal selection, the current that is always under detection needs to be designed as the always-on power domain, and an asymmetric power supply strategy is implemented; The switching between the chip standby mode, sleep mode, and activation mode is triggered by events, the dynamic power domain is turned on and off, and the power supplies of inactive modules and inactive clusters are turned off to suppress leakage current, achieving a balanced optimization of leakage current suppression and wake-up delay.
[0032] Step 4. Design the cache compression of voice data, manifested as: The data is dynamically bit-width reorganized through the cache compression technology for voice bursts. The hierarchical compression technology is adopted according to the importance of the data stream features output by the feature extraction module and the data characteristics. Different bit-width adjustment strategies are adopted according to different data attributes; For example, data such as the higher-order coefficients of Mel Frequency Cepstral Coefficients (MFCC) can achieve a large data compression ratio by retaining the sign bit and truncating the lower bits; the fundamental frequency (F0) feature data can adopt differential coding to obtain a more appropriate data compression ratio; the voice activation flag data can be compressed using Huffman coding; Integrate a Data Dynamic Reorganization Unit (DRU) at the output of the feature extraction module, and implant functions such as zero value statistics (Zero Counter), sign bit prediction logic, and a configurable truncation shifter in the data dynamic reorganization unit to achieve dynamic compression and reorganization of data, reduce data memory occupancy, and improve data efficient operation.
[0033] Among them, the zero value statistics (Zero Counter) function is used to count and compress when consecutive zero characters appear in the data, the sign bit prediction logic is used to retain all sign bits during compression, and the specific function of the configurable truncation shifter is to select different truncation lengths according to the data characteristics.
[0034] Step 5: Optimize the grouping of the smallest addressable units (Banks) of the SRAM to support selective shutdown of different granularity Banks, manifested as: Based on the burst access characteristics of voice data and the access requirements of voice chip operating scenarios for data, optimize the grouping of the smallest addressable units (Banks) of the SRAM. According to the time-domain prediction of the voice silent segment, selectively cut off the power supply of the memory Banks, and finely control the effective row activation granularity according to the internal data bit width of the Bank granularity to ensure the balance between the number of row activations and the actual required number, and solve the problem of the energy efficiency bottleneck of the memory subsystem.
[0035] Step 6: Conduct system modeling, perform power consumption, performance, and function simulations, and end after the simulations pass and meet the expected goals.
[0036] The present invention can solve the problem of too coarse granularity in dynamic energy efficiency management: Traditional Dynamic Voltage and Frequency Scaling (DVFS) schemes usually perform voltage and frequency regulation at the module level and CLUSTER level. Before determining whether frequency modulation and voltage regulation are required, it is necessary to query the working status of the modules and CLUSTERs. This query is often implemented by software. For a SOC system, the larger the scale, the more data of the modules and CLUSTERs need to be queried, and the power consumption and delay consumed by the query increase with the increase of the system scale. In a more complex SOC system, this not only occupies a large amount of power consumption but also seriously affects the running performance of the system. Moreover, due to the complex software control process and large granularity, it faces the problems of large voltage conversion delay and precision compensation.
[0037] The multi-level dynamic gating scheme proposed by the present invention divides clock control into three levels: chip-level gating, module-level gating, and data-level gating through the method of hierarchical gating layer design. When the system is in the standby mode, only the VAD-related circuits are retained to maintain a low-speed clock, and the standby power consumption at this time is extremely low. When the VAD module detects valid voice features, the initial frequency of the main clock tree is started, and module-level gating is achieved based on the differential dynamic regulation of the task load. Different frequency strategies are formulated for different functional modules or clusters (CLUSTERs) to achieve dynamic frequency modulation. State prediction units are deployed in different modules or clusters to detect whether the module or cluster is in the idle state, and the hardware automatically determines whether to automatically turn off the clock and power supply; a zero-value detection circuit is implanted in the disk array of the neural network accelerator. When the sparsity of the input feature map exceeds a certain threshold, the system automatically performs row-by-row gating on the inactive PE units; A reconfigurable clock network is implemented through a segmented clock tree structure to achieve two-way optimization of dynamic energy efficiency management and timing accuracy.
[0038] Low-power voice chips need to meet extremely low wake-up latency to respond to sudden voice commands. However, the existing power clock network delay is difficult to balance fast wake-up and low leakage. Keeping the power domain always on will cause a large standby leakage current. When the power supply is turned off during standby, it takes a long stabilization time when power needs to be restored, which cannot meet the real-time requirements.
[0039] The present invention proposes to combine the timing characteristics of voice feature frames, divide the power supply network into a dynamic power domain and a pseudo-static power domain, implement an asymmetric power supply strategy, and perform dynamic power domain power-on and power-off management technology by triggering the switching of the chip standby mode, sleep mode, and activation mode through events. The power supply network of inactive modules and clusters is turned off to cut off their power supply to suppress leakage current. The balance optimization of leakage current suppression and wake-up latency is achieved.
[0040] There are significant spatio-temporal locality differences between voice feature data and neural network parameters. However, the existing low-power SRAM architecture does not optimize the Bank grouping strategy for the burst access characteristics of voice data. The number of row activations exceeds 40% of the actual demand, and there is a lack of hardware co-design for parameter zero-value compression and feature map sparsification of data, resulting in about 20% - 30% of the dynamic power consumption of the memory being wasted on invalid data transfer.
[0041] The present invention adopts a hierarchical compression technology according to the importance of data stream features and data characteristics output by the feature extraction module, and adopts different bit-width adjustment strategies according to different data attributes to realize dynamic compression and reorganization of data, reduce data memory occupation and improve data efficient operation; based on the time-domain prediction of the voice silent segment, selectively cut off the power supply of the memory bank, and finely control the effective row activation granularity through the internal data bit-width situation of the Bank granularity, ensuring that the number of row activations is close to the actual demand number, and solving the problem of the energy efficiency bottleneck of the memory subsystem.
[0042] For example, the chip system includes a DSP cluster, a CPU cluster, an NPU (embedded neural network processor) cluster, a MEDIA (media) cluster, a MEM (memory) cluster, an ALWAYSON (always-on) cluster, etc. The specific implementation examples are as follows: Step 1: Design the power supply scheme and clock scheme according to the chip performance requirements and functional scenario requirements, and perform power domain division and clock domain division. The power domain division principle follows the voltage requirements of each module's characteristics and whether there is a power-down scenario. Design the power supply scheme according to the separate power supply requirements of analog modules, pin port power supply requirements, digital logic power supply requirements, etc. For example, the pin port uses 3.3v power supply, the codec module uses 2.5v power supply, the flash uses 1.8v power supply, and the digital logic uses 0.9v power supply. The clock domain division is designed according to the principle of high cohesion and low coupling, and the clock tree is designed according to the scenario scheduling situation. For example, place each module supporting DSP calculation and processing inside the DSP cluster, place the peripheral modules supporting CPU scheduling and processing inside the CPU cluster, place the modules supporting NPU calculation, processing and scheduling inside the NPU cluster, and place the modules supporting media subsystem processing inside the MEDIA cluster.
[0043] Step 2: Design a multi-layer dynamic clock gating architecture control system according to the power supply scheme and clock scheme. Through scenario data stream analysis, the DSP cluster internally includes modules such as DSP, DMA (direct memory access module), and scheduler. In some scenarios, the DSP cluster does not work, so a separate power-down design for the DSP cluster can be carried out; in some scenarios, the load of the DSP cluster will change dynamically according to the business situation. To optimize the power consumption of the system while maintaining good performance, a dynamic gating system needs to be designed inside the DSP cluster. This component automatically controls the gating by monitoring the busy and idle states and load conditions of the modules inside the DSP cluster. When it detects that the idle time of all modules inside the DSP cluster reaches a certain threshold, it determines that the module has no business requirements temporarily and can automatically turn off the clock gating of the DSP cluster.
[0044] For finer-grained control of layering, dynamic gating control can be implemented inside the module. For example, the NPU consists of many processing and computing units. According to different load conditions, the processing and computing units can be dynamically called. The unused processing and computing units are automatically gated. Adaptive regulation based on computing characteristics is used to achieve data-level gating. The local clock tree of inactive processing and computing units is dynamically masked based on the sparsity of feature map activation. A zero-value detection circuit is implanted in the disk array of the NPU. When the sparsity of the input feature map exceeds a certain threshold, the system automatically implements row-by-row gating for inactive units. When the sparsity of the input feature map is lower than the threshold, the system automatically activates the inactive units and turns off the gating function to achieve system adaptive adjustment. Similarly, the same principle is adopted for the multi-layer dynamic gating architecture design of other clusters in the system as that of the DSP cluster.
[0045] Step 3: Divide the power supply network into a dynamic power domain and a always-on power domain. In the above system, the DSP cluster, CPU cluster, NPU cluster, MEDIA cluster, and MEM cluster are divided into the dynamic power domain, and the ALWAYSON cluster is divided into the always-on power domain. The power supply of the dynamic power domain supports individual shutdown. The ALWAYSON cluster is the always-on power domain. To reduce the power consumption of the chip, the circuit design in the ALWAYSON cluster follows the minimization principle and only includes necessary monitoring circuits for waking up other clusters for power-on, such as VAD, external interrupt, RTC real-time clock, etc.
[0046] Step 4: Design a data dynamic reorganization unit to perform voice data cache compression. Dynamic bit-width reorganization of data is performed through cache compression technology for voice bursts. Hierarchical compression technology is adopted according to the importance of the data stream features and data characteristics output by the feature extraction module. Different bit-width adjustment strategies are adopted according to different data attributes. For example, for high-order coefficients of Mel Frequency Cepstral Coefficients (MFCC), a larger data compression ratio can be obtained by retaining the sign bit and truncating the low bits. For the fundamental frequency F0 feature, differential coding can be used to obtain a more appropriate data compression ratio. For the voice activation flag, Huffman coding can be used for compression. A data dynamic reorganization unit is integrated at the output end of the feature extraction module, which includes functions such as zero-value statistics, sign-bit prediction, and configurable truncation shifters to achieve dynamic compression and reorganization of data, reduce data memory occupancy, and improve data efficient operation.
[0047] Step 5: Group and optimize the SRAM memories to achieve selective shutdown at different granularities. In the above system, different clusters may have SRAM memories of different sizes. According to the business scenarios and data flow scenarios, the SRAM memories are grouped and split to achieve controllable granularity shutdown and opening. At the same time, the memory clusters contain SRAM memories that can be shared. The appropriate segmentation granularity is determined according to the business scenarios and memory usage size to ensure that the purpose of saving power can be achieved while meeting the computing performance requirements.
[0048] Step 6: Perform full system modeling and simulation, model and simulate according to business flow and data flow, and continuously optimize according to the simulation results until it meets the design intent.
[0049] Compared with the prior art, the present invention has the following advantages: 1. Multi-level hierarchical clock dynamic gating architecture control system divides clock control into three levels: chip-level gating, module-level gating and data-level gating through the method of hierarchical gating layer design, realizing two-way optimization of dynamic energy efficiency management and timing accuracy.
[0050] 2. Chip-level gating is implemented by controlling the startup and shutdown of the overall clock domain of the chip based on the voice activity detection (VAD) module. When the VAD module detects valid voice features, it starts the initial frequency of the main clock tree. When the voice silence duration exceeds a certain time, it triggers the shutdown of the global clock network, leaving only the VAD-related circuits to maintain a low-speed clock, thereby reducing overall clock power consumption.
[0051] 3. Implement module-level gating based on differentiated dynamic regulation of task loads, formulate different frequency strategies for different functional modules or clusters (CLUSTER) to implement dynamic frequency modulation, such as implementing differentiated clock frequency strategies for sub-modules such as noise suppression, feature extraction, and neural network reasoning, and deploy state prediction units in different modules or CLUSTERs to detect whether the module or CLUSTER is in an idle state. When it is determined that there will be no computing tasks in the future, shut down the clock drive of non-critical timing paths in the module in advance.
[0052] 4. Adaptive control based on computing features realizes data-level gating. The local clock tree of inactive PE (Processing Element) is dynamically shielded based on the activation sparsity of feature graphs. A zero-value detection circuit is implanted in the disk array of NPU. When the sparsity of the input feature graph exceeds a certain threshold, the system automatically implements row-by-row gating of inactive PE units. When the sparsity of the input feature graph is lower than the threshold, the system automatically activates the inactive PE units and turns off the gating function to achieve system adaptive adjustment.
[0053] V. A reconfigurable clock network is implemented through a segmented clock tree structure. The clock-driven buffers of the corresponding branches are only turned on when the voice frame is valid, and the clock segment driving of the corresponding clock segment is turned off when it is invalid, so as to achieve dynamic energy efficiency management and balance the wake-up latency and power consumption.
[0054] VI. Different functional areas are divided for the dynamic area and the pseudo-static area, and an asymmetric power supply strategy is implemented. The power supply network is divided into a dynamic power domain and a pseudo-static power domain. Combining the timing characteristics of the voice feature frames, the chip operating modes are divided into standby mode (standby), sleep mode (Sleep), and active mode (Active). The switching of the chip standby mode, sleep mode, and active mode, the turning on and off of the dynamic power domain, and the dynamic power management technology with different operating modes are triggered by events. The power supply networks of inactive modules and CLUSTERs are turned off to suppress leakage current, and the balance optimization of leakage current suppression and wake-up latency is achieved.
[0055] VII. The data is dynamically reorganized in bit width through a cache compression technology for voice bursts. A hierarchical compression technology is adopted according to the importance of the data stream features output by the feature extraction module and the data characteristics, and different bit width adjustment strategies are adopted according to different data attributes to achieve the dynamic compression and reorganization of the data, reduce the data memory occupancy, and improve the efficient operation of the data.
[0056] VIII. Selective power cut-off of the memory bank is implemented based on the time-domain prediction of the voice silent segment. The effective row activation granularity is finely controlled according to the internal data bit width situation at the Bank granularity to ensure the balance between the number of row activations and the actual required number, and solve the problem of the energy efficiency bottleneck of the memory subsystem.
[0057] The foregoing are the preferred embodiments of the present invention. If the preferred implementation manners in each preferred embodiment are not obviously self-contradictory or premised on a certain preferred implementation manner, each preferred implementation manner can be arbitrarily superimposed and combined. The embodiments and the specific parameters in the embodiments are only for clearly expressing the inventor's invention verification process, and are not used to limit the patent protection scope of the present invention. The patent protection scope of the present invention still takes its claims as the criterion. Any equivalent structural changes made by using the content of the specification and drawings of the present invention should be equally included in the protection scope of the present invention.
Claims
1. A design method for ultra-low power consumption of AI voice chip based on dynamic multi-domain power consumption control, characterized in that: The steps include: Step 1. Divide the chip power domain according to the system application scenario, design the system clock and module clock, and obtain the chip power solution and clock solution; Step 2. Design a multi-level hierarchical clock gating architecture control system according to the chip power scheme and clock scheme determined in step 1. Specifically, the clock control is divided into three levels: chip-level gating, module-level gating and data-level gating by a hierarchical gating layer design method. The chip-level gating is gating according to the current working state of the entire chip, the module-level gating is gating according to the current working state of each module, and the data-level gating is gating according to the data receiving and sending state of the current chip; Step 3. Divide the power supply network into a dynamic power domain and a normally-on power domain. Based on the principle of selecting the minimum resources, the current that is always under detection needs to be designed as the normally-on power domain, and the power supply of inactive modules and inactive clusters needs to be turned off; Step 4. Design the cache compression of voice data, specifically, dynamically resize the data bit width through the cache compression technology for voice bursts, use hierarchical compression technology according to the importance of the feature of the data stream output by the feature extraction module and the data characteristics, and use different bit width adjustment strategies according to different data attributes; Step 5. Optimize the minimum addressable unit grouping of the memory, and support the selective shutdown of the minimum addressable units of different granularities; Step 6. Perform system modeling and simulation. The design is completed after the simulation passes and achieves the expected goals.
2. The ultra-low power consumption design method of an AI voice chip based on dynamic multi-domain power consumption control as claimed in claim 1 is characterized in that: In step 1, the clock scheme adopts a segmented clock tree structure, specifically, multiple sub-clocks are obtained by single or multiple frequency division and reduction of the system clock to form a tree structure, and different sub-clocks are applied to different modules or clusters.
3. The design method of ultra-low power consumption of AI voice chip based on dynamic multi-domain power consumption control as claimed in claim 1 is characterized in that: In step 2, the module-level gating is specifically implemented by deploying a state prediction unit in different modules or clusters, the state prediction unit is used to detect whether the module or cluster is in an idle state, and shut down the clock and power supply of the module and cluster in the idle state.
4. The ultra-low power consumption design method of an AI voice chip based on dynamic multi-domain power consumption control as claimed in claim 1, characterized in that: The different bit width adjustment strategies in step 4 include: The high-order coefficient data of Mel-frequency cepstral coefficients are compressed by retaining the sign bit and truncating the low bits; the fundamental frequency feature data are compressed by difference coding to obtain the data compression ratio; the speech activation mark data are compressed by Huffman coding.
5. The design method of ultra-low power consumption of AI voice chip based on dynamic multi-domain power consumption control as claimed in claim 1, characterized in that: The step 4 also includes integrating a data dynamic reorganization unit at the output end of the feature extraction module, and implanting a zero value statistics function, a sign bit prediction logic function, and a configurable truncation shifter function in the data dynamic reorganization unit.
6. The design method of ultra-low power consumption of AI voice chip based on dynamic multi-domain power consumption control as claimed in claim 1, characterized in that: The step 5 is specifically as follows: based on the burst access characteristics of voice data and the data access demand characteristics of the voice chip operation scenario, the minimum addressable unit grouping of the memory is optimized, and the minimum addressable unit is selectively powered off according to the time domain prediction of the voice silence segment.
Citation Information
Cited By
Multi-granularity gated clock processing method, electronic equipment and medium
CN120560489A