Voice keyword recognition method, system and device, storage medium, program product and chip
By adopting RISC-V architecture and pulsating array architecture in the keyword recognition system, modular design and efficient coordination are achieved, which solves the shortcomings of existing systems in terms of flexibility and computational efficiency, and improves the system's adaptability and energy efficiency.
Patent Information
- Application Number
- CN202511212993.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-11
AI Technical Summary
Existing keyword recognition systems are inadequate in terms of flexibility, computational efficiency, and power consumption. They are particularly difficult to adapt to different algorithm and model updates, and traditional architectures struggle to effectively support new algorithms such as extended convolutional neural networks.
It adopts a RISC-V architecture, combined with a custom instruction set and finite state machine design. It optimizes matrix operations through a systolic array architecture, achieves modular design and efficient coordination, separates control logic from computational logic, and supports specialized keyword recognition algorithms and model structure adjustments.
It improves system flexibility and computing efficiency, reduces development costs, optimizes resource utilization and power efficiency, and adapts to changing needs in different scenarios.
Smart Images

Figure CN120932637A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a voice keyword recognition method, system, device, storage medium, program product, and chip. Background Technology
[0002] Keyword Spotting (KWS) is a technology that detects specific words in an audio stream and is widely used in devices such as smart speakers and smartphones. Traditional KWS systems typically rely on cloud processing, which has drawbacks such as privacy concerns, high latency, and dependence on network connectivity. Therefore, implementing on-chip KWS systems has become an industry trend.
[0003] The existing on-chip KWS implementations mainly include the following solutions: (1) Software implementation based on general-purpose processors. This approach is highly flexible, but its power consumption and performance are not ideal. It is difficult to meet the requirements of real-time performance and low power consumption. It is not optimized for KWS tasks and has low execution efficiency. (2) Based on dedicated hardware accelerators, this approach has obvious advantages in performance and power consumption, but lacks flexibility and is difficult to adapt to different algorithm and model updates.
[0004] Furthermore, most existing dedicated processors for keyword recognition employ fixed architectures with high coupling between control and computation logic, making independent optimization difficult and hindering their ability to flexibly adapt to changing needs across different scenarios. This is particularly true when dealing with novel algorithms such as Dilated Convolutional Neural Networks (CNNs), where traditional architectures struggle to provide effective support, thus necessitating urgent improvements. Summary of the Invention
[0005] The purpose of this invention is to provide a voice keyword recognition method, system, device, storage medium, program product, and chip, which solves the problems of insufficient flexibility and low computing efficiency of the KWS system controller architecture in the prior art. Through custom instruction set extension and finite state machine design, it realizes efficient coordination and management of various functional modules of the KWS system, thereby improving the overall system performance and energy efficiency.
[0006] In a first aspect, embodiments of this application provide a speech keyword recognition method, which includes the following steps: an input stage, including a feature extraction step: receiving audio data and extracting audio features; a model inference stage, including a linear transformation step, a ReLU activation step, and a CNN processing step: the linear transformation step includes performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result; the ReLU activation step includes applying the ReLU activation function to the linear transformation result to obtain a first activation result; the CNN processing step includes performing an expanded CNN calculation on the activation result to obtain a CNN result; and an output stage: applying the Sigmoid activation function to the CNN result to obtain a speech keyword recognition result.
[0007] In this embodiment, the input stage feature extraction step is followed by a preprocessing step: the preprocessing step includes: applying mean and variance normalization to the audio features to obtain a normalized result; and performing a linear transformation on the normalized result during the model inference stage to obtain a linear transformation result.
[0008] In some possible implementations, the linear transformation step includes: performing matrix multiplication operations through a pulsating array to calculate the product of audio features and a preset linear layer weight matrix.
[0009] In some possible implementations, the CNN processing step may include padding the first activation result to obtain padded data.
[0010] In some possible implementations, the CNN processing steps include: reading the padded data and CNN weights; performing convolution operations to obtain the convolution result; batch normalizing the convolution result to obtain the batch normalized result; applying the ReLU activation function to the batch normalized result to obtain the second activation result; judging according to preset rules, if the second activation result meets the preset rules, then executing the next layer of CNN processing; if the second activation result does not meet the preset rules, then executing the data padding step and then executing the next layer of CNN processing; repeating the above steps to complete all CNN layer processing and obtain the CNN result.
[0011] Secondly, embodiments of this application provide a speech keyword recognition system, which includes: an interface module for receiving audio data and extracting audio features; a linear transformation module for performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result; a ReLU module for applying the ReLU activation function to the linear transformation result to obtain a first activation result; a CNN module for performing an expanded CNN calculation on the first activation result to obtain a CNN result; a Sigmoid activation module for applying the Sigmoid activation function to the CNN result to obtain a speech keyword recognition result; and a register module for storing the audio features, the preset linear layer weight matrix, the linear transformation result, the first activation result, the CNN result, and the speech keyword recognition result.
[0012] In this embodiment, the system further includes an instruction set extension module for triggering instructions.
[0013] In this embodiment, the system further includes a finite state machine control module for system state transitions and module enable control.
[0014] In this embodiment, the system also includes a data flow management module for coordinating data transmission.
[0015] In this embodiment, the system further includes a CMVN module for performing normalization of audio features by applying mean and variance to obtain a normalized result.
[0016] In some possible implementations, the system further includes a filling module for performing data addition and filling processing on the first activation result to obtain filled data.
[0017] In some possible implementations, the system further includes a batch normalization module for performing batch normalization processing on the convolution results to obtain batch normalization results.
[0018] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method described in the first aspect or any possible implementation of the first aspect.
[0019] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect or any possible implementation of the first aspect.
[0020] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect or any possible implementation of the first aspect.
[0021] In a sixth aspect, embodiments of this application provide a chip including a processor and a transceiver. The transceiver is used to perform data transmission and reception operations, and the processor is used to perform a method implementing the first aspect or any possible implementation of the first aspect.
[0022] The speech keyword recognition method, system, device, storage medium, program product, and chip of this invention have the following advantages compared with the prior art: (1) Enhanced flexibility: Based on the RISC-V architecture, it has low development costs, is easy to integrate and extend, and supports a custom instruction set with dedicated keyword recognition algorithms. The model structure and parameters can be adjusted according to requirements, and the modular design facilitates team collaboration and evolution. (2) Improved computational efficiency; dedicated instructions reduce the overhead of instruction decoding and execution; finite state machine implements optimized control flow; pulsating array architecture improves matrix operation efficiency. (3) Resource utilization is optimized, modular design enables resource sharing, computing units are reused in different operations, control logic and computing logic are separated, which facilitates independent optimization; (4) Improved power efficiency, dedicated control process reduces redundant operations, architecture optimization for KWS tasks, and efficient scheduling reduces data movement and storage overhead. Attached Figure Description
[0023] Figure 1 This is a flowchart of a speech keyword recognition method according to an embodiment of this application; Figure 2 This is a structural block diagram of a speech keyword recognition system according to an embodiment of this application; Figure 3 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0024] Embodiments embodying the features and advantages of the present invention will be described in detail in the following description. It should be understood that the present invention can have various variations in different examples without departing from the scope of the invention, and the descriptions and illustrations herein are for illustrative purposes only and not intended to limit the invention.
[0025] The terms used in this application are explained as follows: KWS: Keyword Spotting, a technique for detecting specific words in speech.
[0026] RISC-V: An open-source Reduced Instruction Set Computer architecture.
[0027] CMVN: Cepstral Mean and Variance Normalization, is an audio feature preprocessing technique.
[0028] Dilated CNN: A network structure that introduces dilated convolutions on top of a standard convolutional neural network, which can obtain a larger receptive field without increasing the number of parameters.
[0029] Systolic Array: A parallel computing architecture consisting of regularly arranged processing units, particularly suitable for matrix operations.
[0030] Fixed-point representation: A method of representing numbers by assigning a fixed number of digits to the integer and fractional parts of a fixed-point number.
[0031] This application provides a speech keyword recognition method, which is applied to speech recognition and embedded processing systems. Figure 1 This is a flowchart of a speech keyword recognition method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps: S101 Input Stage includes the feature extraction step: receiving audio data and extracting audio features; The S102 model inference stage includes a linear transformation step, a ReLU activation step, and a CNN processing step. The linear transformation step includes performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result. The ReLU activation step includes applying the ReLU activation function to the linear transformation result to obtain a first activation result. The CNN processing step includes performing an extended CNN calculation on the activation result to obtain a CNN result. In the S103 output stage, the CNN results are applied to the Sigmoid activation function to obtain the speech keyword recognition results.
[0032] In this embodiment of the application, preferably, the input stage feature extraction step further includes a preprocessing step: the preprocessing step includes: applying mean and variance normalization to the audio features to obtain a normalized result; and performing a linear transformation on the normalized result during the model inference stage to obtain a linear transformation result.
[0033] In one specific embodiment of this application, the linear transformation step includes: performing matrix multiplication operations through a pulsating array to calculate the product of audio features and a preset linear layer weight matrix.
[0034] In this embodiment of the application, specifically, the audio features are a 50×20 matrix, that is, 50 frames, each frame has 20-dimensional features, and the linear layer weight matrix is a 20×20 matrix.
[0035] In this embodiment of the application, preferably, the first activation result is further subjected to data padding before the CNN processing step to obtain padded data.
[0036] In one specific embodiment of this application, the CNN processing steps include: reading the padded data and CNN weights; performing convolution operations to obtain a convolution result; performing batch normalization on the convolution result to obtain a batch normalization result; applying the ReLU activation function to the batch normalization result to obtain a second activation result; judging according to preset rules, if the second activation result meets the preset rules, then performing the next layer of CNN processing; if the second activation result does not meet the preset rules, then performing the data padding step and then performing the next layer of CNN processing; repeating the above steps to complete all CNN layer processing and obtain the CNN result.
[0037] In this embodiment, the dilated CNN is computed as a four-stacked dilated CNN block.
[0038] In this embodiment, the fixed-point number is represented by a 32-bit fixed-point representation (1 sign bit, 7 integer bits, and 24 fractional bits), which simplifies hardware design while ensuring numerical accuracy.
[0039] This application also provides a voice keyword recognition system, which is applied to voice recognition and embedded processing systems. Figure 2 This is a structural block diagram of a speech keyword recognition system according to an embodiment of this application, as shown below. Figure 2As shown, the speech keyword recognition system includes: an interface module 201 for receiving audio data and extracting audio features; a linear transformation module 202 for performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result; a ReLU module 203 for applying the ReLU activation function to the linear transformation result to obtain a first activation result; a CNN module 204 for performing an expanded CNN calculation on the activation result to obtain a CNN result; a Sigmoid activation module 205 for applying the Sigmoid activation function to the CNN result to obtain a speech keyword recognition result; and a register module 206 for storing... The system includes: audio features, a preset linear layer weight matrix, linear transformation results, first activation results, CNN results, and speech keyword recognition results; an instruction set extension module 207 for triggering instructions; a finite state machine control module 208 for system state transitions and module enable control; a data flow management module 209 for coordinating data transmission; a CMVN module 210 for normalizing audio features using mean and variance to obtain normalized results; a padding module 211 for padding the first activation results to obtain padded data; and a batch normalization module 212 for batch normalizing the convolution results to obtain batch normalized results.
[0040] In this embodiment, reference Figure 2 The key data paths include: Feature input path: Interface module 201 → Register module 206 → CMVN module 210 → Register module 206; Linear transformation path: Register module 206 (CMVN result) → Linear transformation module 202 → Systolic array → Register module 206; CNN processing path: Register module 206 → Filling module 211 → CNN module 204 → Systolic array → Batch normalization module 212 → ReLU module 203 → Register module 206; Final output path: Register module 206 (Final CNN result) → Sigmoid activation module 205 → Register module 206 → Interface module 201.
[0041] In this embodiment, the triggering instructions (opcode) of the instruction set extension module 207 include: a CMVN instruction for triggering cepstral mean and variance normalization processing of audio features; a LINEAR instruction for triggering linear transformation operations and performing matrix multiplication calculations; a RELU instruction for performing ReLU activation function operations; a PADDING instruction for performing data padding operations to prepare input for the CNN layer; a CNN instruction for triggering the calculation of the dilated convolutional neural network; a BATCH_NORM instruction for performing batch normalization operations; and a SIGMOID instruction for performing the Sigmoid activation function to generate the final detection result.
[0042] In this embodiment, the finite state machine control module 208 has the following states: IDLE (idle state, waiting for instructions); CMVN (performing CMVN operation); LINEAR (performing linear transformation); RELU (performing ReLU activation); PADDING (performing padding operation); CNN (performing CNN computation); BATCH_NORM (performing batch normalization); and SIGMOID (performing Sigmoid activation).
[0043] In this embodiment, the state transition logic of the finite state machine control module 208 is as follows: The system is initially in the IDLE state. It determines which state to transition to based on the received instructions. After each processing state is completed, it automatically transitions to the next state according to the predefined processing flow. Under specific circumstances (such as the last step of the instruction sequence), the state machine returns to the IDLE state.
[0044] In this embodiment, the control signals for each module are as follows: cmvn_en: CMVN module enable signal; linear_en: Linear transformation module enable signal; relu_en: ReLU module enable signal; padding_en: Padding module enable signal; cnn_en: CNN module enable signal; batch_norm_en: Batch normalization module enable signal; sigmoid_en: Sigmoid module enable signal; systolic_en: Systolic array enable signal; systolic_op: Systolic array operation mode selection (00: matrix multiplication, 01: convolution).
[0045] In this embodiment, the present invention employs a systolic array architecture to accelerate matrix multiplication and convolution operations: Matrix multiplication mode: Used for linear transformations, calculating the product of the input feature and weight matrices, supporting 50×20 and 20×20 matrix multiplication operations.
[0046] Convolution operation mode: used to expand CNN computation, supporting convolution operations with different dilation rates.
[0047] The systolic array selects the operation mode through the systolic_op signal to achieve the reuse of computing resources.
[0048] In this embodiment of the application, the overall data flow is as follows: The input phase includes the feature extraction steps: the interface module 201 receives 50 frames of 20-dimensional audio feature data; the audio feature data is loaded into a specified position in the register module 206 through the interface module 201; the controller sends a start signal to start the processing flow.
[0049] Preprocessing steps: CMVN module 210 reads audio feature data from register module 206, applies mean and variance normalization, and writes the normalization result back to another area of register module 206.
[0050] The model inference stage includes the following linear transformation steps: the linear transformation module 202 reads the normalized result processed by the CMVN module 210; it reads the preset linear layer weight matrix from the register module 206, performs matrix multiplication (50×20 × 20×20) through the pulsating array, and writes the linear transformation result back to the register module 206.
[0051] ReLU activation steps: ReLU module 203 reads the linear transformation result, applies the ReLU activation function (negative values are truncated to zero), and writes the first activation result back to register module 206.
[0052] CNN processing steps: The padding module 211 adds padding to the data as needed. The CNN module 204 performs four stacked convolution operations: reads the padded data and CNN weights; performs convolution operations through a systolic array; the convolution results are processed by the batch normalization module 212, and the ReLU activation function is applied. If necessary, padding is applied again to prepare for the next CNN layer; the above steps are repeated to complete the processing of all CNN layers, and finally the CNN result is written back to the register module 206.
[0053] Output stage: Sigmoid activation module 205 reads the output of the final CNN layer, applies the Sigmoid activation function to generate keyword presence probability values, writes the keyword presence probability values back to register module 206 and outputs them through interface module 201, the controller sends a completion signal and returns to the IDLE state.
[0054] Combination Figure 1 and Figure 2 The detailed execution timing of the embodiments of this application is as follows: (1) Controller Startup → Instruction Decoding: Receive the first instruction (usually a CMVN operation), decode the instruction, and start the execution process.
[0055] (2) CMVN → LINEAR → RELU: CMVN module 210 is enabled (cmvn_en=1), performs normalization, and automatically switches to LINEAR state after completion; linear transformation module 202 is enabled (linear_en=1), and systolic array is enabled (systolic_en=1, systolic_op=00), and automatically switches to RELU state after completion; ReLU module 203 is enabled (relu_en=1), and executes ReLU activation function.
[0056] (3) Processing the first CNN block: When needed, the padding module 211 is enabled (padding_en=1), the CNN module 204 is enabled (cnn_en=1), the systolic array is enabled (systolic_en=1, systolic_op=01), the batch normalization module 212 is enabled (batch_norm_en=1), and the RELU module 203 is enabled again (relu_en=1).
[0057] (4) Processing of the second, third and fourth CNN blocks: The processing flow of the second, third and fourth CNN blocks is the same as that of the first CNN block, with different dilation rates and weight parameters. The first, second, third and fourth CNN blocks are executed in sequence: padding → convolution → batch normalization → ReLU. The data is updated cyclically in register module 206.
[0058] (5) Final output: After the fourth CNN block is processed, the Sigmoid activation module 205 is enabled (sigmoid_en=1), the final result is generated, the controller sets the done signal, the entire inference process is completed, the system returns to the IDLE state, and waits for the next operation.
[0059] Based on the same inventive concept as the above method embodiments, this application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the above method embodiments.
[0060] This invention also provides an electronic device. Figure 3 This is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present invention. Figure 3 As shown, the electronic device 300 includes: an instruction receiver 304, a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302. The various components of the system are coupled together via a bus 303. It is understood that the bus 303 is used to achieve communication between these components. In addition to a data bus, the bus 303 also includes a power bus, a control bus, and a status signal bus, etc. However, for clarity, in... Figure 3 All buses are referred to as bus 303.
[0061] It is understood that memory 301 can be volatile or non-volatile memory, or a combination of both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash EPROM, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), and Sync Link Dynamic Random Access Memory (SLDRAM). The memory 301 described in this embodiment is intended to include, but is not limited to, these and any other suitable types of memory.
[0062] Processor 302 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed through the integrated logic circuitry of the processor 302's hardware or through software instructions. The processor 302 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 302 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, specifically memory 301. The processor reads information from memory 301 and, in conjunction with its hardware, completes the steps of the aforementioned method.
[0063] In this embodiment, when the processor 302 executes the program, it implements the following: an input stage, including a feature extraction step: receiving audio data and extracting audio features; a model inference stage, including a linear transformation step, a ReLU activation step, and a CNN processing step, wherein the linear transformation step includes performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result; the ReLU activation step includes applying the ReLU activation function to the linear transformation result to obtain a first activation result; the CNN processing step includes performing an expanded CNN calculation on the activation result to obtain a CNN result; and an output stage, applying the Sigmoid activation function to the CNN result to obtain a speech keyword recognition result.
[0064] This invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements: an input stage, including a feature extraction step: receiving audio data and extracting audio features; a model inference stage, including a linear transformation step, a ReLU activation step, and a CNN processing step, wherein the linear transformation step includes performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result; the ReLU activation step includes applying the ReLU activation function to the linear transformation result to obtain a first activation result; the CNN processing step includes performing an expanded CNN calculation on the activation result to obtain a CNN result; and an output stage, applying the Sigmoid activation function to the CNN result to obtain a speech keyword recognition result.
[0065] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements: an input stage, including a feature extraction step: receiving audio data and extracting audio features; a model inference stage, including a linear transformation step, a ReLU activation step, and a CNN processing step, wherein the linear transformation step includes performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result; the ReLU activation step includes applying the ReLU activation function to the linear transformation result to obtain a first activation result; the CNN processing step includes performing an expanded CNN calculation on the activation result to obtain a CNN result; and an output stage, applying the Sigmoid activation function to the CNN result to obtain a speech keyword recognition result.
[0066] This invention also provides a chip, including a processor and a transceiver. The transceiver performs data transmission and reception operations, and the processor performs the following: an input stage, including a feature extraction step: receiving audio data and extracting audio features; a model inference stage, including a linear transformation step, a ReLU activation step, and a CNN processing step. The linear transformation step includes performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result; the ReLU activation step includes applying the ReLU activation function to the linear transformation result to obtain a first activation result; the CNN processing step includes performing an expanded CNN calculation on the activation result to obtain a CNN result; and an output stage, applying the Sigmoid activation function to the CNN result to obtain a speech keyword recognition result.
[0067] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can retrieve and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that contains, stores, communicates, propagates, or transmits programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optical scanning of paper or other media, followed by editing, interpretation or other suitable processing as necessary, and then stored in computer memory.
[0068] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals; application-specific integrated circuits (ASICs) having suitable combinational logic gates; programmable gate arrays (PGAs); and field-programmable gate arrays (FPGAs).
[0069] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0070] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.
Claims
1. A method for recognizing speech keywords, characterized in that, Includes the following steps: The input phase includes the feature extraction steps: receiving audio data and extracting audio features; The model inference phase includes linear transformation, ReLU activation, and CNN processing steps: The linear transformation step includes performing a linear transformation on the audio features according to a preset linear layer weight matrix to obtain a linear transformation result; The ReLU activation step includes applying the ReLU activation function to the linear transformation result to obtain a first activation result; The CNN processing steps include performing an expanded CNN calculation on the activation results to obtain the CNN result; Output stage: Apply the Sigmoid activation function to the CNN results to obtain the speech keyword recognition results.
2. The speech keyword recognition method as described in claim 1, characterized in that, The input stage feature extraction step is followed by a preprocessing step: the preprocessing step includes: applying mean and variance normalization to the audio features to obtain a normalized result; the model inference stage performs a linear transformation on the normalized result to obtain a linear transformation result.
3. The speech keyword recognition method as described in claim 1, characterized in that, The linear transformation step includes: performing matrix multiplication operations through a pulsating array to calculate the product of audio features and a preset linear layer weight matrix.
4. The speech keyword recognition method as described in claim 1, characterized in that, Before the CNN processing step, the data of the first activation result is padded to obtain the padded data.
5. The speech keyword recognition method as described in claim 4, characterized in that, The CNN processing steps include: Read the padded data and CNN weights; Perform a convolution operation to obtain the convolution result; The convolution results are batch normalized to obtain the batch normalized result. The batch normalization result is then applied to the ReLU activation function to obtain the second activation result. According to the preset rules, if the second activation result meets the preset rules, the next layer of CNN processing is executed; if the second activation result does not meet the preset rules, the data imputation step is executed and then the next layer of CNN processing is executed. Repeat the above steps to complete the processing of all CNN layers and obtain the CNN result.
6. A voice keyword recognition system, characterized in that, Include: The interface module is used to receive audio data and extract audio features; The linear transformation module is used to perform linear transformation on the audio features according to the preset linear layer weight matrix to obtain the linear transformation result; The ReLU module is used to apply the ReLU activation function to the linear transformation result to obtain the first activation result. The CNN module is used to perform an expanded CNN computation on the first activation result to obtain the CNN result; The Sigmoid activation module is used to apply the Sigmoid activation function to the CNN results to obtain speech keyword recognition results. The register module is used to store the audio features, the preset linear layer weight matrix, the linear transformation result, the first activation result, the CNN result, and the speech keyword recognition result.
7. The speech keyword recognition system as described in claim 6, characterized in that, It also includes: an instruction set extension module for triggering instructions.
8. The speech keyword recognition system as described in claim 6, characterized in that, It also includes: a finite state machine control module, used for system state transitions and module enable control.
9. The speech keyword recognition system as described in claim 6, characterized in that, It also includes a data flow management module for coordinating data transmission.
10. The speech keyword recognition system as described in claim 6, characterized in that, It also includes: the CMVN module, which performs normalization of audio features by applying mean and variance to obtain normalized results.
11. The speech keyword recognition system as described in claim 6, characterized in that, It also includes a fill module, which performs data addition and fill processing on the first activation result to obtain the filled data.
12. The speech keyword recognition system as described in claim 6, characterized in that, It also includes a batch normalization module, which performs batch normalization on the convolution results to obtain the batch normalized result.
13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
16. A chip, characterized in that, The chip includes a processor and a transceiver, the transceiver being used to perform data transmission and reception operations, and the processor being used to perform the steps of implementing the method according to any one of claims 1 to 5.