Method and Apparatus of Wake-up Word and Key-Word Recognition
Patent Information
- Application Number
- KR1020230058252
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-04
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2043-05-04
Smart Images

Figure R1020230058252_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a method and apparatus for recognizing starter words and commands. Background Technology
[0003] Keywork Spotting (KWS) systems are an important technology used as a means of communication between humans and machines in various applications, such as home appliances, automobiles, and medical equipment.
[0004] In many cases, KWS combines machine activation via wake-up-word (WUW) recognition with command classification. However, supporting diverse applications while maintaining low power, small footprint, and high performance is a challenging task for KWS hardware.
[0005] Human-machine interfaces (HMI) are being widely developed in fields such as virtual reality (VR), augmented reality (AR), automatic speech recognition (ASR), and gesture recognition (GR) for effective human-machine interaction.
[0006] Among ASR technologies, KWS, in particular, includes WUW recognition and command classification networks and is a universal technology widely used in robot vacuum cleaners, wearable devices, and smart speakers.
[0007] To implement a high-performance KWS system, the following two aspects must be considered.
[0008] The goal is to ensure high performance in diverse environments and real-time capability on hardware with limited resources. In the first aspect, it is necessary to efficiently process various functions—such as machine activation via WUW recognition and command classification—with high performance on a single device.
[0009] From the second perspective, limited area / power consumption and real-time performance must be guaranteed without server intervention. Since these two issues must be resolved simultaneously and involve a trade-off, various solutions are being proposed.
[0010] Recently, various studies have demonstrated the efficiency of deep neural network (DNN)-based KWS systems in terms of performance and power consumption [17,18,19], and convolutional neural networks (CNN) are being applied to many vision and audio systems.
[0011] However, DNN algorithms are difficult to apply to resource-constrained embedded platforms due to their high computational complexity and large memory requirements. To improve resource efficiency, this paper adopts a quantization network model, a widely used technique, as the primary algorithm.
[0012] Voice recognition technologies, such as those found in robot vacuums or smart speakers, require WUW recognition to activate powered-off machines and a command classification network to issue commands to activated machines in order to prevent malfunctions.
[0013] There are two ways for network accelerators to support this series of processes.
[0014] The first is to implement a network with high bit precision so that WUW recognition and instruction classification pass through the same network.
[0015] The second approach is to design and integrate separate quantization networks with bit precision optimized for each application. Since WUW recognition and instruction classification differ in the number of keywords to be classified, the network complexity and the total number of parameters also differ.
[0016] From an accelerator perspective, it is important to find a compromise between memory usage and area efficiency for the two solutions mentioned above.
[0017] Since WUW recognition performs a simple classification task of distinguishing between a small number of target keywords and 'unknown', research is mainly being conducted on increasing the quantization strength of the network to maximize power efficiency [1,2,3].
[0018] Dandan Song et al. [1] propose a BNN-based VAD and a wakeup system based on a binarized weight network (BWN). The model performs classification for two keywords with 32 times less memory and 5 times less latency compared to a GPU. Weiwei Shan et al. [2] propose a 28nm KWS system including mel-frequency cepstral coefficient (MFCC) preprocessing and a BNN accelerator.
[0019] Although it has a relatively complex network configuration, it minimizes memory size through a DSCNN-based BNN technique. This model performs classification for 1 / 2 keywords. BNN has the advantage of low power consumption and small area due to bit-level operations during hardware implementation, while possessing sufficient accuracy to perform WUW recognition depending on the network configuration.
[0020] Quantized networks with higher bit precision than BNNs are being studied for accuracy compensation in the task of classifying multiple instructions [4,5,6]. Bo Liu et al. [4] propose an adaptive BWN model in which the bit precision of the input data and the computational flow are modified by SNR prediction. Due to the pipelining structure and binarized weights, the size of the SRAM is very small, and the average energy consumption of the entire system is reduced through module-by-module approximation computing.
[0021] Command classification is performed for 10 keywords. Yu Gong et al. [6] propose an optional quantized convolutional neural network (QCNN) that corrects accuracy by quantizing data and weights into 4 / 8 bits depending on the SNR of the input data.
[0022] This model performs command classification for 10 keywords. This paper proposes a TNN with an expanded range of possible weight values to overcome the accuracy limitations of BNNs while maintaining a small memory size. It is optimized for parallel computation by maintaining bit-wise operations, can be implemented with fewer resources through a computational flow similar to BNNs, and enables network-specific memory optimization through mode switching with BNNs.
[0023] Accordingly, the present invention proposes a starter word and command recognition device and method that detects a valid voice signal through a sensor in an MPU, converts it into a Mel-spectrogram, and transmits it to an NNU to perform WUW recognition and command classification.
[0025] [Prior Art Literature]
[0026] [1] Song, D.; Yin, S.; Ouyang, P.; Liu, L.; Wei, S. Low Bits: Binary Neural Network for Vad and Wakeup. In Proceedings of the 2018 5th International Conference on Information Science and Control Engineering (ICISCE), Zhengzhou, China, 20-22 July 2018; pp. 306-311.
[0027] [2] Shan, W.; Yang, M; Wang, T.; Lu, Y.; Cai, H.; Zhu, L.; Xu, J.; Wu, C.; Shi, L.; Yang, J. A 510-nW Wake-Up Keyword-Spotting Chip Using Serial-FFT-Based MFCC and Binarized Depthwise Separable CNN in 28-nm CMOS. IEEE J. Solid-State Circuits, 2020, 56, 151-164.
[0028] [3] Zhu, L.; Shan, W.; Xu, J.; Lu, Y. AAD-KWS: a sub-μW keyword spotting chip with a zero-cost, acoustic activity detector from a 170nW MFCC feature extractor in 28nm CMOS. In Proceedings of the 51st IEEE European Solid-State Device Research Conference (ESSDERC), Grenoble, France, 13-22 September 2021; pp. 99-102.
[0029] [4] Liu, B.; Cai, H.; Wang, Z.; Sun, Y.; Shen, Z.; Zhu, W.; Li, Y.; Gong, Y.; Ge, W.; Yang, J.; Shi, L.; A 22nm, 10.8 μ W / 15.1 μ W Dual Computing Modes High Power-Performance-Area Efficiency Domained Background Noise Aware Keyword- Spotting Processor. IEEE Trans. Circuits and Systems , 2020, 67 , 4733-4746.
[0030] [5] Gong, Y.; Li, Y.; Ding, X.; Yang, H.; Zhang, Z.; Zhang, X.; Ge, W.; Wang, Z.; Liu, B.; QCNN Inspired Reconfigurable Keyword Spotting Processor With Hybrid Data-Weight Reuse Methods. IEEE Access , 2020, 8 , 205878-205893.
[0031] [6] Giraldo, J.S.P.; Lauwereins, S.; Badami, K.; Hamme, H.V.; Verhelst, M. 18μW SoC for near-microphone Keyword Spotting and Speaker Verification. In Proceedings of the 2019 Symposium on VLSI Circuits, Kyoto, Japan, 09-14 June 2019, pp. C52-C53.
[0032] [7] Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adan, H. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv 2017, arXiv:1704.04861.
[0033] [8] Gupta, H.; Gupta, D. LPC and LPCC method of feature extraction in speech recognition system. In Proceedings of the 6th International Conference - Cloud System and Big Data Engineering (Confluence), Noida, India, 14-15 January 2016, pp. 498-502
[0034] [9] Hermansky, H. Perceptual linear predictive (PLP) analysis of speech,” Acoust. Soc. Amer. J. , 1990, 87 , 1738-1752.
[0035]
[10] Hermansky, H.; Morgan, N.; Bayya, A.; Kohn, P. The challenge of inverse-E: The RASTA-PLP method. In Proceedings of the 25th Asilomar Conference on Signals, Systems & Computers, Pacific Grove, CA, USA, 04-06 November 1991, pp. 800-804.
[0036]
[11] Veton, Z.K.; Hussien, A.E. Robust speech recognition system using conventional and hybrid features of MFCC, LPCC, PLP, RASTA-PLP and hidden Markov model classifier in noisy conditions,'' J. Computer and Communications , 2015, 3 , 1-9.
[0037]
[12] Ioffe, S.; Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv 2015, arXiv:1502.03167
[0038]
[13] Courbariaux, M.; Hubara, I.; Soudry, D.; El-Yaniv, R.; Bengio, Y. Binarized neural networks: Training deep neural networks with weights and activations constrained to +1 or -1. arXiv 2016, arXiv:1602.02830.
[0039]
[14] Miyashita, D.; Kousai, S.; Suzuki, T.; Deguchi, J. A neuromorphic chip optimized for deep learning and CMOS technology with timedomain analog and digital mixed-signal processing. IEEE J. Solid-State Circuits , 2017, 52 , 2679-2689.
[0040]
[15] Warden, P. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv 2018, arXiv:1804.03209
[0041]
[16] Adjoudani, A.; Beck, E.C.; Burg, A.P.; Djuknic, G.M.; Gvoth, T.G.; Haessig, D.; Wolniansky, P.W. Prototype experience for MIMO BLAST over third-generation wireless system. IEEE J. Sel. Areas Commun . 2003, 21 , 440-451.
[0042]
[17] Ando, K.; Ueyoshi, K.; Orimo, K.; Yonekawa, H.; Sato, S.; Nakahara, H.; Takamaeda-Yamazaki, S.; Ikebe, M.; Asai, T.; Juroda, T.; Motomura, M. BRein Memory: A Single-Chip Binary / Ternary Reconfigurable in-Memory Deep Neural Network Accelerator Achieving 1.4 TOPS at 0.6 W. IEEE J. Solid-State Circuits , 2018, 53 , 983-994.
[0043]
[18] Bankman, D.; Yang, L.; Moons, B.; Verhelst, M.; Murmann, B. An always-on 3.8 J / 86% CIFAR-10 mixed-signal binary CNN processor with all memory on chip in 28-nm CMOS. IEEE J. Solid-State Circuits , 2019, 54 , 158-172.
[0044]
[19] Choi, S.; Lee, J.; Lee, K.; Yoo, H.-J. A 9.02 mW CNN-stereo-based real-time 3D hand-gesture recognition processor for smart mobile devices. In Proceedings of the 2018 IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 11-15 February 2018, pp. 220-222 해결하려는 과제
[0046] The purpose of the present invention is to provide a method and device for recognizing starter words and commands that can solve the problems of the past. means of solving the problem
[0048] A wake-up word and command recognition device according to an embodiment of the present invention for solving the above problem includes: an MPU that preprocesses voice data into a Mel-spectrogram; and an NNU that applies the Mel-spectrogram to a binary classification network to recognize a wake-up word of voice data from the Mel-spectrogram, converts it to a ternary classification network, and then applies a subsequently input Mel-spectrogram to the ternary classification network to classify commands within the voice data.
[0050] In one embodiment, an A / D converter is further included to convert an analog voice signal into the voice data, which is a digital signal.
[0052] In one embodiment, the MPU includes an energy-based VAD that verifies the validity of data continuously input through the pmod terminal of the ZCU104 board; an STFT unit that generates a spectrogram analyzing frequency changes over time of the continuous data; and a Mel-filtering unit that performs Mel-filtering operations on the spectrogram to reflect the structural characteristics of human hearing.
[0054] In one embodiment, the VAD is configured to detect actual voice intervals within the voice data, and is characterized by recognizing the starting point of a valid voice signal when the difference between the energy level within a specific interval and the energy level of the previous interval exceeds a threshold value.
[0056] In one embodiment, the VAD includes an address controller and a memory internally, and the address controller stores data in the memory with address values from 0 to 127 for every 128 inputs.
[0058] In one embodiment, the STFT unit starts a 256-point FFT by filling data with address values between 128 and 255 after validation, and then stores a value at an arbitrary address for every 128 data inputs and performs a 256-point FFT when there is 50% overlap.
[0060] In one embodiment, the STFT unit is characterized by performing the STFT process through FFT operations every 256 points using a fixed point radix-4 algorithm of the DIF method.
[0062] In one embodiment, the STFT unit receives serialized data as input, performs a radix-4 operation using a buffer, and stores partial result values of the FFT in a 32x128 local memory.
[0064] In one embodiment, the STFT unit is characterized by deriving 129 non-duplicate complex values during FFT operation and deriving absolute values required for Mel filtering through bit-shift and arithmetic operations.
[0066] In one embodiment, the Mel-filtering unit is characterized by including a Mel-filter bank composed of a square filter converted by quantizing a triangular filter.
[0068] In one embodiment, the NNU includes an XNOR PE that stores input data, two pop-counters, an accumulator, an activation block, a concatenator, and a max pooling block, and each block is characterized by being designed with variable parameters to accommodate various network topologies.
[0070] In one embodiment, the NNU is designed as a layer accelerator structure, which is a reusable parallel computing structure with low resource usage, and is characterized by supporting a layer-by-layer adaptive parallel computing method.
[0072] In one embodiment, the NNU is characterized by operating to reduce memory usage by binarizing the network input and binarizing or ternaryizing the weights according to the target.
[0074] In one embodiment, the NNU is characterized by including an SRAM for reading and storing parameters assigned to a layer from an external DDR.
[0076] In one embodiment, the NNU is characterized by replacing the convolution operation with bitwise XNOR and PoP-Count operations.
[0078] In one embodiment, the NNU is characterized by supporting a function to set at least one of the quantization level, padding, pooling, bias, and BN of each layer.
[0080] In one embodiment, the Pop-counter counts the number of [1] and [-1] data in the XNOR result of 128-bit data, the Accumulator accumulates the pop-counter result value in the output channel, the activation block binarizes the final result value, the Concatenator groups the binarized result value into 128-bit units per channel, and the max pooling block reduces the output image size to 1 / 4 depending on the selection.
[0082] A method for recognizing wake-up words and commands according to an embodiment of the present invention for solving the above problem comprises: a step of preprocessing voice data into a Mel-spectrogram in an MPU; and a step of applying the Mel-spectrogram to a binary classification network in an NNU to recognize a wake-up word of voice data from the Mel-spectrogram, converting to a ternary classification network, and then applying a subsequently input Mel-spectrogram to the ternary classification network to classify commands within the voice data.
[0084] In one embodiment, the method further includes the step of converting an analog voice signal into the voice data, which is a digital signal, through an A / D converter.
[0086] In one embodiment, the preprocessing step includes a step of recognizing as the starting point of a valid voice signal when the difference between the energy level within a specific interval and the energy level of the previous interval exceeds a threshold value when the actual voice interval is detected in the voice data in the VAD.
[0088] In one embodiment, the preprocessing step is characterized by including the step of starting a 256-point FFT by filling data with address values between 128 and 255 after validation in the STFT unit, and then performing a 256-point FFT when there is 50% overlap by storing a value at an arbitrary address for every 128 data inputs.
[0090] In one embodiment, the preprocessing step is characterized by including a step of performing an STFT process through FFT operations every 256 points using a fixed point radix-4 algorithm of the DIF method in the STFT unit.
[0092] In one embodiment, the preprocessing step is characterized by including the step of receiving serialized data from the STFT unit as input, performing a radix-4 operation using a buffer, and storing partial result values of the FFT in a 32x128 local memory.
[0094] In one embodiment, the preprocessing step is characterized by including the step of deriving 129 non-duplicate complex values during FFT operation in the STFT unit, and deriving absolute values required for MEL filtering through bit-shift and arithmetic operations.
[0096] In one embodiment, the step of classifying the instruction comprises counting the number of [1] and [-1] data in the XNOR result of 128-bit data in the Pop-counter of the NNU, accumulating the result value of the pop-counter in the output channel in the Accumulator of the NNU, binarizing the final result value in the activation block of the NNU, grouping the binarized result value into 128-bit units by channel in the Concatenator of the NNU, and reducing the output image size to 1 / 4 according to selection in the max pooling block of the NNU. Effects of the invention
[0098] A starter word and command recognition device according to one embodiment of the present invention is composed of a mel-processing unit (MPU) and a neural network unit (NNU) to classify real-time voice signals and is implemented in a Xilinx UltraScale+ ZCU104 FPGA board environment, and has the advantage of being able to achieve high accuracy in WUW recognition (97.1%) and command classification using 10 keywords (90.5%). Brief explanation of the drawing
[0100] Figure 1 is an example diagram comparing CNN and DS-CNN. FIG. 2a is an architecture of a starter word and command recognition device according to one embodiment of the present invention. FIG. 2b is a flowchart illustrating a starter word and command recognition method according to an embodiment of the present invention. Figure 3 is an example of a Mel-spectrogram of a dataset used to test the performance of the device shown in Figure 2a. Figure 4 is a diagram illustrating the difference between a triangular filter bank and a square filter bank. Figure 5 is an example diagram of a DS-BTNN network structure. Figure 6 is the hardware architecture of the MPU shown in Figure 2a. Figure 7 is the hardware architecture of the NNU shown in Figure 2a. FIG. 8 is an example of an FPGA platform with an advanced extensible interface (AXI) bus interface for verifying the performance of the device shown in FIG. 2a. Figure 9 is a figure showing the verification environment on an FPGA platform. Specific details for implementing the invention
[0101] Embodiments of the present invention are described below with reference to the attached drawings so that those skilled in the art can easily implement the invention. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.
[0102] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected" but also cases where they are "electrically connected" with other elements interposed between them. Furthermore, when a part is described as "including" a component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components, and it should be understood that this does not preclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0103] Terms such as “about,” “substantially,” etc., used throughout the specification, are used to mean at or near the stated value when inherent manufacturing and material tolerances are presented in the stated meaning, and are used to prevent unscrupulous infringers from unfairly exploiting the disclosure in which precise or absolute values are mentioned to aid in understanding the invention. Terms such as “step” or “step of” used throughout the specification of the invention do not mean “step for”.
[0104] In this specification, the term "part" includes a unit realized by hardware, a unit realized by software, and a unit realized using both. Additionally, one unit may be realized using two or more pieces of hardware, and two or more units may be realized by one piece of hardware. Meanwhile, "part" is not limited to software or hardware, and "part" may be configured to reside in an addressable storage medium or configured to run on one or more processors. Accordingly, as an example, "part" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided within the components and "parts" may be combined into a smaller number of components and "parts" or further separated into additional components and "parts." In addition, the components and '~parts' may be implemented to play one or more CPUs within the device or secure multimedia card.
[0105] Some of the operations or functions described herein as being performed by a terminal, device, or device may instead be performed by a server connected to said terminal, device, or device. Likewise, some of the operations or functions described as being performed by a server may also be performed by a terminal, device, or device connected to said server.
[0106] In this specification, some of the operations or functions described as mapping or matching with a terminal may be interpreted to mean mapping or matching the terminal's unique number or personal identification information, which is the terminal's identifying data.
[0108] Hereinafter, a starter word and command recognition device and method according to an embodiment of the present invention will be described in more detail based on the attached drawings.
[0109] First, prior to describing the present invention, the terms and / or algorithms mentioned in the present invention are briefly explained.
[0110] 1. Mel-Processing Algorithm
[0111] Recently, there are four main methods for digital feature extraction techniques primarily used in voice data, as follows.
[0112] MFCC, linear prediction coding coefficient (LPCC), perceptual linear production (PLP), relative spectral analysis perceptual linear prediction (RASTA-PLP) [8,9,10].
[11] evaluates the advantages and disadvantages of these approaches through experimental comparative analysis.
[0113] Experimental results indicate that the MFCC technique is a highly efficient feature extraction technique for data containing background noise or low SNR, as it has excellent robustness and low computational complexity.
[0114] The present invention adopts the MFCC technique, and MFCC has the following feature extraction process.
[0115] pre-emphasis, framing, windowing, short-time Fourier transform (STFT), mel-filtering, logarithmic operation, discrete cosine transforms (DCT).
[0116] STFT is used to overcome the limitations of FFT.
[0117] In the STFT operation, the input signal is divided into consecutive frames, and after multiplying by a Hamming window function to minimize signal discontinuity, the FFT is performed. The frames resulting from the STFT each contain frequency information, and since this frequency information changes in the time domain due to consecutive frames, they also contain time information. An image having two-dimensional information in the time-frequency domain is called a spectrogram, and the equation for the STFT is as follows (1).
[0118] (1)
[0119] Mel-filtering is an operation that maps a mel-filter bank containing the structural characteristics of human hearing, which is sensitive to low frequencies and insensitive to high frequencies, and outputs data in the form of a mel-spectrogram. A logarithmic operation is applied to reduce the distribution range of the mel-filter, and finally, an MFCC is obtained through a discrete cosign transform (DCT).
[0121] 2. Binarized / Ternarized Neural Network
[0122] CNN models are largely composed of convolutional layers (CL) and fully connected layers (FCL).
[0123] CL is a learnable layer that performs convolution operations directly on input data to extract features without loss of input data.
[0124] FCL is a layer that performs a linear transformation on 1-D vectors through a weight matrix and uses the final result generated by CL as input data to perform classification for the final label. The pooling layer reduces data size, thereby reducing computational complexity and preventing overfitting. The activation layer plays an important role in solving complex problems by reflecting non-linear characteristics in the network. Batch normalization (BN)
[12] not only increases the learning speed by maintaining the distribution of input data but also prevents the gradient vanishing problem. CNNs have been widely used in various image classification competitions due to their high robustness against image distortion and variation.
[0125] Generally, the data used in the actual computation of a CNN, such as output values, weights, and biases, consists of floating-point data. In addition, operations such as multiplication and division used in CL, FCL, and BN not only consume a large amount of memory but also increase power consumption due to their high computational complexity. Consequently, it was difficult to apply CNNs on resource-constrained embedded platforms, and various lightweight CNNs were introduced. BNN
[13] is a representative lightweight CNN, characterized by compressing the CNN's activations and weights so that they consist of 1s and -1s instead of multi-bit floating-point data.
[0126] The multiply-accumulation used in CL is replaced by 1-bit XNOR and pop-count. While BN is performed by complex operations in multi-bit NN, BNN is simplified into the process of adding an offset to the result. Since the four parameters used in the network inference operation of BN are fixed and σ always has a positive value, the sign direction is changed according to γ, so it can be expressed by the following equations (2) and (3)
[14] .
[0127]
[0129] BNNs significantly reduce memory usage by compressing weights and input data into a single bit and can perform hardware-optimized parallel operations using bitwise operations such as XNOR and pop-count. However, they have limitations in complex networks, such as multi-keyword detection, due to the degradation in accuracy resulting from this lightweight design. To address this issue, we propose a TNN that maintains input data in binary while converting weights to ternary. TNNs offer higher accuracy than BNNs due to their high bit precision, while retaining the bit-wise operation method and featuring very similar computational processes between the two networks.
[0131] 3. Depth-wise Separable Convolutional Neural Network
[0132] Low-bit quantization methods like BNN and TNN compress memory usage to the extreme, but cause problems with reduced network accuracy.
[0133] To compensate for the degradation in accuracy, a complex network structure was constructed, and DS-CNN, a lightweight technique widely used in MobileNet that achieves high accuracy with the same parameters while reducing memory usage, was applied.
[0134] DS-CNN is a network that separates the local and global characteristics of general convolution operations into separate layers. Depthwise (DW) convolution performs operations by matching a single input channel to an output channel, thereby excluding inter-channel correlations and reflecting local characteristics. Pointwise (PW) convolution is equivalent to 1x1 convolution and reflects inter-channel correlations, i.e., global characteristics. Figure 1 shows CNN and DS-CNN, and Table 1 below shows that the number of parameters and computational load in a specific layer have decreased by approximately eight times.
[0135] [Table 1]
[0136]
[0137] FIG. 2a is an architecture of a starter word and command recognition device according to an embodiment of the present invention, and FIG. 2b is a flowchart illustrating a starter word and command recognition method according to an embodiment of the present invention.
[0138] FIG. 3 is an example of a Mel-spectrogram of a dataset used to test the performance of the device shown in FIG. 2a, FIG. 4 is a diagram illustrating the difference between a triangular filter bank and a square filter bank, FIG. 5 is an example of a DS-BTNN network structure, FIG. 6 is the hardware architecture of the MPU shown in FIG. 2a, FIG. 7 is the hardware architecture of the NNU shown in FIG. 2a, FIG. 8 is an example of an FPGA platform with an advanced extensible interface (AXI) bus interface for verifying the performance of the device shown in FIG. 2a, and FIG. 9 is a diagram showing the verification environment on the FPGA platform.
[0140] First, as illustrated in FIG. 2a, a starter word and command recognition device (100) according to one embodiment of the present invention includes an MPU (120) and an NNU (130).
[0141] The above MPU (120) is configured to preprocess voice data into a Mel-spectrogram, and the NNU (130) may be configured to apply the Mel-spectrogram to a binary classification network to recognize the wake-up word of the voice data from the Mel-spectrogram, convert it to a ternary classification network, and then apply a subsequently input Mel-spectrogram to the ternary classification network to classify commands within the voice data.
[0142] In addition, the present invention may further include an A / D converter (110) that converts an analog voice signal into the voice data, which is a digital signal.
[0143] The above MPU (120) includes an energy-based VAD (121) that verifies the validity of data continuously input through the pmod terminal of the ZCU104 board; an STFT unit (122) that generates a spectrogram analyzing frequency changes over time of continuous data; and a Mel-filtering unit (123) that performs Mel-filtering operations on the spectrogram to reflect the characteristics of the human auditory structure.
[0144] The above VAD (121) is configured to detect actual voice intervals within the voice data, and recognizes the starting point of a valid voice signal when the difference between the energy level within a specific interval and the energy level of the previous interval exceeds a threshold value.
[0145] Additionally, the above VAD (121) includes an address controller and memory internally, and the address controller may be configured to store data in memory with address values from 0 to 127 for every 128 inputs.
[0146] Next, the STFT unit (122) starts a 256-point FFT by filling the data with address values between 128 and 255 after validation, and then performs a 256-point FFT when there is 50% overlap by storing a value at an arbitrary address for every 128 data inputs.
[0147] The above STFT unit (122) can perform the STFT process through FFT operations every 256 points using a fixed point radix-4 algorithm of the DIF method.
[0148] The above STFT unit (122) can receive serialized data as input, perform a radix-4 operation using a buffer, and store partial result values of the FFT in 32x128 local memory.
[0149] The above STFT unit (122) can derive 129 non-duplicate complex values during FFT operation and derive absolute values required for MEL filtering through bit-shift and arithmetic operations.
[0150] The above Mel-filtering unit (123) may include a Mel-filter bank composed of a square filter converted by quantizing a triangular filter.
[0151] The above NNU (130) includes an XNOR PE that stores input data, two pop-counters, an accumulator, an activation block, a concatenator, and a max pooling block, and each block can be designed with variable parameters to accommodate various network topologies.
[0152] The above NNU (130) is designed with a layer accelerator structure, which is a reusable parallel computing structure with low resource usage, and can support a layer-by-layer adaptive parallel computing method.
[0153] The above NNU (130) can operate to reduce memory usage by binarizing network inputs and binarizing or ternaryizing weights according to the target.
[0154] The above NNU (130) includes an SRAM for reading and storing parameters assigned to a layer from an external DDR.
[0155] The above NNU (130) may be configured to replace convolution operations with bitwise XNOR and PoP-Count operations.
[0156] The above NNU (130) can support the function of setting at least one of the quantization level, padding, pooling, bias, and BN of each layer.
[0157] The Pop-counter of the above NNU (130) counts the number of [1] and [-1] data in the XNOR result of the 128-bit data, the Accumulator accumulates the pop-counter result value in the output channel, the activation block binarizes the final result value, the Concatenator groups the binarized result value into 128-bit units per channel, and the max pooling block performs the function of reducing the output image size to 1 / 4 depending on the selection.
[0159] FIG. 2b is a flowchart illustrating a starter word and command recognition method according to an embodiment of the present invention.
[0160] Referring to FIG. 2b, the starting word and command recognition method (S700) according to an embodiment of the present invention preprocesses voice data into a Mel-spectrogram (S710) in the MPU (120).
[0161] The above S700 may further include a step of converting an analog voice signal into the voice data, which is a digital signal, through an A / D converter.
[0162] The above S710 process includes a process of recognizing a valid voice signal starting point when the difference between the energy level within a specific interval and the energy level of the previous interval exceeds a threshold value, when the VAD detects an actual voice interval within the voice data.
[0163] Additionally, the STFT unit may include a step of starting a 256-point FFT by filling data with address values between 128 and 255 after validation, and then performing a 256-point FFT when there is 50% overlap by storing a value at a random address for every 128 data inputs.
[0164] In addition, the STFT section may include a step of performing an STFT process through FFT operations every 256 points using a fixed point radix-4 algorithm of the DIF method.
[0165] In addition, the method may include a step of receiving serialized data from the STFT unit as input, performing a radix-4 operation using a buffer, and storing partial result values of the FFT in a 32x128 local memory.
[0166] In addition, when performing FFT operations in the STFT unit, the method may include a step of deriving 129 non-duplicate complex values and deriving absolute values required for Mel filtering through bit-shift and arithmetic operations.
[0167] When the above S710 process is completed, the NNU (130) applies the Mel-spectrogram to a binary classification network to recognize the wake-up word of the voice data from the Mel-spectrogram, switches to a ternary classification network, and then applies the subsequently input Mel-spectrogram to the ternary classification network to classify the command within the voice data (S720).
[0168] The above S720 process is characterized by including the steps of counting the number of [1] and [-1] data in the XNOR result of 128-bit data in the Pop-counter of the NNU, accumulating the result value of the pop-counter in the Accumulator of the NNU to the output channel, the activation block of the NNU binarizing the final result value, grouping the binarized result value into 128-bit units by channel in the Concatenator of the NNU, and reducing the output image size to 1 / 4 according to selection in the max pooling block of the NNU.
[0170] The structures of the MPU and NNU presented in the present invention will be described in more detail below.
[0171] First, the MPU (120) of the present invention is composed of an energy-based VAD (121) that verifies the validity of data continuously input through the pmod terminal of the ZCU104 board as shown in FIG. 6, an STFT unit (122) for analyzing frequency changes over time of continuous data, and a Mel-filtering unit (123) for reflecting human auditory structural characteristics.
[0172] Energy-based VAD (121) is an algorithm widely used to extract actual speech intervals from continuous data.
[0173] It signals the start of a valid voice signal when the difference between the energy level within a specific section and the energy level of the previous section exceeds a threshold value.
[0174] The address controller inside the VAD (121) stores data with address values from 0 to 127 for every 128 inputs.
[0175] After validation, the 256-point FFT is started by filling the data with address values between 128 and 255, and then a 50% overlapped 256-point FFT is performed by storing the value at an appropriate address for every 128 data inputs.
[0176] The STFT process is performed by conducting FFT operations every 256 points using the fixed-point radix-4 algorithm of the DIF method. Serialized data is received as input, radix-4 operations are performed using a buffer, and the partial product of the FFT is stored in 32x128 local memory. Since the input data is in real number format, the FFT yields 129 non-repeating complex values. An approximation technique was used to replace complex operations such as multiplication and squaring, which are used in calculating absolute values required for MEL filtering, with simple operations such as bit-shift and addition
[16] , and the equation is as follows.
[0177] [ceremony]
[0178]
[0179] Subsequently, the spectrogram obtained from STFT is passed through a Mel-filter bank to perform filtering operations.
[0180] By quantizing a triangular filter and converting it into a square filter, filtering can be performed using a simple accumulation operation.
[0181] The LUT stores the start and end indices of the FFT result values corresponding to each filter.
[0182] Then, data is read from SRAM based on the index and accumulated. The final result is stored in DDR and used as input data for the NNU.
[0184] Next, the NNU (130) may be configured to perform WUW recognition by applying the Mel spectrogram output from the MPU (120) to a network set to binary mode, and when the input voice data is recognized as WUW, the network is changed to ternary mode and then perform command classification from subsequent input data.
[0185] The above NNU (130) is composed of a DS-BTNN in the form of a combination of DS-CNN and BTNN, and the DS-BTNN network structure is composed of 9 layers (CL + (DW CL + PW CL)x3 + FCLx2) as shown in FIG. 5, and all layers except the output layer include Batch Normalization (BN) and sign activation, and max pooling is performed in the first layer.
[0186] The DS-BTNN network structure introduced here is designed to allow users to select configurations such as bit precision (1 / 2 bit), padding, BN, and kernel size, thereby supporting various networks.
[0187] The input layer uses a binarized / ternarized weight network (BTWN), which quantizes only the weights rather than the input data, because Mel-spectrograms contain a large amount of information relative to their small data size.
[0189] As shown in FIG. 8, the NNU (130) consists of an XNOR PE that stores input data, two pop-counters, an accumulator, an activation block, a concatenator, and a max pooling block.
[0190] Each block is designed with variable parameters to accommodate various network topologies. NNU is designed with a reusable layer accelerator structure that uses fewer resources than a streaming accelerator, allowing for flexible application across diverse applications. Furthermore, it maximizes acceleration effects through a parallel processing structure that processes data at the channel level and supports adaptive parallel processing per layer, enabling optimized parallel processing at every layer.
[0191] NNU(130) operates by repeating the DS-BTNN layer a specific number of times.
[0192] When each layer operation starts, the parameter SRAM in the processor containing the NNU retrieves the data allocated to the layer from the external DDR. Therefore, the size of the on-chip SRAM is set to fit the layer using the most parameters and currently occupies 18.7kB.
[0193] NNU (130) significantly reduces memory usage by binarizing network inputs and ternaryizing weights according to the target.
[0194] Convolution operations are replaced with bitwise XNOR and pop-count operations, significantly reducing the amount of computation. It supports variable kernel sizes (1x1, 3x3, 5x5), variable I / O channel sizes, and various operation modes, enabling the execution of not only standard CNNs but also DW convolution, PW convolution, and fully connected operations. In addition to the aforementioned features, the quantization level, padding, pooling, bias, and BN of each layer can be freely configured.
[0195] The operation process of the NNU is as follows. The Pop-counter counts the number of [1] and [-1] data in the XNOR result of the 128-bit data. The Accumulator accumulates the pop-count result value in the output channel, and the activation block binarizes the final result value.
[0196] The Concatenator groups the binarized results into 128-bit chunks by channel, and the max pooling block reduces the output image size to one-quarter depending on the selection. The NNU can function as the FCL of the final output layer by scaling the input data and omitting the activation block.
[0198] The dataset adopted for performance comparison of the present invention is the Google speech command dataset (GSCD)
[15] .
[0199] The dataset contains over 105,000 1-second speech data points for 30 keywords. For comparison with reference papers, the words used in command classification consist of a total of 12 keywords: 10 words—"up," "stop," "yes," "right," "go," "down," "on," "off," "no," and "left"—and two others—"silence" and "unknown." For the command classification task, the training set consists of 27,520 data points for learning, and the test set consists of 6,880 data points for validation. For WUW (Wake-Unknown Word) recognition, the keyword was reduced to one, other keywords were treated as "unknown," and the ratio between the data was adjusted; the entire dataset consists of 3,840 training set points and 960 test set points.
[0200] Training was performed on the NVIDIA GeForce RTX A6000 using the cross-entropy loss function Adam optimizer.
[0201] The batch size was set to 256 for command classification and 128 for WUW recognition, and the epoch was set to 100. The learning rate was set to 0.005 for the first 40 epochs, 0.001 for the next 40 epochs, and 0.0005 thereafter.
[0202] The network applied quantization techniques (BNN / TNN) and was configured to be complex to compensate for the resulting decrease in accuracy. Experiments were conducted using various networks to analyze changes in accuracy and parameters based on the number of nodes, CL and FCL configurations, and whether the DSCNN technique was applied. While the number of nodes was kept almost fixed to meet the required accuracy, the number of parameters varied significantly depending on the configuration of CL and FCL and whether the DSCNN technique was applied.
[0203] Looking at the comparison results in [Table 2], it can be seen that the proposed Network No. 4 achieved high accuracy and an optimized number of parameters simultaneously.
[0204] [Table 2]
[0205]
[0206] The final network structure proposed in this invention, DS-BTNN, is composed of a combination of DS-CNN and BTNN, and its structure is shown in Figure 5. The entire network consists of 9 layers (CL + (DW CL + PW CL)x3 + FCLx2), and all layers except the output layer include Batch Normalization (BN) and sign activation, and max pooling is performed in the first layer.
[0208] The verification results of the device shown in FIG. 2a will be explained below.
[0209] Verification of the starter and command recognition device (100) of the present invention is implemented through an FPGA platform having an advanced extensible interface (AXI) bus interface.
[0210] As shown in Fig. 8, the FPGA platform is configured as a KWS system that performs voice signal classification.
[0211] This system includes two processors, each with an I2S2 interface, a microprocessor, double data rate (DDR) memory, an NNU, and an MPU.
[0212] The I2S2 sensor includes an ADC that converts serial voice signals input from a microphone into 16-bit data, and bundles data of a certain size through the I2S2 interface and transmits it to the MPU.
[0213] Each processor includes DDR and I2S2 interfaces, a master interface for communication with the processor, a slave interface for communication with the microprocessor, and cache RAM for storing input / output data.
[0214] In addition, both processors have registers for mode settings, including quantization levels and BN. Furthermore, this system can efficiently perform data transfer using a 64-bit AXI bus interface for high-bandwidth and real-time communication between the processor and other components.
[0215] In addition, the proposed KWS system was designed at the register transfer level (RTL) using Verilog hardware description language (HDL), and then implemented and verified based on Xilinx’s UltraScale+ ZCU104 FPGA.
[0216] The output results from each platform were compared and verified using vectors from a Python-based simulator.
[0217] [Table 3] is a table summarizing the resource usage of the proposed KWS system.
[0218]
[0219] The proposed KWS system was synthesized using a total of 27,315 LUTs and 25,534 FFs, and uses 31 DSPs.
[0220] These figures correspond to 11.7% of the total LUTs and 5.42% of the total FFs based on the UltraScale+ ZCU104 FPGA board, and 0.87% of the total for the DSP.
[0221] Synthesis results confirmed that it can operate at a maximum of 170 MHz. In addition, to compare with a reference paper implemented as an ASIC, synthesis was performed in a CMOS 40 nm process environment to derive the area.
[0222] [Table 4] shows a comparison between the KWS system presented in the present invention and other KWS systems.
[0223] All subjects for comparison used GSCD and were implemented on ASICs with different quantization strategies. The comparison focused on the number of target keywords, network accuracy, on-chip SRAM size, and normalized area.
[0224] [Table 4]
[0225]
[0227] Refer to Table 4. In the case of [2], 1 to 2 keywords are recognized using BNN. While it supports a small number of keywords and has a small memory size and area due to the application of quantization networks and DS-CNN techniques, there is a limitation in that the number of keywords to be detected cannot be increased due to the decrease in accuracy of BNN. The proposed paper can perform 10 command recognitions by applying TNN to achieve high accuracy with a computational flow similar to BNN.
[0228] [4, 5, 6] support 10 keyword classification functions, identical to the proposed invention. [4] has a full BWN structure with only the weights binarized. By applying a pipelining structure rather than layer-unit operations, the on-chip SRAM size is minimized, and the network structure is concise due to the high bit precision of the data, but it has a larger area compared to the proposed processor.
[0229] [5] presents a system that adjusts the quantization bits of the input data and weights to 4 / 8 depending on the noise in the input data. It has an advantage in terms of accuracy because it uses up to 8 bits of input data and weights for the same dataset, but the memory size is 69kB and the total chip area is 1.42mm 2 It consumes relatively large resources compared to the present paper. In addition, the present paper can operate with minimized parameters by partially applying BNN to 1 to 2 keywords as well as 10 keywords.
[0231] The present invention takes the form of a system in which a voice input sensor, an MPU, and an NNU are integrated to perform the intended purpose of an ASR system on a single device.
[0232] Data determined to be valid speech is converted into Mel-spectrograms and used as input data for the network; if the speech is recognized after initially passing through a BNN for WUW recognition, it passes through a TNN for command classification through an iterative process to determine the command.
[0233] To apply a complex network structure to a limited hardware structure, the following two network lightweighting techniques were used.
[0234] DS-CNN divides conventional convolutional layers into separate layers with local and global characteristics for computation, and utilizes a quantization network technique that quantizes input data and weights to a low bit rate. While the network achieves low memory occupancy through lightweight design, it has been expanded into a deep and complex structure to compensate for accuracy; this minimizes overhead through parallel computation structures, pipelining, and the DS-CNN architecture.
[0235] In addition, the network is designed to support networks with various purposes, in addition to verified keyword types and numbers, by allowing representative network components such as padding, number of layers, BN application, kernel size, number of channels, and quantization level to be variable.
[0236] In the case of reference papers, BNNs suffered from severe accuracy degradation and struggled to classify 10 keywords, so a quantized network method with high bit-precision was adopted. In the case of the present invention, the load on complex network configuration was mitigated through DS-CNN and a parallel hardware structure while simultaneously ternaryizing the weights.
[0237] In addition, by utilizing the similarity in the computational processes for binarization and ternaryization, it was possible to apply memory-optimized quantization networks according to network objectives.
[0238] NNU's DS-BTNN layers are designed with a reusable block structure, enabling layer-by-layer configuration, and the on-chip SRAM size is determined by the maximum number of parameters used in a single layer.
[0239] It demonstrated accuracy of 97.1% and 90.5% in BNN-based WUW recognition and TNN-based instruction classification, respectively, and with 18.7kB of SRAM, 0.558mm in a 40nm process environment 2 The area of was derived.
[0241] The “part” used in one embodiment of the present invention may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.
[0242] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0243] A method according to an embodiment of the present invention may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0244] The embodiments described above are merely examples, and those skilled in the art to which the present invention pertains will be able to make various modifications and variations within the scope of the essential characteristics of the present invention. Accordingly, the embodiments disclosed in the present invention are intended to explain, not limit, the technical concept of the present invention, and the scope of the technical concept of the present invention is not limited by these embodiments. The scope of protection of the present invention shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present invention. Explanation of the symbols
[0245] 100: Starter and command recognition device 110: A / D converter 120: MPU 130: NNU
Claims
Claim 1 MPU that preprocesses voice data into Mel-spectrograms; A wake-up-word and command recognition device comprising an NNU that applies the above Mel-spectrogram to a binarized classification network to recognize a wake-up-word of voice data from the above Mel-spectrogram, converts to a ternary classification network, and subsequently applies an input Mel-spectrogram to the ternary classification network to classify commands within the voice data, wherein the MPU includes an energy-based VAD that verifies the validity of data continuously input through the pmod terminal of the ZCU104 board; an STFT unit that generates a spectrogram analyzing frequency changes over time of continuous data; and a Mel-filtering unit that performs Mel-filtering operations on the spectrogram to reflect the characteristics of human auditory structure, wherein the STFT unit performs the STFT process through FFT operations every 256 points using a DIF-based fixed-point radix-4 algorithm, receives serialized data as input, performs radix-4 operations using a buffer, and stores partial result values of the FFT in a 32x128 local memory. Claim 2 A command and instruction recognition device according to claim 1, further comprising an A / D converter that converts an analog voice signal into the voice data, which is a digital signal. Claim 3 delete Claim 4 A starter and command recognition device according to claim 1, wherein the energy-based VAD is configured to detect actual voice intervals within the voice data, and is characterized by recognizing as the starting point of a valid voice signal when the difference between the energy level within a specific interval and the energy level of the previous interval exceeds a threshold value. Claim 5 A starter and command recognition device according to claim 1, wherein the energy-based VAD includes an address controller and a memory internally, and the address controller stores data in the memory with address values from 0 to 127 for every 128 inputs. Claim 6 A starter and instruction recognition device according to claim 1, wherein the STFT unit starts a 256-point FFT by filling data with address values between 128 and 255 after validation, and subsequently performs a 256-point FFT when there is 50% overlap by storing a value at an arbitrary address for every 128 data inputs. Claim 7 delete Claim 8 delete Claim 9 A command and instruction recognition device according to claim 1, wherein the STFT unit derives 129 non-duplicate complex values during FFT operation and derives absolute values required for MEL filtering through bit-shift and arithmetic operations. Claim 10 A starter and command recognition device according to claim 9, wherein the above-mentioned Mel-filtering unit comprises a Mel-filter bank composed of a square filter converted by quantizing a triangular filter. Claim 11 A starter and command recognition device according to claim 10, wherein the NNU comprises an XNOR PE for storing input data, two pop-counters, an accumulator, an activation block, a concatenator, and a max pooling block, and each block is designed with variable parameters to accommodate various network topologies. Claim 12 In claim 11, the above NNU is designed as a layer accelerator structure, which is a reusable parallel computing structure with low resource usage, and is characterized by supporting a layer-by-layer adaptive parallel computing method, thereby forming a command and instruction recognition device. Claim 13 A starter and instruction recognition device according to claim 12, wherein the above NNU operates to reduce memory usage by binarizing network inputs and binarizing or ternaryizing weights according to a target. Claim 14 A command and instruction recognition device according to claim 13, wherein the above NNU further comprises a parameter SRAM for retrieving data assigned to a layer from an external DDR. Claim 15 In claim 14, the above NNU is a command and instruction recognition device characterized by replacing convolution operations with bit-unit XNOR and PoP-Count operations. Claim 16 A starter and instruction recognition device according to claim 15, wherein the above NNU supports a function to set at least one of the quantization level, padding, pooling, bias, and BN of each layer. Claim 17 In claim 16, the pop-counter counts the number of [1] and [-1] data in the XNOR result of 128-bit data, the accumulator accumulates the pop-counter result value in an output channel, the activation block binarizes the final result value, the concatenator groups the binarized result value into 128-bit units per channel, and the max pooling block reduces the output image size to 1 / 4 depending on the selection, characterized by a starter and instruction recognition device. Claim 18 A method for recognizing a starter word and a command, performed by a starter word and command recognition device comprising a Mel-Processing Unit (MPU) and a Neural Network Unit (NNU), wherein, in the MPU, voice data is preprocessed into a Mel-spectrogram; A method for recognizing wake-up words and commands, comprising the steps of: applying the Mel-spectrogram to a binary classification network to recognize a wake-up word of voice data from the Mel-spectrogram; converting to a ternary classification network; and subsequently applying the input Mel-spectrogram to the ternary classification network to classify commands within the voice data; wherein the preprocessing step comprises the step of performing an STFT process by performing an FFT operation every 256 points using a fixed-point radix-4 algorithm of the DIF method in the STFT unit; and wherein the preprocessing step comprises the step of receiving serialized data as input in the STFT unit, performing a radix-4 operation using a buffer, and storing partial result values of the FFT in a 32x128 local memory. Claim 19 A method for recognizing a starter and command according to claim 18, further comprising the step of converting an analog voice signal into the voice data, which is a digital signal, through an A / D converter. Claim 20 In claim 19, the preprocessing step includes a step of recognizing a valid voice signal starting point when the difference between the energy level within a specific interval and the energy level of the previous interval exceeds a threshold value when detecting an actual voice interval within the voice data in the VAD. Claim 21 A method for recognizing a starter and an instruction according to claim 20, characterized in that the preprocessing step includes the step of starting a 256-point FFT by filling data with address values between 128 and 255 after validation in the STFT section, and then performing a 256-point FFT when there is 50% overlap by storing a value at an arbitrary address for every 128 data inputs. Claim 22 delete Claim 23 delete Claim 24 A method for recognizing starter words and commands according to claim 18, wherein the preprocessing step comprises deriving 129 non-duplicate complex values during FFT operation in the STFT unit and deriving absolute values required for MEL filtering through bit-shift and arithmetic operations. Claim 25 In claim 24, the step of classifying the above-mentioned instruction comprises counting the number of [1] and [-1] data from the XNOR result of 128-bit data in the pop-counter of the NNU, accumulating the pop-counter result value in the output channel in the accumulator of the NNU, binarizing the final result value in the activation block of the NNU, grouping the binarized result value into 128-bit units by channel in the concatenator of the NNU, and reducing the output image size to 1 / 4 according to selection in the max pooling block of the NNU, characterized in that it is a method for recognizing a starter and instruction.