Keyword spotting method and apparatus based on magnetic tunnel junction arrays

By using a speech recognition method based on magnetic tunnel junction arrays, the neural network weights are mapped to the resistance states of the magnetic tunnel junction array. Combined with Mel-frequency cepstral coefficient features and current readout operations, the problem of high hardware resource consumption in edge computing is solved, and efficient and accurate speech keyword recognition is achieved.

WO2026113130A1PCT designated stage Publication Date: 2026-06-04INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD
Filing Date
2025-01-15
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing speech keyword recognition methods suffer from low accuracy and high hardware resource consumption in edge computing scenarios with limited hardware resources. Traditional convolutional neural networks and binarization networks have disadvantages in terms of storage and energy efficiency, and cannot be effectively applied to magnetic tunnel junction arrays.

Method used

A keyword-based speech recognition method based on magnetic tunnel junction array is adopted. The neural network weight data is mapped to the high and low resistance states of the magnetic tunnel junction array through statistical perception training. Mel-frequency cepstral coefficient feature extraction and current readout operation are used for speech keyword recognition, realizing the integrated application of in-memory computing and neuromorphic computing.

Benefits of technology

It achieves high-speed and energy-efficient voice data processing in edge computing scenarios, reduces hardware resource consumption, improves recognition accuracy, and enhances the network's robustness to hardware defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072376_04062026_PF_FP_ABST
    Figure CN2025072376_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A keyword spotting method and apparatus based on magnetic tunnel junction arrays. The method comprises: using a statistically-aware training method to train a neural network, so as to obtain weight data of the neural network, mapping the weight data to high and low resistance state data of magnetic tunnel junction arrays, on the basis of an actual network structure of the neural network, selecting a magnetic tunnel junction array of a corresponding size, and pre-programming magnetic tunnel junctions into corresponding resistance states (S101); extracting a Mel-frequency cepstral coefficient (MFCC) feature from a speech keyword signal (S102); and during performing keyword spotting, continuously inputting binarized speech MFCC features into the magnetic tunnel junction array in the form of voltage amplitudes, acquiring a multiply-accumulate operation result by means of measuring an output current of a target column of the magnetic tunnel junction array, and finally outputting a hardware-based keyword spotting result after performing normalization and layer-by-layer inference (S103).
Need to check novelty before this filing date? Find Prior Art

Description

Keyword Speech Recognition Method and Device Based on Magnetic Tunneling Array

[0001] This application claims priority to Chinese Patent Application No. 2024117061576, filed on November 26, 2024, entitled "Keyword Speech Recognition Method and Apparatus Based on Magnetic Tunnel Array", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This document relates to the field of computer technology, and in particular to a keyword speech recognition method and device based on a magnetic tunneling array. Background Technology

[0003] With the continuous development of artificial intelligence (AI), image and speech recognition and processing technologies, as a crucial foundation for human-computer interaction, are one of the important research directions in the field of AI. Keyword recognition (KWS) is a core step in speech information processing. In edge computing scenarios with limited hardware resources, the accuracy and hardware resource consumption of KWS are the most concerning performance indicators. Researching KWS methods with lower hardware resource overhead while maintaining accuracy is a key direction for improving speech recognition performance and expanding its application scenarios.

[0004] Convolutional Neural Networks (CNNs) can automatically extract features through the convolution process and have excellent accuracy and other performance characteristics, making them widely used in the field of neural network object recognition. However, traditional CNNs, in order to meet the complex task requirements, have complex structures, a large number of parameters, and require a lot of computing and storage resources, making them difficult to apply in edge computing scenarios with extremely limited hardware resources. The memory wall and power wall problems limit the large-scale integration of CNNs into computing hardware platforms based on the traditional von Neumann architecture and their application in edge computing scenarios. With the key dimensions of semiconductor devices gradually entering the sub-3nm node and the rapid development of artificial intelligence technology, in-memory computing architecture has attracted much attention as a new architecture that breaks through the limitations of the von Neumann architecture. Magnetic random access memory (MRAM) has the advantages of simple read and write, low read and write power consumption, high stability, high storage density, and data non-volatility. It is fully compatible with CMOS technology and is a strong competitor to the mainstream in-memory computing devices, and has been extensively studied in fields such as neuromorphic computing. Meanwhile, Deep Separable Neural Networks (DS-CNN) significantly simplify the network structure and reduce computational and storage overhead by step-by-step reducing the dimensionality of the three-dimensional convolutional kernels of traditional convolutional neural networks, while keeping the accuracy of the neural network basically unchanged. It has great potential for hardware applications in the field of edge computing.

[0005] Existing technologies provide a binary speech keyword recognition neural network based on MFCC feature extraction, which can be deployed on FPGAs for hardware acceleration. However, the network model described therein is a traditional deep neural network (DNN), with a larger number of weights and higher storage overhead than the present invention. Furthermore, FPGAs cannot achieve an in-memory computing architecture compared to magnetic tunnel junction arrays, resulting in significant disadvantages in terms of energy efficiency and read / write speeds.

[0006] Existing technology also provides a low-occupancy speech keyword recognition technique based on convolutional recurrent neural networks (CRNNs), which can identify and classify input speech keyword information. However, the network structure still requires ~230K network weights, resulting in a large storage footprint, and it is not implemented in hardware. Summary of the Invention

[0007] The purpose of this invention is to provide a keyword speech recognition method and apparatus based on a magnetic tunneling array, aiming to solve the above-mentioned problems in the prior art.

[0008] This invention provides a keyword speech recognition method based on a magnetic tunneling junction array, comprising:

[0009] The neural network is trained using a statistical sensing training method to obtain the weight data of the neural network. The weight data is then mapped to the high and low resistance state data of the magnetic tunnel junction array. Based on the actual network structure of the neural network, a magnetic tunnel junction array of appropriate size is selected, and the magnetic tunnel junction is pre-written with the corresponding resistance state.

[0010] Extracting Mel-frequency cepstral coefficients (MFCC) features from speech keyword signals;

[0011] When performing speech keyword recognition, the binarized speech MFCC features are continuously input into the magnetic tunnel junction array in the form of voltage amplitude. The multiplication and accumulation calculation results are obtained by measuring the output current of the target column of the magnetic tunnel junction array. After normalization and layer-by-layer reasoning, the hardware keyword recognition results are finally output.

[0012] This invention provides a keyword speech recognition device based on a magnetic tunneling junction array, comprising:

[0013] The weight mapping module is used to train a neural network using a statistical sensing training method to obtain the weight data of the neural network, map the weight data to the high and low resistance state data of the magnetic tunnel junction array, select a magnetic tunnel junction array of appropriate size according to the actual network structure of the neural network, and pre-write the magnetic tunnel junction to the corresponding resistance state.

[0014] The speech feature extraction module is used to extract Mel-frequency cepstral coefficients (MFCCs) features from speech keyword signals;

[0015] The speech keyword recognition module is used to continuously input the binarized speech MFCC features in the form of voltage amplitude into the magnetic tunnel junction array when performing speech keyword recognition. The multiplication and accumulation calculation results are obtained by measuring the output current of the target column of the magnetic tunnel junction array. After normalization and layer-by-layer reasoning, the hardware keyword recognition results are finally output.

[0016] This invention also provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the keyword speech recognition method based on a magnetic tunneling array described above.

[0017] This invention also provides a computer-readable storage medium storing an information transmission implementation program, which, when executed by a processor, implements the steps of the keyword speech recognition method based on a magnetic tunneling array described above.

[0018] The embodiments of the present invention are compatible with existing CMOS integration processes based on the magnetic tunnel junction core structure, which is conducive to large-scale fabrication and thus helps to realize the integrated application of in-memory computing and neuromorphic computing, such as voice data processing in high-speed, high-energy-efficiency edge computing scenarios. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 is a flowchart of the keyword speech recognition method based on magnetic tunneling array according to an embodiment of the present invention;

[0021] Figure 2 is a schematic diagram of the training process of obtaining DS-CNN network weights by the statistical perception training method of this invention and the performance index of DS-CNN in the host computer simulation experiment.

[0022] Figure 3 is a schematic diagram of the neural network structure of DS-CNN according to an embodiment of the present invention;

[0023] Figure 4 is a schematic diagram of the defect matrix generation method applied to statistical perception training according to an embodiment of the present invention;

[0024] Figure 5 is a schematic diagram of the neural network weight mapping rule according to an embodiment of the present invention;

[0025] Figure 6 is a schematic diagram of the process of MFCC feature extraction of speech keyword signals according to an embodiment of the present invention;

[0026] Figure 7 is a schematic diagram illustrating the principle of convolution operation based on magnetic tunnel junction array in an embodiment of the present invention;

[0027] Figure 8 is a schematic diagram of the hardware implementation of obtaining the final keyword recognition result in the speech keyword recognition stage of an embodiment of the present invention;

[0028] Figure 9 is a schematic diagram of the speech keyword recognition process based on magnetic tunneling array according to an embodiment of the present invention;

[0029] Figure 10 is a schematic diagram of a keyword speech recognition device based on a magnetic tunneling array according to an embodiment of the present invention;

[0030] Figure 11 is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.

[0032] Method Implementation Examples

[0033] According to an embodiment of the present invention, a keyword speech recognition method based on a magnetic tunneling junction array is provided. Figure 1 is a flowchart of the keyword speech recognition method based on a magnetic tunneling junction array according to an embodiment of the present invention. As shown in Figure 1, the keyword speech recognition method based on a magnetic tunneling junction array according to an embodiment of the present invention specifically includes:

[0034] Step S101: Train the neural network using a statistical sensing training method to obtain the weight data of the neural network. Map the weight data to the high and low resistance state data of the magnetic tunnel junction array. Based on the actual network structure of the neural network, select a magnetic tunnel junction array of appropriate size and pre-write the corresponding resistance state into the magnetic tunnel junction. Specifically, this includes:

[0035] Using a statistical sensing training method, weight data for a binary, deeply separable neural network robust to non-ideal hardware is obtained. The trained weight matrix is ​​mapped onto a magnetic tunnel junction array. A single magnetic tunnel junction is used to form a single weight. Different weights are mapped to different resistance state configurations of the magnetic tunnel junction. Based on the structure of the neural network and the number of weights, a magnetic tunnel junction array of corresponding size is selected. The magnetic tunnel junction is written into the corresponding resistance state through spin-transfer torque or spin-orbit torque. Wherein, the magnetic tunnel junction is in a parallel state and mapped to a weight "1"; the magnetic tunnel junction is in an antiparallel state and mapped to a weight "0".

[0036] The magnetic tunnel junction array is at least one of the following: a passive STT-MRAM array, an active STT-MRAM array, an active SOT-MRAM array, or a three-dimensional stacked SOT-MRAM array.

[0037] Step S102, extracting Mel-frequency cepstral coefficients (MFCC) features from the speech keyword signal; specifically including:

[0038] Mel-frequency cepstral coefficients (MFCC) feature extraction is performed on the input speech keyword signal. The window length, window step size, and number of MFCC feature channels are selected in the MFCC processing according to the network structure of the neural network and actual needs, so as to adjust the size of the final MFCC two-dimensional feature matrix input to the neural network.

[0039] Step S103: During speech keyword recognition, the binarized speech MFCC features are continuously input into the magnetic tunnel junction array in the form of voltage amplitude. The multiplication and accumulation calculation results are obtained by measuring the output current of the target column of the magnetic tunnel junction array. After normalization and layer-by-layer inference, the hardware keyword recognition results are finally output. Specifically, this includes:

[0040] The speech MFCC feature matrix P is mapped to the read voltage matrix V, where each column of the read voltage matrix V is sequentially input into the column of the corresponding convolution kernel of the magnetic tunnel junction array;

[0041] Based on Ohm's law and Kirchhoff's law, the convolution operation is equivalent to the read operation of the magnetic tunnel junction array. The output current of each column of the magnetic tunnel junction array corresponding to the convolution result is summed to calculate the output current of each column of the magnetic tunnel junction under the corresponding read voltage combination, and an intermediate feature output matrix is ​​generated.

[0042] The intermediate feature output matrix is ​​binarized to form a new input matrix P1, which is mapped to the read voltage matrix V1. The read voltage input and current accumulation reading operations of the subsequent columns of the magnetic tunnel junction array are performed again to obtain intermediate results P2 and V2. This process continues until the inference of each layer of the network is completed, and finally the speech keyword recognition result is output.

[0043] As described above, the process begins by training a neural network using a statistical sensing training method, mapping the network weights to the high and low resistance states of a magnetic tunnel junction (MTJ). Then, Mel-frequency cepstral coefficients (MFCCs) are extracted from the speech keyword signal and input as voltage amplitude values ​​to the MTJ array, mapped to read voltage amplitudes. Based on the actual network structure, an MTJ array of appropriate size is selected, and the corresponding resistance states are pre-written into the MTJs. During the actual speech keyword recognition task, the binarized speech MFCC features are continuously input to the MTJ array as voltage amplitude values. The multiplication and accumulation calculation results are obtained by measuring the output current of the target column of the MTJ array. After normalization and layer-by-layer inference, the final hardware keyword recognition result is output. This technical solution is compatible with existing CMOS integration processes based on the MTJ core structure, facilitating large-scale fabrication and enabling integrated applications of in-memory computing and neuromorphic computing, such as high-speed, high-energy-efficiency speech data processing in edge computing scenarios.

[0044] The technical solutions of the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0045] The technical solution of this invention specifically includes: a weight mapping stage, a speech feature extraction stage, and a speech keyword recognition stage. In the weight mapping stage, weight data of the neural network is obtained through statistical perceptual training and mapped to high and low impedance state data of the magnetic tunnel junction array. In the speech feature extraction stage, MFCC features of the speech keyword signal are extracted to generate an input feature matrix. In the speech keyword recognition stage, the readout operation of the magnetic tunnel junction is used to replace the convolution operation between the weight matrix and the input speech feature data; the sum of the readout currents of each column of the magnetic tunnel junction array is the new intermediate feature value obtained by convolution and the recognition result.

[0046] First, in the weight mapping stage, a statistical sensing training method is used to obtain binarized, deeply separable neural network weights robust to non-ideal hardware. Then, the trained weight matrix is ​​mapped onto a magnetic tunnel junction array. A single magnetic tunnel junction is used to form a single weight. Different weights are mapped to different resistance state configurations of the magnetic tunnel junction. Specifically, a parallel state (P, low resistance state) is mapped to a weight "1"; an antiparallel state (AP, high resistance state) is mapped to a weight "0". Based on the neural network structure and the number of weights, a magnetic tunnel junction array of corresponding size is selected, and the magnetic tunnel junctions are written into the corresponding resistance states using spin-transfer torque or spin-orbit torque.

[0047] Secondly, in the speech feature extraction stage, MFCC features are extracted from the input speech keyword signal. Based on the network structure and actual needs, parameters such as window length, window step size, and the number of MFCC feature channels in the MFCC processing are selected to adjust the size of the final MFCC two-dimensional feature matrix input to the neural network.

[0048] In the speech keyword recognition stage, the speech MFCC feature matrix P is mapped to the read voltage matrix V. Each column of the read voltage matrix V is sequentially input to the column containing the corresponding convolution kernel of the magnetic tunnel junction array. According to Ohm's law and Kirchhoff's laws, the convolution operation, i.e., the multiplication-addition operation, is equivalent to the read operation of the magnetic tunnel junction array. The convolution result corresponds to the sum of the output currents I of each column of the magnetic tunnel junction array. out =∑ i,j V ij *G ij The output current of each column of the magnetic tunnel junction is measured under the corresponding read voltage combination, and an intermediate feature output matrix is ​​generated. The intermediate feature output matrix is ​​re-binarized to form a new input matrix P1, which is mapped to the read voltage matrix V1. The read voltage input and current accumulation reading operations of the subsequent columns of the magnetic tunnel junction array are performed again to obtain intermediate results such as P2 and V2, until the inference of each layer of the network is completed, and the array finally outputs the speech keyword recognition result.

[0049] In practical applications, the embodiments of the present invention can implement different network structures; use different speech feature extraction methods; use different magnetic tunnel junction arrays; and use different read voltage mapping ranges and step sizes.

[0050] Figure 2 illustrates the training process of the DS-CNN network weights obtained by the statistical perception training method in this embodiment of the invention, and the performance indicators of DS-CNN in the host computer simulation experiment. Inputting MFCC feature matrix training data after Wasat processing for different defect levels, the neural network is trained using gradient descent for backpropagation to obtain network weights robust to hardware non-ideals. The array return result is the output value of the fully connected layer, representing the return values ​​of target words and non-target words respectively. After regularization, these are converted into the discrimination probabilities of target words and non-target words, ultimately yielding the recognition result. In this example, the same non-target keyword group was used for training on different target keywords, and the recognition accuracy of different target keywords is shown in the figure.

[0051] Figure 3 illustrates the neural network structure of DS-CNN in this embodiment of the invention. DS-CNN consists of several convolutional layers, depthwise separable convolutional layers, pooling layers, and fully connected layers. The number and allocation of network layers, the operator size and stride of each layer, and the network connection method are adjusted according to specific application requirements. In this specific embodiment, the network structure shown in Figure 2 is used, consisting of one convolutional layer, one separable convolutional layer, one pooling layer, and one fully connected layer. The illustrated structure requires a total of 790 binarized weights.

[0052] Figure 4 illustrates the defect matrix generation method applied to statistical sensing training in this embodiment of the invention. When training the neural network weights, a random defect matrix within the upper limit of the defect quantity (Wsat) is generated, randomly defecting the training dataset to simulate non-ideal factors such as random single-device failures in the hardware array. By repeatedly training with multiple batches of defect matrices of different Wasats, the robustness of the neural network weights to non-ideal hardware factors is improved.

[0053] Figure 5 illustrates the neural network weight mapping rules of this embodiment. When hardware-encapsulating the weights, a single magnetic tunnel structure is represented by a single weight, and different weights are mapped to different resistance state configurations of the magnetic tunnel junction. Specifically, a magnetic tunnel junction in a parallel state (P, low resistance state) is mapped to a weight "1"; a magnetic tunnel junction in an antiparallel state (AP, high resistance state) is mapped to a weight "0". In this example, each weight of the DS-CNN in Figure 2 occupies 1 bit of storage resource, i.e., one magnetic tunnel junction. The total weights of the complete neural network occupy less than 1 Kbit. The weights of each layer of the DS-CNN in this embodiment are written into a 32×32 magnetic tunnel junction array as shown in Figure 3.

[0054] Figure 6 illustrates the process of MFCC feature extraction for speech keyword signals according to an embodiment of the present invention. The MFCC feature extraction of the speech signal consists of pre-emphasis, framing, windowing, frequency domain transformation, Mel-Cepstral Filtering (MCF), obtaining log-DCT values, and calculating the first and second order differences. The engineering parameters in the MFCC feature extraction process can be modified according to the characteristics of the speech keyword signal, such as its duration and sampling rate. The final output of the MFCC feature extraction process is an N×D MFCC feature matrix, where N is the number of frames after signal framing, and D is the dimension of the Mel-Cepstral Filter. In this specific embodiment, a speech keyword signal with a length of 1 second and a sampling rate of 16000Hz is framed, with 1024 sampling points per frame, a step size of 512, and N = 30; the dimension of the Mel-Cepstral Filter is D = 40, resulting in a 30×40 MFCC feature matrix. To simplify the neural network task requirements and reduce hardware resource overhead, the first 10 dimensions of the features are used as input, resulting in a 30×10 MFCC feature matrix.

[0055] Figure 7 illustrates the principle of convolution operation based on magnetic tunnel junction array in an embodiment of the present invention. In this specific embodiment, taking the first layer of conventional convolution as an example, its convolution kernel is a 3×3×10 three-dimensional operator, with each dimension being a 3×3 convolution operator. This can be represented by the following variables:

[0056] Depending on the actual needs, the size of each convolutional kernel and the specific weight values ​​can be adjusted. Therefore, the scope of this patent also includes different convolutional kernels and neural network weight values.

[0057] According to the neural network weight mapping rule of this invention, 3×3 one-dimensional convolution weights are mapped to the corresponding resistance states of the magnetic tunnel junction array. First, the weights are reduced to one dimension:

[0058] The conductance matrix corresponding to the magnetic tunnel junction:

[0059] As mentioned earlier, the weight values ​​w in the weight matrix S ij The 1 corresponds to the conductance matrix G P (Low resistance state), w ij The zero-corresponding derivative matrix G AP (High-impedance state). As shown in Figure 2, the MFCC feature matrix P of the speech keywords is mapped to the read voltage matrix V. According to Ohm's law and Kirchhoff's law, the convolution operation of the MFCC feature matrix P with a one-dimensional convolution kernel is directly equivalent to the read operation of the magnetic tunnel junction array under the corresponding read voltage. out =∑ i,j V ij *G ij By sequentially inputting the voltages of each column of the voltage matrix and reading the magnetic tunneling current corresponding to the current convolution kernel, the result of this convolution can be obtained and used as the intermediate input feature for the next layer of convolution.

[0060] Figure 8 illustrates the hardware implementation diagram of the speech keyword recognition stage of this embodiment of the invention, showing how the magnetic tunneling junction array outputs intermediate features, performs layer-by-layer reasoning, and finally obtains the keyword recognition result. The input MFCC feature matrix P is mapped to a read voltage matrix V, and the intermediate result is output after convolution operation of the corresponding layer operator of the neural network on the magnetic tunneling junction array. In this specific embodiment, the FPGA controls peripheral circuit components such as digital-to-analog converters and analog-to-digital converters to perform read and write operations on the magnetic tunneling junction. The 30×10 MFCC feature matrix is ​​input column by column into the magnetic tunneling junction array after weight mapping in the form of a read voltage matrix. During the first layer of conventional convolution, a column of read voltage is applied multiple times, and the accumulated current of each column of the magnetic tunneling junction mapped to the first layer convolution kernel is read to obtain the intermediate features. This operation is repeated for each column of the voltage matrix to obtain the intermediate feature matrix P1. The intermediate feature matrix P1 is mapped to a new read voltage matrix V1 and re-input into the magnetic tunneling junction array. For each column of voltage, the column current corresponding to the second layer depth-separable convolution kernel is read to obtain the intermediate feature matrix P2. Similarly, the results are processed through pooling and fully connected layers, with the output of the fully connected layer being the final result of speech keyword recognition.

[0061] Figure 9 illustrates the flowchart of the speech keyword recognition based on a magnetic tunnel junction array according to an embodiment of the present invention. The entire process is divided into a weight mapping stage, a speech feature extraction stage, and a speech keyword recognition stage. In the weight mapping stage, statistical perception training is performed according to the defect matrix generated as shown in Figure 4, and then the DS-CNN weights are mapped to the magnetic tunnel junction array according to the weight mapping principle shown in Figure 5, and the corresponding resistance states are written. In the speech feature extraction stage, the input speech keyword signal is extracted into a 30×10 MFCC feature matrix according to the MFCC speech feature extraction process shown in Figure 6. In the speech keyword recognition stage, the MFCC feature matrix is ​​input into the magnetic tunnel junction array in the manner of reading voltage mapping according to the convolution principle shown in Figure 7. Multi-layer neural network convolution operations are performed according to the hardware structure shown in Figure 8, and finally the speech keyword recognition result is obtained.

[0062] The beneficial effects of the embodiments of the present invention are as follows:

[0063] In terms of hardware, the neural network convolution operation is implemented through the current readout operation of the magnetic tunnel junction array. One convolution operation can be completed within a single clock cycle, which improves parallelism and convolution efficiency. The single recognition latency is <...μs and the energy consumption is <...J.

[0064] On the software side, a binarized DS-CNN neural network model is used, optimizing the number of network weights to <1k and the storage of network weights to <1k bit, which greatly reduces the hardware resource overhead of speech keyword recognition while maintaining a recognition accuracy of >90%.

[0065] During network training, statistical sensing training methods were applied to enhance the robustness of network weights to array physical defects and the versatility of weights for different defective devices at the software level.

[0066] Device Example 1

[0067] According to an embodiment of the present invention, a keyword speech recognition device based on a magnetic tunneling junction array is provided. Figure 10 is a schematic diagram of the keyword speech recognition device based on a magnetic tunneling junction array according to an embodiment of the present invention. As shown in Figure 10, the keyword speech recognition device based on a magnetic tunneling junction array according to an embodiment of the present invention specifically includes:

[0068] The weight mapping module 100 is used to train a neural network using a statistical sensing training method to obtain the weight data of the neural network, map the weight data to the high and low resistance state data of the magnetic tunnel junction array, select a magnetic tunnel junction array of appropriate size according to the actual network structure of the neural network, and pre-write the corresponding resistance state of the magnetic tunnel junction; specifically used for:

[0069] Using a statistical sensing training method, weight data for a binary, deeply separable neural network robust to non-ideal hardware is obtained. The trained weight matrix is ​​mapped onto a magnetic tunnel junction array. A single magnetic tunnel junction is used to form a single weight. Different weights are mapped to different resistance state configurations of the magnetic tunnel junction. Based on the structure of the neural network and the number of weights, a magnetic tunnel junction array of corresponding size is selected. The magnetic tunnel junction is written into the corresponding resistance state through spin-transfer torque or spin-orbit torque. Specifically, a parallel magnetic tunnel junction is mapped to a weight of "1", and an antiparallel magnetic tunnel junction is mapped to a weight of "0".

[0070] The speech feature extraction module 102 is used to extract Mel-frequency cepstral coefficients (MFCC) features from the speech keyword signal; specifically, it is used for:

[0071] MFCC feature extraction is performed on the input speech keyword signal. The window length, window step size, and number of MFCC feature channels in the MFCC processing are selected according to the network structure of the neural network and actual needs, so as to adjust the size of the final MFCC two-dimensional feature matrix input to the neural network.

[0072] The speech keyword recognition module 104 is used to continuously input the binarized speech MFCC features as voltage amplitudes into the magnetic tunnel junction array during speech keyword recognition. It obtains the multiplication-accumulation calculation results by measuring the output current of the target column of the magnetic tunnel junction array, and after normalization and layer-by-layer inference, finally outputs the hardware keyword recognition results. Specifically, it is used for:

[0073] The speech MFCC feature matrix P is mapped to the read voltage matrix V, where each column of the read voltage matrix V is sequentially input into the column of the corresponding convolution kernel of the magnetic tunnel junction array;

[0074] Based on Ohm's law and Kirchhoff's law, the convolution operation is equivalent to the read operation of the magnetic tunnel junction array. The output current of each column of the magnetic tunnel junction array corresponding to the convolution result is summed to calculate the output current of each column of the magnetic tunnel junction under the corresponding read voltage combination, and an intermediate feature output matrix is ​​generated.

[0075] The intermediate feature output matrix is ​​binarized to form a new input matrix P1, which is mapped to the read voltage matrix V1. The read voltage input and current accumulation reading operations of the subsequent columns of the magnetic tunnel junction array are performed again to obtain intermediate results P2 and V2. This process continues until the inference of each layer of the network is completed, and finally the speech keyword recognition result is output.

[0076] The aforementioned magnetic tunnel junction array is at least one of the following: passive STT-MRAM array, active STT-MRAM array, active SOT-MRAM array, or three-dimensional stacked SOT-MRAM array.

[0077] The embodiments of the present invention are device embodiments corresponding to the above method embodiments. The specific operation of each module can be understood with reference to the description of the method embodiments, and will not be repeated here.

[0078] Device Example 2

[0079] An embodiment of the present invention provides an electronic device, as shown in FIG11, including: a memory 110, a processor 112, and a computer program stored in the memory 110 and executable on the processor 112. When the computer program is executed by the processor 112, it implements the steps described in the method embodiment.

[0080] Device Example 3

[0081] This invention provides a computer-readable storage medium storing an information transmission implementation program, which, when executed by a processor 112, implements the steps described in the method embodiment.

[0082] The computer-readable storage media described in this embodiment include, but are not limited to, ROM, RAM, disk, or optical disk.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A keyword speech recognition method based on a magnetic tunnel junction array, characterized by, include: The neural network is trained using a statistical sensing training method to obtain the weight data of the neural network. The weight data is then mapped to the high and low resistance state data of the magnetic tunnel junction array. Based on the actual network structure of the neural network, a magnetic tunnel junction array of appropriate size is selected, and the magnetic tunnel junction is pre-written with the corresponding resistance state. Extracting Mel-frequency cepstral coefficients (MFCC) features from speech keyword signals; When performing speech keyword recognition, the binarized speech MFCC features are continuously input into the magnetic tunnel junction array in the form of voltage amplitude. The multiplication and accumulation calculation results are obtained by measuring the output current of the target column of the magnetic tunnel junction array. After normalization and layer-by-layer reasoning, the hardware keyword recognition results are finally output.

2. The method of claim 1, wherein, A neural network is trained using a statistical sensing training method to obtain the network's weight data. This weight data is then mapped to the high and low resistance states of a magnetic tunnel junction array. Based on the actual network structure, a magnetic tunnel junction array of appropriate size is selected. The pre-writing of the magnetic tunnel junction with the corresponding resistance states specifically includes: Using a statistical sensing training method, weight data for a binary, deeply separable neural network robust to non-ideal hardware is obtained. The trained weight matrix is ​​mapped onto a magnetic tunnel junction array. A single magnetic tunnel junction is used to form a single weight. Different weights are mapped to different resistance state configurations of the magnetic tunnel junction. Based on the structure of the neural network and the number of weights, a magnetic tunnel junction array of corresponding size is selected. The magnetic tunnel junction is written into the corresponding resistance state through spin-transfer torque or spin-orbit torque. Wherein, the magnetic tunnel junction is in a parallel state and mapped to a weight "1"; the magnetic tunnel junction is in an antiparallel state and mapped to a weight "0".

3. The method of claim 1, wherein, Extracting Mel-frequency cepstral coefficient features from speech keyword signals and mapping them to read voltage amplitudes by inputting them into a magnetic tunneling junction array as voltage amplitudes specifically includes: MFCC feature extraction is performed on the input speech keyword signal. The window length, window step size, and number of MFCC feature channels in the MFCC processing are selected according to the network structure of the neural network and actual needs, so as to adjust the size of the final MFCC two-dimensional feature matrix input to the neural network.

4. The method of claim 1, wherein, When performing speech keyword recognition, the binarized speech MFCC features are continuously input into the magnetic tunnel junction array in the form of voltage amplitude. The multiplication and accumulation calculation results are obtained by measuring the output current of the target column of the magnetic tunnel junction array. After normalization and layer-by-layer inference, the final hardware keyword recognition results are output, including: The speech MFCC feature matrix P is mapped to the read voltage matrix V, where each column of the read voltage matrix V is sequentially input into the column of the corresponding convolution kernel of the magnetic tunnel junction array; Based on Ohm's law and Kirchhoff's law, the convolution operation is equivalent to the read operation of the magnetic tunnel junction array. The output current of each column of the magnetic tunnel junction array corresponding to the convolution result is summed to calculate the output current of each column of the magnetic tunnel junction under the corresponding read voltage combination, and an intermediate feature output matrix is ​​generated. The intermediate feature output matrix is ​​binarized to form a new input matrix P1, which is mapped to the read voltage matrix V1. The read voltage input and current accumulation reading operations of the subsequent columns of the magnetic tunnel junction array are performed again to obtain intermediate results P2 and V2. This process continues until the inference of each layer of the network is completed, and finally the speech keyword recognition result is output.

5. The method of claim 1, wherein, The magnetic tunnel junction array is at least one of the following: passive STT-MRAM array, active STT-MRAM array, active SOT-MRAM array, and three-dimensional stacked SOT-MRAM array.

6. A keyword speech recognition device based on an array of magnetic tunnel junctions, characterized by, include: The weight mapping module is used to train a neural network using a statistical sensing training method to obtain the weight data of the neural network, map the weight data to the high and low resistance state data of the magnetic tunnel junction array, select a magnetic tunnel junction array of appropriate size according to the actual network structure of the neural network, and pre-write the magnetic tunnel junction to the corresponding resistance state. The speech feature extraction module is used to extract Mel-frequency cepstral coefficients (MFCCs) features from speech keyword signals; The speech keyword recognition module is used to continuously input the binarized speech MFCC features in the form of voltage amplitude into the magnetic tunnel junction array when performing speech keyword recognition. The multiplication and accumulation calculation results are obtained by measuring the output current of the target column of the magnetic tunnel junction array. After normalization and layer-by-layer reasoning, the hardware keyword recognition results are finally output.

7. The apparatus according to claim 6, characterized in that, The weight mapping module is specifically used for: Using a statistical sensing training method, weight data for a binary, deeply separable neural network robust to non-ideal hardware is obtained. The trained weight matrix is ​​mapped onto a magnetic tunnel junction array. A single magnetic tunnel junction is used to form a single weight. Different weights are mapped to different resistance state configurations of the magnetic tunnel junction. Based on the structure of the neural network and the number of weights, a magnetic tunnel junction array of corresponding size is selected. The magnetic tunnel junction is written into the corresponding resistance state through spin-transfer torque or spin-orbit torque. Specifically, the magnetic tunnel junction is in a parallel state and mapped to a weight of "1"; the magnetic tunnel junction is in an antiparallel state and mapped to a weight of "0". The speech feature extraction module is specifically used for: MFCC feature extraction is performed on the input speech keyword signal. The window length, window step size, and number of MFCC feature channels in the MFCC processing are selected according to the network structure of the neural network and actual needs, so as to adjust the size of the final MFCC two-dimensional feature matrix input to the neural network. The voice keyword recognition module is specifically used for: The speech MFCC feature matrix P is mapped to the read voltage matrix V, where each column of the read voltage matrix V is sequentially input into the column of the corresponding convolution kernel of the magnetic tunnel junction array; Based on Ohm's law and Kirchhoff's law, the convolution operation is equivalent to the read operation of the magnetic tunnel junction array. The output current of each column of the magnetic tunnel junction array corresponding to the convolution result is summed to calculate the output current of each column of the magnetic tunnel junction under the corresponding read voltage combination, and an intermediate feature output matrix is ​​generated. The intermediate feature output matrix is ​​binarized to form a new input matrix P1, which is mapped to the read voltage matrix V1. The read voltage input and current accumulation reading operations of the subsequent columns of the magnetic tunnel junction array are performed again to obtain intermediate results P2 and V2. This process continues until the inference of each layer of the network is completed, and finally the speech keyword recognition result is output.

8. The apparatus of claim 6, wherein, The magnetic tunnel junction array is at least one of the following: passive STT-MRAM array, active STT-MRAM array, active SOT-MRAM array, and three-dimensional stacked SOT-MRAM array.

9. An electronic device, comprising: include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the keyword speech recognition method based on a magnetic tunneling array as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an information transmission implementation program, which, when executed by a processor, implements the steps of the keyword speech recognition method based on a magnetic tunneling array as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Voice keyword identification method and apparatus based on deep neural network

    CN105679316A

  • Memristor-based neural network training method and memristor-based neural network training device

    CN110796241A

  • Magnetoresistive memory unit, preparation method, array circuit and binary neural network chip

    CN116403623A

  • Imaging Lens System

    KR1020220101057A

  • Electric power pack for electric sprayer

    KR102514075B1