Neuromorphic device for realizing a neural network and its operating method
The neuromorphic device addresses cloud-based speech recognition limitations by implementing a neural network on-chip for real-time, secure, and efficient speech recognition without cloud dependency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-10-12
- Publication Date
- 2026-03-30
AI Technical Summary
Current speech recognition systems rely on cloud environments, leading to limitations in real-time performance and instability due to communication issues, making them unreliable for accurate and immediate speech recognition.
A neuromorphic device with an on-chip memory and processor that performs speech recognition using a neural network with layers for Fourier transform, Mel generation, and output, enabling local processing without cloud dependency.
The solution provides low dependence on servers, reduced communication costs, improved security, and enhanced speed by preventing data transmission bottlenecks, ensuring reliable speech recognition.
Smart Images

Figure 0007837086000001 
Figure 0007837086000002 
Figure 0007837086000003
Abstract
Description
Technical Field
[0001] The present invention relates to a neuromorphic device for realizing a neural network and an operation method thereof, and more particularly, to a system for performing speech recognition using an edge AI chip without using a cloud server or a physical server and an operation method thereof.
Background Art
[0002] With the development of Internet technology, the interaction between humans and computer devices has become more frequent. In the process of interaction between a computer device by human speech, it has become very important to accurately recognize natural language and determine the user's intention.
[0003] However, current speech recognition is performed in a cloud environment. In the case of a cloud-based speech recognition system, the input speech must be transmitted to the cloud, so there are limitations in real-time speech recognition. In particular, there is a problem that when the communication connection is unstable at the time of speech recognition or a problem occurs in the communication connection, speech recognition itself becomes impossible.
[0004] The above-mentioned background art is technical information that the inventor possessed for deriving the present invention or acquired in the process of deriving the present invention, and it is not necessarily known art that was publicly disclosed to the general public before the filing of the present invention.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The purpose of this disclosure is to provide neuromorphic devices for realizing neural networks and methods for operating them. The problems that this disclosure seeks to solve are not limited to the technical problems described above, and other technical problems not described will be clearly understood by those ordinary skill in the art from the description of the invention and will be further made clear from embodiments of this disclosure. It will also be understood that the problems and benefits that this disclosure seeks to solve can be achieved by the means and combinations thereof shown in the claims. [Means for solving the problem]
[0006] As a means to solve the aforementioned technical problems, a first aspect of the present disclosure provides a neuromorphic device comprising: a memory storing at least one program; an on-chip memory including a crossbar array circuit; and at least one processor that drives the neural network by executing the at least one program, wherein the at least one processor receives an audio signal, inputs the audio signal to a neural network trained on predetermined training data, and outputs a speech recognition result, the neural network comprising: a Fourier transform layer that converts the audio signal from a time band to a frequency band to generate a frequency signal; a Mel generation layer that generates a Mel spectrogram from the frequency signal; and an output layer that outputs the speech recognition result based on the speech features of the Mel spectrogram.
[0007] A second aspect of this disclosure provides a method for operating a neuromorphic device that implements a neural network, comprising the steps of receiving an audio signal and inputting the audio signal to a neural network trained on predetermined training data to output a speech recognition result, wherein the neural network includes a Fourier transform layer that converts the audio signal from a time band to a frequency band to generate a frequency signal, a Mel generation layer that generates a Mel spectrogram from the frequency signal, and an output layer that outputs the speech recognition result based on the speech features of the Mel spectrogram.
[0008] A third aspect of this disclosure can provide a computer-readable recording medium that stores a program for performing the method of the second aspect on a computer.
[0009] In addition to those, other methods, other apparatuses for realizing the present invention, and computer-readable recording media containing programs for performing the said methods can be further provided.
[0010] Other aspects, features, and advantages not described herein will become apparent from the attached drawings, claims, and the following detailed description of the invention. [Effects of the Invention]
[0011] According to the aforementioned problem-solving method of this disclosure, it is possible to provide a neuromorphic device that has low dependence on servers, reduced communication costs, and improved security for personal information.
[0012] Furthermore, the problem-solving means of this disclosure makes it possible to provide a neuromorphic device with improved speed by preventing bottleneck phenomena caused by data transmission and reception.
[0013] The effects of the embodiments are not limited to those described above, and any other effects not described will be clearly understood by those with ordinary skill in the art from the description of the present invention. [Brief explanation of the drawing]
[0014] [Figure 1] This is a diagram illustrating a neuromorphic chip structure according to one embodiment. [Figure 2] This is a diagram illustrating the architecture of a fully-connected neural network (FCNN) according to one embodiment. [Figure 3] This block diagram shows the hardware configuration of a neuromorphic device according to one embodiment. [Figure 4A] This is an illustrative diagram comparing the Von Neumann structure and the PIM (Processing-In Memory) structure. [Figure 4B] This is an illustrative diagram comparing the Von Neumann structure and the PIM (Processing-In Memory) structure. [Figure 5A] This is a diagram illustrating the operation method of a neural network according to one embodiment. [Figure 5B] This is a diagram illustrating the operation method of a neural network according to one embodiment. [Figure 6A] This figure compares vector-matrix multiplication in one embodiment with operations performed in a neural network. [Figure 6B] This figure compares vector-matrix multiplication in one embodiment with operations performed in a neural network. [Figure 7] This figure illustrates an example of a convolution operation being performed in a neural network according to one embodiment. [Figure 8] This is an illustrative diagram showing the implementation of a neural network according to one embodiment. [Figure 9]An exemplary diagram showing the realization of a crossbar array circuit according to an embodiment. [Figure 10] A flowchart of a method for operating a neuromorphic device according to an embodiment. [Figure 11] A block diagram of a neuromorphic device according to an embodiment of the present invention.
Best Mode for Carrying Out the Invention
[0015] A neuromorphic device for realizing a neural network, including a memory storing at least one program, an on-chip memory including a crossbar array circuit, and at least one processor for driving the neural network by executing the at least one program. The at least one processor receives an audio signal, inputs the audio signal to the neural network learned based on predetermined learning data, and outputs a speech recognition result. The neural network may include a Fourier transform layer for converting the audio signal from a time domain to a frequency domain to generate a frequency signal, a mel generation layer for generating a mel spectrogram from the frequency signal, and an output layer for outputting the speech recognition result based on the speech features of the mel spectrogram.
Mode for Carrying Out the Invention
[0016] In describing the present invention, when it is determined that a specific description of related known technologies obscures the gist of the present invention, the detailed description thereof may be omitted. Unless otherwise specifically defined, all terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention pertains.
[0017] Phrases such as "according to an embodiment", "relating to an embodiment", and "by the realization of an embodiment" in this specification do not necessarily all refer to the same embodiment.
[0018] Since embodiments can be modified in various ways and may take many forms, some embodiments are illustrated and described in detail in the drawings. However, this should not be understood as limiting the embodiments to any particular form of disclosure, but rather as including all modifications, equivalents, or substitutes that fall within the concept and technical scope of the embodiments. The terms used in this specification are for illustrative purposes only and are not intended to limit the embodiments.
[0019] The terminology used in the embodiments has been selected as widely used and general terms as possible, taking into account the functions of these embodiments. However, this may vary depending on the intentions of engineers in the technical field to which the embodiments belong, precedents, and the emergence of new technologies. In certain cases, the applicant may have arbitrarily selected some terms, in which case their meaning will be described in detail in the relevant section. Therefore, the terminology used in the embodiments should not be merely names of terms, but should be defined based on the meaning of those terms and the overall content of the embodiments.
[0020] Some embodiments of this disclosure can be represented by functional block configurations and various processing steps. Some or all of such functional blocks can be implemented by various numbers of hardware and / or software configurations that perform a particular function. For example, a functional block of this disclosure can be implemented by one or more microprocessors or by a circuit configuration for a given function.
[0021] Furthermore, for example, the functional blocks of this disclosure can be implemented in various programming or scripting languages. Functional blocks can also be implemented in algorithms that run on one or more processors. In addition, this disclosure may employ prior art for electronic environment configuration, signal processing, and / or data processing.
[0022] Terms such as "database," "element," "means," and "configuration" can be used broadly and are not limited to mechanical and physical configurations. Furthermore, terms such as "-part" and "-module" as described in the specification refer to a unit that processes at least one function or operation, which can be implemented in hardware or software, or a combination of hardware and software.
[0023] The connecting lines or members shown in the drawings between components are merely illustrative examples of functional and / or physical or circuit connections. In actual devices, connections between components can be represented by various alternative or additional functional, physical, or circuit connections.
[0024] Furthermore, while ordinal terms such as "first," "second," etc., used herein may be used to describe various components, the components themselves should not be limited by these terms. The terms are used solely for the purpose of distinguishing one component from another.
[0025] Furthermore, some components in the drawings may be shown with slightly exaggerated sizes or proportions. Additionally, components shown in one drawing may not be shown in other drawings.
[0026] Throughout the specification, “Embodiments” are arbitrary classifications that facilitate the description of the invention in this disclosure, and each embodiment does not need to be mutually exclusive. For example, a configuration disclosed in one embodiment can be applied to and / or implemented in other embodiments, and can be applied to and / or implemented with modifications without departing from the scope of this disclosure.
[0027] Furthermore, the terms used in this disclosure are for illustrative purposes only and do not limit the embodiments. In this disclosure, singular forms include plural forms unless otherwise specified.
[0028] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings, so as to be easily implemented by a person with ordinary skill in the art. However, embodiments of the present disclosure can be realized in a variety of different forms and are not limited to the embodiments described herein.
[0029] The present invention will now be described in detail based on the above and with reference to the drawings.
[0030] Figure 1 is a diagram illustrating a neuromorphic chip structure according to one embodiment.
[0031] Referring to Figure 1, Neuromorphic Chip 1 is hardware that generates circuits mimicking the morphology of neurons to simulate human brain function. In other words, Neuromorphic Chip 1 refers to a computer chip that mimics the structure of the nervous system.
[0032] Because Neuromorphic Chip 1 consists only of the circuits necessary for neural network computation, it can achieve gains of several hundred times or more in terms of power, area, and speed. Unlike conventional computers, the human brain does not consume much power even when processing a large amount of data, so Neuromorphic Chip 1 mimics this way of operating the brain by configuring structures that connect neurons and synapses in parallel, and saving energy by disconnecting the connections when data is not being processed.
[0033] For example, in conventional computers with a von Neumann architecture, data is processed sequentially as it is input, making it excellent for executing precisely crafted programs. However, it suffers from problems such as limited power consumption and low efficiency in areas like pattern recognition and real-time recognition.
[0034] In contrast, the Neuromorphic Chip 1 uses analog operation, where various states change progressively, rather than digital data such as 0s and 1s. That is, the parallel-configured artificial neurons operate in an event-driven manner without clock operation. Therefore, it can efficiently process unstructured characters, speech, and images that conventional computers have difficulty intuitively recognizing. Specifically, by distributing neurons (nerve cells) and synapses (connecting lines) using silicon transistor circuits and memory elements, data can be processed in parallel.
[0035] In one embodiment, when input data such as images, videos, or audio is input to the neuromorphic chip 1 shown in Figure 1, the neuromorphic chip 1 performs calculations on the input data and outputs predetermined output data. The predetermined output data may include audio / image / video recognition results obtained by feature classification of the input data. For example, when an audio signal is input as input data, the predetermined output data may include speech recognition results obtained by binary classification of whether or not the input data contains a pre-set keyword.
[0036] On the other hand, the data input to the neuromorphic chip 1 is not limited to the aforementioned images, videos, or audio, but may also include data in various forms such as text.
[0037] Figure 2 is a diagram illustrating the architecture of a fully-connected neural network (FCNN) according to one embodiment.
[0038] Referring to Figure 2, a fully connected layer, known as FCNN, refers to a convolutional neural network where all neurons in one layer are connected to all neurons in the next layer. This layer is used to classify data using a flattened matrix in the form of a one-dimensional array.
[0039] In FCNN, each node has a node value, and each neuron has a weight and a bias. When moving from one layer to another, the node value of the next layer is obtained by multiplying the weight of each node by the bias. If this calculated value satisfies certain conditions, the output value, after passing through an activation function, is input as the node value of the next layer and activated. If this calculated value does not satisfy the conditions, an activation function intervenes to deactivate the node after passing through an activation function.
[0040] Therefore, since the output value differs depending on the type of activation function, it is important to use the appropriate activation function as needed. Typical examples include the ReLU function and the Softmax function.
[0041] Since FCNN can only accept one-dimensional array data as input, a drawback is that if the input data is a three-dimensional image consisting of length, width, and channels (color), it must be flattened into one-dimensional data before being input to the FCNN. In other words, because the spatial information of the image is ignored, it has the disadvantage of not being able to extract features embedded in the shape. Therefore, FCNN is highly useful when the input data can be realized as one-dimensional array data, such as audio data. According to one embodiment, the neural network described later may be an FCNN.
[0042] Figure 3 is a block diagram showing the hardware configuration of a neuromorphic device according to one embodiment.
[0043] The neuromorphic device 300 can be implemented in various devices such as PCs (personal computers), server devices, mobile devices, and embedded devices. Specific examples include, but are not limited to, smartphones, tablet devices, AR (Augmented Reality) devices, IoT (Internet of Things) devices, autonomous vehicles, robotics, and medical devices that perform speech recognition, image recognition, and image classification using neural networks. Furthermore, the neuromorphic device 300 may also be a dedicated hardware accelerator (HW accelerator) installed in such devices. The neuromorphic device 300 may also be a dedicated module for driving neural networks, such as an NPU (neural processing unit), TPU (Tensor Processing Unit), or neural engine, but is not limited to these.
[0044] The neuromorphic device 300 may include a processor and memory. The neuromorphic device 300 shown in Figure 3 only shows the components according to this embodiment, and it will be obvious to those ordinary skill in the art that the neuromorphic device 300 may further include other general-purpose components in addition to those shown in Figure 3.
[0045] A neuromorphic device 300 according to one embodiment may include an input / output interface 310.
[0046] An input / output interface 310 according to one embodiment can transmit data from inside the neuromorphic device 300 to an external source and receive data from the external source back into the neuromorphic device 300. The signals of the input / output interface 310 may be unidirectional or bidirectional, single-ended or differential-mode, and may conform to one of other input / output interface standards.
[0047] An input / output interface 310 according to one embodiment may further include an audio receiver. For example, it may further include an audio ADC or DAC. The audio receiver can digitize the sound input from the microphone, and the audio receiver may also include an amplifier and support multiplex sampling rates.
[0048] In one embodiment, the input / output interface 310 can receive external audio signals from the neuromorphic device 300. The input / output interface 310 can also receive input, which is the output value of the neural network 330, as the speech recognition result, and transmit it to the arithmetic circuit 350 or another external unit (not shown).
[0049] An input / output interface 310 according to one embodiment may further include GPIO (General-Purpose Input / Output), I2S (Integrated Interchip Sound, Inter-IC Sound), I2C (Inter-Integrated Circuit), SPI (Serial Peripheral Interface), UART (Universal asynchronous receiver / transmitter), PWM (Pulse Width Modulation), and the like.
[0050] The neuromorphic device 300 according to one embodiment may include analog devices 320 and 340.
[0051] Analog devices 320 and 340 can provide stable power supply, system monitoring, and supervision functions for the edge AI chip.
[0052] Specifically, the analog devices 320 and 340 are devices that operate on parameters that appear in response to continuous physical quantities such as voltage, resistance, rotation, and pressure in monitoring, control, data acquisition, and automatic control. The analog devices 320 and 340 may include analog display devices and analog input / output devices, and may further include analog-to-digital converters and digital-to-analog converters.
[0053] Furthermore, analog devices 320 and 340 can provide power management and supervision functions for the neuromorphic device 300. Analog devices 320 and 340 further include a Low Voltage Detector (LVD), which can activate a Power-On Reset (POR) when the digital power supply (VDD) falls below a safe operating level to prevent unstable operation of the neuromorphic device 300.
[0054] A neuromorphic device 300 according to one embodiment may include an arithmetic circuit 350.
[0055] The arithmetic circuit 350 may include an MCU (Micro Controller Unit), DMA (Direct Memory Access), etc.
[0056] The MCU can play a control and management role in enabling efficient algorithm execution in a PIM structure. Because the PIM structure offers high efficiency in MAC (Multiply and Accumulate) operations, it may be suitable for computation of neural networks according to one embodiment of the present invention. The PIM structure will be described later based on Figure 4b.
[0057] DMA enables high-speed data transmission between peripherals and memory, or between memory modules. Data can be moved quickly via DMA without the CPU, allowing the CPU to be used for other tasks. DMA resides within SRAM and can contribute to accelerating data movement in neural networks.
[0058] One embodiment of the neuromorphic device 300 can be realized with an edge AI chip. Edge AI refers to the technology of executing AI algorithms on hardware devices using edge computing, which is based on data generated by the system. AI processing is mainly performed in cloud-based data centers that require enormous computing capacity, so it is highly dependent on servers. In contrast, by using edge AI, the execution of AI algorithm calculations is performed locally, so the dependence on the cloud (server) is reduced, communication costs are reduced, and privacy is protected because sensitive personal information is not sent to the cloud. Therefore, by configuring the neuromorphic device 300 with an edge AI chip, it is possible to reduce costs and improve security, as well as realize a highly responsive system because the calculations are processed immediately within the same hardware.
[0059] Figures 4a and 4b are illustrative diagrams comparing the Von Neumann structure and the PIM (Processing-In Memory) structure.
[0060] Referring to Figure 4a, the von Neumann structure is a computer structure proposed by John von Neumann, and is a typical three-stage structure of a stored-program computer consisting of main memory, a central processing unit, and an input / output device.
[0061] The von Neumann architecture has the advantage of greatly improving versatility in computing devices because, when changing tasks, hardware (such as wires) does not need to be rearranged, and only the software (program) needs to be changed. However, because it involves sequentially executing a sequence of instructions, each instruction consisting of operations that change values in designated memory locations, it causes serious problems in the design of high-speed computers. This is known as the von Neumann bottleneck phenomenon.
[0062] To solve the von Neumann bottleneck phenomenon, alternatives have been proposed, such as the Harvard architecture, which divides memory into parts where instructions are stored and parts where data is stored; the PIM architecture, which performs not only data storage but also data computation in memory; and neuromorphic computing, which uses an integrated circuit in the form of an artificial neural network that mimics the brain structure of higher animals, with many units that combine computation and memory functions connected in parallel in a network, and then operates each unit in an event-driven manner.
[0063] Referring to Figure 4b, it can be seen that the PIM structure consists of a processor and memory with computing capabilities.
[0064] Unlike conventional von Neumann architectures, where all data in memory is moved to the processor for calculations, the PIM architecture performs calculations in memory upon receiving a processor instruction and sends only the result data to the processor. This eliminates the need to move large amounts of data, effectively resolving the aforementioned von Neumann bottleneck phenomenon. It also has the advantage of significantly reducing power consumption.
[0065] Returning to Figure 3, the neuromorphic device 300 according to one embodiment may include a neural network 330.
[0066] In one embodiment, the neural network 330 can output output data by extracting features related to input data using multiple layers. For example, the neural network 330 can receive an audio signal as input and output a speech recognition result. As described above, the neural network 330 can filter or classify the features of the input audio signal to determine the type of audio signal, such as whether or not the audio signal contains a specific keyword, and output the determination result as a speech recognition result.
[0067] Here, the neural network 330 may be composed of orthogonal matrices, with inputs in the row direction and outputs generated in the column direction. However, the configuration of the neural network 330 is not limited to this.
[0068] In one embodiment, the neural network 330 may be a model trained on predetermined training data. As mentioned above, each layer constituting the neural network 330 may have weights and / or bias values. Since the number of nodes in one layer is equal to the size of the input tensor, there may be hundreds or thousands of nodes in each layer. When multiple such layers are stacked, the number of weights connected to the nodes of the next layer becomes infinitely large, making it difficult to set all the weights and / or bias values. Therefore, through learning, it is possible to detect the layer (i.e., weights and / or biases) that is most optimized for the desired classification for a huge amount of data.
[0069] For example, as a training method for neural network 330, deep learning can detect and set the most suitable weights in the neural network by combining three methods: loss function or cost function, optimization, and backpropagation.
[0070] The neural network 330 can perform calculations using only on-chip memory, without the need for external memory. For example, the neural network 330 can perform calculations on each layer using only on-chip memory in a PIM (Performance-Input-Motion) basis, without the need for external memory (such as off-chip memory), thereby enabling calculations to be performed without memory updates while processing audio signals. Specifically, the neural network 330 can perform PIM-based calculations with each memory cell directly connected to the processor.
[0071] However, while PIM requires high memory bandwidth, memory is sensitive to high temperatures, which can limit the power consumption of computing devices. Therefore, producing high-performance PIM chips may require new hardware structures, which can increase manufacturing costs. Thus, PIM structures may be advantageous for performing relatively small amounts of calculations or simple calculations. In contrast, the neural network 330 may be configured with multi-bit memory to overcome the shortcomings of such PIM structures. For example, the neural network 330 may be configured with memory capable of implementing 7 bits and 128 analog memory states. By configuring the neural network 330 with a large capacity, unlike typical PIM chips which exhibit problems such as heat generation and performance degradation, it can process vast amounts of data with low power consumption and high performance even during long-term use.
[0072] Figures 5a and 5b are diagrams illustrating the operation method of a neural network according to one embodiment.
[0073] Referring to Figure 5a, the neural network may include multiple cores, each of which may be implemented as a RCA (Resistive Crossbar Memory Array). Specifically, each core may include multiple presynaptic neurons 510, multiple postsynaptic neurons 520, and synapses 530 that provide the respective connections between the multiple presynaptic neurons 510 and the multiple postsynaptic neurons 520.
[0074] In one embodiment, the core of the neural network includes four presynaptic neurons 510, four postsynaptic neurons 520, and sixteen synapses 530, although their numbers can vary. If the number of presynaptic neurons 510 is N (where N is a natural number greater than or equal to 2) and the number of postsynaptic neurons 520 is M (where M is a natural number greater than or equal to 2, and may be the same as or different from N), then N*M synapses 530 may be arranged in a matrix.
[0075] Specifically, it is possible to provide wiring 512 that connects to each of multiple presynaptic neurons 510 and extends in a first direction (e.g., lateral direction), and wiring 522 that connects to each of multiple postsynaptic neurons 520 and extends in a second direction (e.g., longitudinal direction) that intersects with the first direction. For the sake of explanation, the wiring 512 extending in the first direction will be called a row line, and the wiring 522 extending in the second direction will be called a column line. Multiple synapses 530 are arranged at each intersection of the row lines 512 and the column lines 522, and the corresponding row lines 512 and the corresponding column lines 522 can be connected to each other.
[0076] The presynaptic neuron 510 can generate signals, such as signals corresponding to specific data, and transmit them to the row wiring 512, while the postsynaptic neuron 520 can receive and process synaptic signals via the synaptic element 530 through the column wiring 522. The presynaptic neuron 510 may correspond to an axon, and the postsynaptic neuron 520 may correspond to a neuron. However, whether a neuron is presynaptic or postsynaptic may be determined by its relative relationship to other neurons. For example, if the presynaptic neuron 510 receives synaptic signals in relation to other neurons, it can function as a postsynaptic neuron. Similarly, if the postsynaptic neuron 520 transmits signals in relation to other neurons, it can function as a presynaptic neuron. The presynaptic neuron 510 and the postsynaptic neuron 520 can be realized in various circuits such as CMOS.
[0077] The connection between the presynaptic neuron 510 and the postsynaptic neuron 520 can be established via a synapse 530. Here, the synapse 530 is an element whose electrical conductivity or weight changes in response to an electrical pulse, such as voltage or current, applied to both ends.
[0078] The core of the neural network may be configured using multi-level memory such as ReRAM (Resistive RAM) where synapses 530 are composed of variable resistors, FeRAM (Ferroelectric RAM) where synapses 530 are composed of ferroelectric material, PRAM (Phase-change RAM) where synapses 530 are composed of chalcogenide glass that changes from an amorphous state to a crystalline state when heat is applied, MRAM (Magnetic RAM) where synapses 530 are composed of magnetic elements, or NAND / NOR flash memory.
[0079] Synapse 530 may include, for example, a variable resistor element. The variable resistor element is an element that can switch between different resistance states depending on the voltage or current applied across its terminals, and may have a single-film structure or a multi-film structure containing various materials that can have multiple resistance states, such as metal oxides such as transition metal oxides and perovskite-based materials, phase change materials such as chalcogenide-based materials, ferroelectric materials, ferromagnetic materials, etc.
[0080] The core synapse 530 can be implemented with various characteristics that distinguish it from variable resistor elements in memory, such as exhibiting analog behavior where there is no abrupt change in resistance during set and reset operations, and conductivity changes gradually according to the number of electrical pulses input. This is because the characteristics required for variable resistor elements in memory and the characteristics required for synapse 530 in the core of a neural network are different.
[0081] Specifically, the neuromorphic chip can have a variable voltage applied to the core synapse 530 analog device. This allows the resistance or weight of the synapse 530 to change progressively.
[0082] The operation of the neural network described above can be explained as follows with reference to Figure 5b. For the sake of explanation, the rows 512 can be referred to as the first row 512A, second row 512B, third row 512C, and fourth row 512D from top to bottom, and the column rows 522 can be referred to as the first column row 522A, second column row 522B, third column row 522C, and fourth column row 522D from left to right.
[0083] Referring to Figure 5b, in the initial state, all synapses 530 may be in a state of relatively low conductivity, i.e., a high-resistance state. If at least some of the multiple synapses 530 are in a low-resistance state, further initialization operations may be required to bring them into a high-resistance state. Each of the multiple synapses 530 may have a predetermined threshold required for changes in resistance and / or conductivity. More specifically, if a voltage or current smaller than the predetermined threshold is applied across each synapse 530, the conductivity of the synapse 530 does not change, and if a voltage or current larger than the predetermined threshold is applied to the synapse 530, the conductivity of the synapse 530 can change.
[0084] In this state, in order to perform the operation of outputting specific data as the result of a specific column wiring 522, an input signal corresponding to the specific data can be entered into the row wiring 512 in accordance with the output of the presynaptic circuit 510. For example, the input signal may appear as the application of an electrical pulse to each of the row wirings 512. Alternatively, the column wirings 522 may be driven with an appropriate voltage or current for output.
[0085] As another example, it is not necessary to define a column wiring 522 that outputs specific data. In that case, by applying an electrical pulse corresponding to the specific data to the low wiring 512 and measuring the current flowing through each of the column wirings 522, the column wiring 522 that first reaches a predetermined threshold current, for example, the third column wiring 522C, may be the column wiring 522 that outputs that specific data.
[0086] The method described above allows different data to be output to different column wirings 522.
[0087] Figures 6a and 6b are diagrams illustrating a comparison between vector-matrix multiplication and operations performed in a neural network according to one embodiment.
[0088] Figures 6a and 6b are diagrams illustrating a comparison between vector-matrix multiplication and operations performed in a neural network according to one embodiment.
[0089] First, referring to Figure 6a, the convolution operation between the input data and the kernel may be performed using vector-matrix multiplication. For example, the pixel data of the input data can be represented by matrix X610, and the kernel value can be represented by matrix W611. The pixel data of the output data can be represented by matrix Y612, which is the result of the multiplication operation between matrix X610 and matrix W611.
[0090] Referring to Figure 6b, a vector multiplication operation may be performed using the core of the neural network. In comparison with Figure 6a, the pixel data of the input data is received as the input value of the core, and the input value may be a voltage of 620. Furthermore, the kernel value is stored in the core's synapses, i.e., memory cells, and the kernel value stored in the memory cells may be a conductance of 621. Therefore, the output value of the core can be represented by a current of 622, which is the result of the multiplication operation between the voltage of 620 and the conductance of 621.
[0091] Figure 7 illustrates an example of a convolution operation being performed in a neural network according to one embodiment.
[0092] The neural network can receive the audio signal 710, and the neural network's core 700 can be implemented using RCA (Resistive Crossbar Memory Arrays). Here, the audio signal 710 is converted from a time band to a frequency band by a Fourier transform layer implemented in the neural network's core 700, converted to a mel-spectrogram by a mel-generation layer implemented in the neural network's core 700, and a convolution operation can be performed on the pixel data corresponding to the mel-spectrogram. Therefore, in the following explanation of Figure 7, the pixel data of the audio signal 710 can mean either the converted frequency band pixel data or the generated mel-spectrogram pixel data.
[0093] In one embodiment, if the core 700 is a matrix of size N × M (where N and M are natural numbers greater than or equal to 2), the number of pixel data for the audio signal 710 may be less than or equal to the number of columns M in the core 700. The pixel data for the audio signal 710 may be parameters in floating-point format or fixed-point format.
[0094] The neural network can receive pixel data in digital signal form and convert the received pixel data into an analog signal form voltage using the DAC (Digital Analog Converter) 720. The pixel data of the audio signal 710 can have various bit resolution values, such as 1-bit, 4-bit, or 8-bit resolution. In one embodiment, the neural network can convert the pixel data into a voltage using the DAC 720 and then receive the voltage as the input value 701 of the core 700.
[0095] Furthermore, the neural network's core 700 may store learned kernel values. The kernel values may be stored in the core's memory cells, and the kernel values stored in the memory cells may be conductances 702. Here, the neural network can calculate an output value by performing a vector multiplication operation between an input value 701, for example, an input voltage, and conductance 702, and the output value can be represented by a current 703. In other words, the neural network can use the core 700 to output the same result value as the result of a convolution operation between an audio signal and a kernel.
[0096] Since the output value 703 output from core 700, for example, the output current, is an analog signal, the neural network can use an ADC (Analog Digital Converter) 730 to use the output current 703 as input data for other cores. The neural network can use the ADC 730 to convert the analog signal output current 703 into a digital signal. In one embodiment, the neural network can use the ADC 730 to convert the output current 703 into a digital signal having the same bit resolution as the pixel data of the audio signal 710. For example, if the pixel data of the audio signal 710 has a 1-bit resolution, the neural network can use the ADC 730 to convert the output current 703 into a 1-bit resolution digital signal.
[0097] The neural network can apply activation functions to the digital signals converted by the ADC730 using the activation unit 740. While sigmoid, tanh, and ReLU (Rectified Linear Unit) functions can be used as activation functions, the activation functions applicable to digital signals are not limited to these.
[0098] A digital signal to which an activation function has been applied can be used as an input value for another core 750. When a digital signal to which an activation function has been applied is used as an input value for another core 750, the process described above can be similarly applied to the other core 750.
[0099] On the other hand, core 700 and the other core 750 are not physically separated, but rather each core 700 and 750 has a variable resistance element value of the synapse changed according to the weight and / or bias value of each core 700 and 750.
[0100] Figure 8 is an example diagram illustrating the implementation of a neural network according to one embodiment.
[0101] Referring to Figure 8, the neural network may include a Fourier Transform layer 810 that converts the received audio signal from a time band to a frequency band to generate a frequency signal, a Mel generation layer 820 that generates a Mel spectrogram from the frequency signal, one or more hidden layers 830, 840 that classify the features of the audio signal based on the Mel spectrogram, and an output layer 850 that outputs the speech recognition result. The order of the layers may be the same as, but is not limited to, that shown in Figure 8.
[0102] In one embodiment, the Fourier transform layer 810 can perform a Fourier transform on a time-band audio signal to generate a frequency-band frequency signal. As mentioned above, an audio signal can have a time band because it can be data obtained by digitizing the sound input from a microphone to an input / output interface (or an audio receiver included in the input / output interface). Therefore, the Fourier transform layer 810 can convert the audio signal into frequency-band data by Fourier operations so that it is suitable for input to other layers 820-850 for classification of audio signals. For example, the Fourier transform layer 810 can generate a frequency signal by performing a Discrete Fourier Transform (DFT), which is a Fourier transform on a discontinuous discrete function.
[0103] In one embodiment, the neural network further includes a digital-to-analog converter (DAC) that converts an audio signal corresponding to a digital signal into an analog signal, and the received audio signal may be converted into an analog signal by the digital-to-analog converter and then input to the Fourier transform layer 810.
[0104] In one embodiment, the Mel generation layer 820 can receive a frequency signal input from the Fourier transform layer 810 and generate a Mel spectrogram corresponding to the audio signal.
[0105] A spectrogram is a graphical representation of the spectrum of an audio signal. The x-axis of a spectrogram represents time, and the y-axis represents frequency. The value of the frequency at each time point can be represented by color, depending on its magnitude. Here, since the result of applying a Fourier transform to an audio signal is a complex value, it is possible to remove the phase information by taking the absolute value of the complex value and generate a spectrogram that contains only the magnitude information.
[0106] On the other hand, a Mel spectrogram is a spectrogram whose frequency intervals have been readjusted using the Mel Scale. The human auditory organ is more sensitive in the low frequency range than in the high frequency range, and the Mel Scale reflects this characteristic by showing the relationship between physical frequencies and the frequencies that humans actually perceive. A Mel spectrogram can be generated by applying a filter bank based on the Mel Scale to a spectrogram.
[0107] In one embodiment, one or more hidden layers 830, 840 can receive input of a Mel spectrogram corresponding to an audio signal from a Mel generation layer 820 and classify the audio features of the audio signal or Mel spectrogram.
[0108] One or more hidden layers 830, 840 can classify a Mel spectrogram according to a predetermined criterion by treating the tensor corresponding to the input Mel spectrogram as a presynaptic neuron and applying weights to the presynaptic neuron. Therefore, the more hidden layers there are, the more sophisticated the processing of the input data can be, improving the functionality of the algorithm.
[0109] In one embodiment, one or more hidden layers 830, 840 can be located between the Mel generation layer 820 and the output layer 850. Thus, one or more hidden layers 830, 840 can receive a Mel spectrogram input from the Mel generation layer 820, classify the features, output the classification results, and transmit them to the output layer 850.
[0110] In one embodiment, the output layer 850 can receive input of speech features or the results of classifying speech features from the Mel generation layer 820 or one or more hidden layers 830, 840 and output a speech recognition result. For example, the output layer 850 can output a speech recognition result by determining whether the results of the classification of speech features by one or more hidden layers 830, 840 correspond to one of a set of pre-set keywords. Specifically, the neuromorphic device can store a set of pre-set keywords. Here, the hidden layers 830, 840 classify the features of the audio signal, the output layer 850 receives the classification results and can determine whether they correspond to one of a set of pre-set keywords.
[0111] For example, the multiple keywords may be command words consisting of a predetermined number of syllables, such as "Turn on the lights," "Turn off the lights," "Brighten the lights," "Dim the lights," "Block blue light," and "Warm lighting." Thus, the input / output interface receives an audio signal including ambient noise and natural language, and the Fourier transform layer 810, the Mel generation layer 820, and one or more hidden layers 830, 840 classify the features of the audio signal, allowing the output layer 850 to determine whether the classification result corresponds to one of the pre-set multiple keywords. That is, if the neuromorphic device finds that the received audio signal corresponds to one of the pre-set multiple keywords, it outputs that keyword as the speech recognition result via the output layer 850. If the received audio signal does not correspond to one of the pre-set multiple keywords, it outputs a null value or receives the next audio signal without outputting anything and repeats the calculations by each layer described above. In one embodiment, the multiple keywords may be set to around 20 or less.
[0112] In one embodiment, if the output layer 850 detects that the audio signal is a negative signal similar to at least one of a set of predefined keywords, it may output a null value as the speech recognition result, or it may receive the next audio signal without outputting anything and repeat the calculations performed by each of the layers described above. A negative signal can mean a keyword that is similar to any one of the set of predefined keywords and is highly likely to be recognized as a similar keyword by the neural network. For example, if the predefined keyword is "auto care", then "clock" may be a negative signal. Therefore, negative signals may be predefined to prevent the neural network from misrecognizing a negative signal similar to a keyword as a keyword.
[0113] In one embodiment, the neural network further includes an analog-to-digital converter (ADC) that converts speech recognition results corresponding to analog signals into digital signals, and the feature classification results or speech recognition results output from the output layer may be converted into digital signals by the analog-to-digital converter and then output via an input / output interface.
[0114] In one embodiment, if the neural network determines, based on the Mel spectrogram output from the Mel generation layer 820, that the audio signal does not correspond to any of the pre-set keywords, it may output a null value as the speech recognition result without inputting the Mel spectrogram to the hidden layers 830 and 840, or it may receive the next audio signal without outputting and repeat the calculations by each of the layers described above. In other words, when the neural network receives an audio signal, it performs calculations on the data using default values up to the Fourier transform layer 810 and the Mel generation layer 820, and if it determines, based on the Mel spectrogram, that the audio signal is not human speech or does not correspond to any of the pre-set keywords, it may not send data to the hidden layers 830, 840 and / or the output layer 850. Since it is not necessary for the hidden layers 830, 840 and / or the output layer 850 to perform calculations on all audio signals, the amount of computation can be reduced, thereby increasing the computational efficiency of the neuromorphic device and, if it has a PIM structure, preventing excessive heat generation.
[0115] Specifically, the neural network can receive an audio signal and, if the received input signal is not included in the trained model data, it can interrupt the computation and switch to a low-power mode (sleep mode). In other words, if the input signal is merely noise, including everyday ambient noise, the hidden layers 830 and 840 will not perform calculations, thereby minimizing unnecessary power consumption. Conversely, if the received input signal is included in the trained model data, the neural network can perform classification and output speech recognition results using the next hidden layers 830 and 840 and the output layer 850, and then switch to a low-power mode (sleep mode).
[0116] On the other hand, each layer 810-850 of the neural network is not physically separated, but rather represents a core in which the value of the variable resistor element of the synapse is changed according to the weights of each layer. Referring again to Figure 3 for explanation, each layer 810-850 of the neural network may all be implemented in the neural network 330.
[0117] For example, when an audio signal is received via the input / output interface 310, the audio signal, converted to an analog signal by the digital-to-analog converter 320, is input to a neural network 330, where the values of the synaptic variable resistors are set according to the weights of the Fourier transform layer 810, thereby generating a frequency signal in which the time bandwidth of the audio signal is changed to a frequency bandwidth. Subsequently, the frequency signal may be converted to a digital signal by the analog-to-digital converter 340 and input to the arithmetic circuit 350. Alternatively, the frequency signal, converted to an analog signal by the digital-to-analog converter 320 from the output of the arithmetic circuit 350, is input to a neural network 330, where the values of the synaptic variable resistors are set according to the weights of the Mel generation layer 820, thereby generating a Mel spectrogram of the frequency signal. Subsequently, the Mel spectrogram may be converted to a digital signal by the analog-to-digital converter 340 and input to the arithmetic circuit 350. Similarly, the process of classifying the features of an audio signal based on a neural network 330 in which the values of the synaptic variable resistor elements are set in accordance with the weights of one or more hidden layers 830, 840 and output layer 850, and outputting the speech recognition result, is redundant with the process described above and will therefore not be explained.
[0118] In one embodiment, the neural network may be trained using predetermined training data. For example, the predetermined training data may be audio data generated from a combination of a set of pre-defined keywords and noise sets. In the field of speech recognition, training data used to train a neural network can be used to build a neural network with good recognition performance by repeatedly training with data that mainly consists of a mixture of pre-defined keywords and noise such as ambient noise. However, building a neural network with good keyword recognition performance using keyword audio in a highly noisy environment as training data requires an enormous amount of training data and training processes. In contrast, if the neural network is first trained using a set of pre-defined keyword audio as training data, and then the neural network is trained using a dataset generated from a combination of the aforementioned pre-defined keyword audio and noise data as training data, the amount of training data is smaller compared to the aforementioned training method, making it more efficient and reducing the cost of generating training data.
[0119] In one embodiment, one or more hidden layers 830, 840 may consist of weights adjusted based on predetermined training data. The calculations of one or more hidden layers 830, 840 are performed with the resistance values of the variable resistor elements in the synapses of the neural network acting as weights. Here, each weight may be adjusted based on the aforementioned predetermined training data, i.e., audio data generated from a combination of a set of keyword speech and noise sets.
[0120] Figure 9 is an illustrative diagram of a crossbar array circuit according to one embodiment.
[0121] Referring to Figure 9, the crossbar array circuit 900 that realizes the neural network may include a plurality of sub-circuits 910, 920, and 930. The sub-circuits 910, 920, and 930 are circuits consisting of a combination of multiple cores that constitute the crossbar array circuit 900, and the synaptic weights may be set so that each sub-circuit 910, 920, and 930 corresponds to each layer that constitutes the neural network.
[0122] In one embodiment, the crossbar array circuit 900 may include a first sub-circuit 910, a second sub-circuit 920, and a third sub-circuit 930. Here, the first sub-circuit 910 can implement a Fourier transform layer, the second sub-circuit 920 can implement a Mel generation layer, and the third sub-circuit 930 can implement an output layer.
[0123] In one embodiment, weights corresponding to the Fourier transform layer may be stored in the synapses of the first subcircuit 910. For example, the first subcircuit 910 can generate a frequency signal by receiving a time-bandwidth audio signal as input data via a presynaptic neuron and transmitting a synaptic signal, which has passed through the synapse where the weights corresponding to the Fourier transform layer are stored, as output data via a postsynaptic neuron.
[0124] In one embodiment, weights corresponding to the Mel generation layer may be stored in the synapses of the second subcircuit 920. For example, the second subcircuit 920 can generate a Mel spectrogram by receiving a frequency signal as input data via a presynaptic neuron and transmitting a synaptic signal, which has passed through the synapse where the weights corresponding to the Mel generation layer are stored, as output data via a postsynaptic neuron.
[0125] In one embodiment, weights corresponding to the output layer may be stored in the synapses of the third subcircuit 930. For example, the third subcircuit 930 can output speech recognition results by receiving a Mel spectrogram or Mel spectrogram features as input data via a presynaptic neuron, and transmitting a synaptic signal via a postsynaptic neuron as output data through a synapse where weights corresponding to the output layer are stored.
[0126] Similarly, the crossbar array circuit 900 may further include other subcircuits (not shown) that implement hidden layers, in addition to the first to third subcircuits 910, 920, and 930. These other subcircuits (not shown) that implement hidden layers may be circuits in which weights corresponding to the hidden layers are stored in synapses. For example, a subcircuit (not shown) can classify the speech features of a Mel spectrogram by receiving a Mel spectrogram as input data via a presynaptic neuron and transmitting a synaptic signal via a postsynaptic neuron as output data through a synapse in which weights corresponding to the hidden layers are stored.
[0127] On the other hand, the reception of input data for each subcircuit 910, 920, and 930 can be performed not by presynaptic neurons, but via row or column wiring connected to the subcircuit that realizes the previous layer. Similarly, the transmission of output data for each subcircuit 910, 920, and 930 can be performed not by postsynaptic neurons, but via row or column wiring connected to the subcircuit that realizes the next layer. The method by which each subcircuit 910, 920, and 930 sends and receives input and output data is not limited to this. Also, although Figure 9 shows that the crossbar array circuit 900 contains three subcircuits 910, 920, and 930, the number of subcircuits is not limited to this, nor is the region, synaptic configuration, and location of subcircuits 910, 920, and 930 in the crossbar array circuit 900.
[0128] The crossbar array circuit 900 according to the aforementioned embodiment is included in on-chip memory and can realize a neural network through PIM-based computation. Such a neuromorphic device not only reduces the chip volume but also provides high performance relative to power consumption, allowing for very efficient output of speech recognition results for a relatively small number of keywords.
[0129] Figure 10 is a flowchart illustrating a method for operating a neuromorphic device according to one embodiment.
[0130] In step 1010, the neuromorphic device (hereinafter referred to as "device") can receive an audio signal.
[0131] In step 1020, the device can input an audio signal into a neural network trained based on predetermined training data and output a speech recognition result.
[0132] In one embodiment, the neural network may include a Fourier transform layer that converts an audio signal from a time band to a frequency band to generate a frequency signal, a Mel generation layer that generates a Mel spectrogram from the frequency signal, and an output layer that outputs a speech recognition result based on the speech features of the Mel spectrogram.
[0133] In one embodiment, the neural network may further include one or more hidden layers that classify the audio features of the Mel spectrogram, where the one or more hidden layers may be located between the Mel generation layer and the output layer.
[0134] In one embodiment, the output layer can output a speech recognition result by determining whether the audio signal corresponds to one of a set of predefined keywords, based on the classification of speech features by the hidden layer.
[0135] In one embodiment, the predetermined training data is audio data generated from a combination of a set of keywords and noise sets, and one or more hidden layers may consist of weights adjusted based on the predetermined training data.
[0136] In one embodiment, the neural network can output a null value as the speech recognition result if the audio signal corresponds to a negative signal that is similar to at least one of a set of keywords.
[0137] In one embodiment, the device may have a memory cell (or on-chip memory) and a processor directly connected, and perform PIM (Processing-in-Memory) based calculations.
[0138] In one embodiment, the device can convert an audio signal corresponding to a digital signal into an analog signal and input it to a Fourier transform layer, and convert the speech recognition result, which corresponds to an analog signal output from the output layer, into a digital signal.
[0139] In one embodiment, the device may, in response to determining that the audio signal does not correspond to any of a set of keywords based on the Mel spectrogram, refrain from inputting the Mel spectrogram to the hidden layer.
[0140] Figure 11 is a block diagram of a neuromorphic device according to one embodiment of the present invention. In the following explanation, explanations that overlap with the explanations of Figures 1 to 10 above will be omitted.
[0141] The processor plays a role in controlling the overall functions for executing the neuromorphic device 1100. For example, the processor 1110 controls the neuromorphic device 1100 overall by executing a program stored in the memory 1120 within the neuromorphic device 1100. The processor 1110 can be implemented as a CPU (central processing unit), GPU (graphics processing unit), AP (application processor), etc., provided within the neuromorphic device 1100, but is not limited to these.
[0142] Memory 1120 is hardware that stores various data processed within the neuromorphic device 1100. For example, memory can store data processed by the neuromorphic device 1100 and data being processed. Memory 1120 can also store applications, drivers, etc., driven by the neuromorphic device 1100. Memory includes RAM (random access memory) such as DRAM (dynamic random access memory) and SRAM (static random access memory), ROM (read-only memory), EEPROM (electrically erasable programmable read-only memory), CD-ROM, Blu-ray or other optical disc storage, HDD (hard disk drive), SSD (solid state drive), or flash memory.
[0143] Unlike the neural network 1 shown in Figure 1, the actual neural network driven by the neuromorphic device 1100 can be implemented with a more complex architecture. As a result, the processor 1110 can perform operations with an extremely large number of operations (operation counts), ranging from hundreds of millions to tens of billions, and the frequency with which the processor 1110 accesses memory 1130 for these operations can increase dramatically.
[0144] Therefore, memory 1120 can be on-chip memory. In one embodiment, the neuromorphic device 1100 is equipped with memory 1120 only in the form of on-chip memory and can perform calculations without accessing external memory 1130. For example, memory 1120 may be SRAM implemented in the form of on-chip memory. In that case, unlike above, types of memory mainly used in external memory 1130, such as DRAM, ROM, HDD, SSD, etc., do not have to be used in memory 1120.
[0145] The processor 1110 can read and write neural network data, such as audio data (audio signals), from the memory 1120, and execute the neural network using the read and written data. When the neural network is executed, the processor 1110 can repeatedly perform calculations on the audio signal to generate output data. That is, it can receive the audio signal at a predetermined interval and repeatedly perform the calculations related to the layer described above at each predetermined interval.
[0146] In one embodiment, the processor 1110 is directly connected to the memory 1120 and can drive a neural network by performing PIM-based calculations. Here, the memory 1120 may include a crossbar array circuit that receives instructions from the processor 1110 and performs calculations within the memory.
[0147] In one embodiment, the neuromorphic device 1100 may be a server. The server can be implemented as a computer device or a group of computer devices that communicate over a network to provide commands, codes, files, content, services, etc. The server can receive data necessary for speech recognition and perform speech recognition based on the received data.
[0148] On the other hand, embodiments of the present invention can be realized in the form of a computer program that can be executed on a computer by various components, and such a computer program can be recorded on a computer-readable medium. Here, the medium includes magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical recording media such as floptical disks; and hardware devices specially configured to store and execute program instructions, such as ROMs, RAMs, and flash memory.
[0149] On the other hand, the computer program may be specifically designed and configured for the present invention, or it may be publicly known and available to those skilled in the field of computer software. Examples of computer programs include not only machine code generated by a compiler, but also high-level language code executed by a computer using an interpreter or the like.
[0150] According to one embodiment, the methods according to various embodiments of the present disclosure can be provided in a computer program product. The computer program product can be traded as a commodity between a seller and a buyer. The computer program product can be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or through an application store (e.g., Play Store). TMThe computer program product can be distributed online (e.g., by download or upload) via a network or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated on a storage medium readable by equipment such as the memory of a manufacturer's server, an application store server, or an intermediary server.
[0151] With respect to the steps constituting the method according to the present invention, unless otherwise stated, the steps may be performed in any order that suits them. The present invention is not necessarily limited to the order in which the steps are described above. All use of examples or exemplary terms in the present invention is solely for the purpose of illustrating the invention in detail, and the scope of the present invention is not limited by such examples or exemplary terms unless otherwise limited by the claims. Furthermore, those skilled in the art will understand that the present invention can be constructed within the scope of the claims or their equivalents with various modifications, combinations, and changes, depending on the design conditions and factors.
[0152] Therefore, the concept of the present invention should not be limited to the embodiments described above, and it can be said that not only the scope of the appended claims, but also all scopes equivalent to or equivalently modified from those claims, fall within the scope of the concept of the present invention.
Claims
1. In neuromorphic devices that realize neural networks, The system includes at least one processor that drives the neural network, It includes an on-chip memory that includes a crossbar array circuit that receives instructions from at least one processor and performs in-memory operations, The aforementioned at least one processor is Receive an audio signal, The audio signal is input to the neural network trained based on predetermined training data, and the speech recognition result is output. The aforementioned neural network is The system includes a Fourier transform layer that converts the audio signal from a time band to a frequency band to generate a frequency signal, a Mel generation layer that generates a Mel spectrogram from the frequency signal, and an output layer that outputs the speech recognition result based on the speech features of the Mel spectrogram. The aforementioned crossbar array circuit is The system includes a crossbar array circuit comprising at least a first subcircuit corresponding to the Fourier transform layer, a second subcircuit corresponding to the Mel generation layer, and a third subcircuit corresponding to the output layer, The aforementioned at least one processor is A neuromorphic device that converts the audio signal, which corresponds to a digital signal, into an analog signal and inputs it to the first sub-circuit corresponding to the Fourier transform layer, and converts the speech recognition result, which corresponds to the analog signal output from the third sub-circuit corresponding to the output layer, into a digital signal.
2. The aforementioned neural network is The mel spectrogram includes one or more hidden layers that classify the audio features, The one or more hidden layers are located between the MEL generation layer and the output layer, and the at least one processor is The neuromorphic device according to claim 1, wherein, in response to the determination that the audio signal does not correspond to any of a plurality of pre-set keywords based on the Mel spectrogram, the Mel spectrogram is not input to the hidden layer.
3. The aforementioned at least one processor is The neuromorphic device according to claim 2, which outputs a null value as the speech recognition result in response to the determination that the audio signal does not correspond to any of the preset keywords.
4. The aforementioned at least one processor is The neuromorphic device according to claim 2, which, in response to the determination that the audio signal does not correspond to any of the preset keywords, receives the next audio signal without outputting the speech recognition result.
5. The aforementioned at least one processor is The neuromorphic device according to claim 2, which switches to a low-power mode in response to the determination that the audio signal does not correspond to any of the preset keywords.
6. The predetermined training data is The neuromorphic device according to claim 1, which is audio data generated based on at least one of a set of pre-configured keywords and noise sets.
7. The aforementioned neural network is The neuromorphic device according to claim 6, wherein audio data of the aforementioned set of pre-set keywords is used as training data for primary learning, and thereafter, an audio dataset generated from a combination of the aforementioned set of pre-set keywords and the noise set is used as training data for secondary learning.
8. The aforementioned at least one processor is The neuromorphic device according to claim 1, wherein if the audio signal corresponds to a negative signal similar to at least one of a plurality of pre-set keywords, a null value is output as the speech recognition result, or the next audio signal is received without outputting a null value.
9. In a method for operating a neuromorphic device that implements a neural network, The steps include receiving an audio signal and The process includes the step of inputting the audio signal into a neural network trained based on predetermined training data and outputting a speech recognition result, The aforementioned neural network is The system includes a Fourier transform layer that converts the audio signal from a time band to a frequency band to generate a frequency signal, a Mel generation layer that generates a Mel spectrogram from the frequency signal, and an output layer that outputs the speech recognition result based on the speech features of the Mel spectrogram. The neuromorphic device is The system includes a crossbar array circuit comprising at least a first subcircuit corresponding to the Fourier transform layer, a second subcircuit corresponding to the Mel generation layer, and a third subcircuit corresponding to the output layer, The steps include converting the audio signal, which corresponds to a digital signal, into an analog signal and inputting it to the first sub-circuit corresponding to the Fourier transform layer, A method comprising the step of converting the speech recognition result, which corresponds to an analog signal output from the third subcircuit corresponding to the output layer, into a digital signal.
10. A computer-readable recording medium that stores a program for executing the method according to claim 9 on a computer.
Citation Information
Patent Citations
Nerve cell element, recognition using neural network, and its learning method
JP2000352994A
Hotword recognition speech synthesis
JP2020528566A
Information processing device, information processing method, recognition model and program
JP2021026130A
Neural network-implementing device and method of operating the same
JP2021193565A