Neuromorphic device for implementing neural network and operation method thereof
The neuromorphic device addresses the limitations of cloud-based voice recognition by implementing a neural network on-chip for secure, low-cost, and efficient speech recognition, converting audio signals to frequency domain and generating mel spectrograms for improved local processing.
Patent Information
- Application Number
- JP2025172146
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-08
- Filing Date
- 2025-10-10
- Publication Date
- 2026-02-03
AI Technical Summary
Current voice recognition systems rely on cloud-based solutions, which limit real-time performance and can fail due to unstable connections, and they do not effectively address the need for secure, low-cost, and efficient speech recognition without server dependency.
A neuromorphic device with an on-chip memory and processor that implements a neural network, including a Fourier transform layer, mel generation layer, and output layer, enabling on-device speech recognition by converting audio signals to frequency domain and generating mel spectrograms for accurate recognition.
The neuromorphic device reduces dependency on servers, lowers communication costs, enhances security, and improves recognition speed by performing calculations locally, thus overcoming data transmission bottlenecks.
Smart Images

Figure 2026016459000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a neuromorphic device that realizes a neural network and an operating method thereof, and more particularly to a system that performs speech recognition using an edge AI chip without using a cloud server or a physical server, and an operating method thereof. [Background technology]
[0002] With the development of Internet technology, interactions between humans and computers are becoming more frequent. In the process of human speech-based interactions with computers, it is becoming increasingly important to accurately recognize natural language and determine user intent.
[0003] However, current voice recognition is performed in a cloud environment, and in the case of a cloud-based voice recognition system, the input voice must be sent to the cloud, which limits real-time voice recognition. In particular, if the communication connection is unstable or a problem occurs at the time of voice recognition, the voice recognition itself may become impossible.
[0004] The above-mentioned background art is technical information that the inventor possessed in order to derive the present invention or that he acquired in the process of deriving the present invention, and is not necessarily publicly known art that was disclosed to the general public prior to the filing of the present invention. Summary of the Invention [Problem to be solved by the invention]
[0005] An object of the present disclosure is to provide a neuromorphic device that realizes a neural network and a method for operating the same. The problems to be solved by the present disclosure are not limited to the technical problems described above, and other technical problems not described will be clearly understood by those skilled in the art from the description of the present invention and will be further clearly understood by the embodiments of the present disclosure. It will also be understood that the problems and advantages to be solved by the present disclosure can be achieved by the means and combinations thereof set forth in the claims. [Means for solving the problem]
[0006] As a means for solving the above-described technical problems, a first aspect of the present disclosure can provide a neuromorphic device including: a memory storing at least one program; an on-chip memory including a crossbar array circuit; and at least one processor that drives the neural network by executing the at least one program, wherein the at least one processor receives an audio signal, inputs the audio signal to a neural network that has been trained based on predetermined training data, and outputs a speech recognition result, and the neural network includes: a Fourier transform layer that converts the audio signal from a time domain to a frequency domain to generate a frequency signal; a mel generation layer that generates a mel spectrogram from the frequency signal; and an output layer that outputs the speech recognition result based on speech features of the mel spectrogram.
[0007] A second aspect of the present disclosure can provide a method for operating a neuromorphic device that implements a neural network, the method including steps of receiving an audio signal, inputting the audio signal to a neural network trained based on predetermined training data, and outputting a speech recognition result, the neural network including a Fourier transform layer that converts the audio signal from a time band to a frequency band to generate a frequency signal, a mel generation layer that generates a mel spectrogram from the frequency signal, and an output layer that outputs the speech recognition result based on speech features of the mel spectrogram.
[0008] A third aspect of the present disclosure can provide a computer-readable recording medium having recorded thereon a program for causing a computer to execute the method of the second aspect.
[0009] In addition to these, other methods and devices for realizing the present invention, and computer-readable recording media having recorded thereon programs for executing the methods can also be provided.
[0010] Other aspects, features, and advantages beyond those described above will be apparent from the accompanying drawings, the claims, and the following detailed description of the invention. [Effects of the Invention]
[0011] According to the above-described means for solving the problems of the present disclosure, it is possible to provide a neuromorphic device that has low dependency on a server, reduces communication costs, and improves security for personal information.
[0012] Furthermore, according to the means for solving the problems of the present disclosure, it is possible to provide a neuromorphic device that prevents bottlenecks caused by data transmission and reception and improves speed.
[0013] The effects of the embodiments are not limited to those described above, and other effects not described will be clearly understood by those skilled in the art from the description of the present invention. [Brief explanation of the drawings]
[0014] [Figure 1] 1A and 1B are diagrams illustrating a neuromorphic chip structure according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating the architecture of a fully-connected neural network (FCNN) according to one embodiment. [Figure 3] FIG. 2 is a block diagram showing the hardware configuration of a neuromorphic device according to one embodiment. [Figure 4A] This is an illustrative diagram for comparing a Von Neumann structure and a PIM (Processing-In Memory) structure. [Figure 4B] This is an illustrative diagram for comparing a Von Neumann structure and a PIM (Processing-In Memory) structure. [Figure 5A] FIG. 2 is a diagram illustrating a method of operating a neural network according to an embodiment. [Figure 5B] FIG. 2 is a diagram illustrating a method of operating a neural network according to an embodiment. [Figure 6A] FIG. 10 is a diagram for comparing vector-matrix multiplication according to one embodiment with operations performed in a neural network. [Figure 6B] FIG. 10 is a diagram for comparing vector-matrix multiplication according to one embodiment with operations performed in a neural network. [Figure 7] FIG. 10 is a diagram illustrating an example in which a convolution operation is performed in a neural network according to an embodiment. [Figure 8] FIG. 1 illustrates an exemplary implementation of a neural network according to one embodiment. [Figure 9]1 is an exemplary diagram illustrating an implementation of a crossbar array circuit according to one embodiment. [Figure 10] 1 is a flowchart of a method of operating a neuromorphic device according to one embodiment. [Figure 11] FIG. 1 is a block diagram of a neuromorphic device according to one embodiment of the present invention. BEST MODE FOR CARRYING OUT THE INVENTION
[0015] A neuromorphic device that realizes a neural network may include: a memory storing at least one program; an on-chip memory including a crossbar array circuit; and at least one processor that drives the neural network by executing the at least one program, wherein the at least one processor receives an audio signal, inputs the audio signal to the neural network that has been trained based on predetermined training data, and outputs a speech recognition result, and the neural network may include: a Fourier transform layer that converts the audio signal from a time domain to a frequency domain to generate a frequency signal; a mel generation layer that generates a mel spectrogram from the frequency signal; and an output layer that outputs the speech recognition result based on speech features of the mel spectrogram. DETAILED DESCRIPTION OF THE INVENTION
[0016] In describing the present invention, if it is determined that a specific description of related publicly known technology would obscure the gist of the present invention, the detailed description may be omitted, and unless otherwise defined, all terms used in this specification have the same meaning as commonly understood by a person having ordinary skill in the art to which the present invention belongs.
[0017] The appearances of phrases such as "in one embodiment," "in accordance with one embodiment," and "by implementing one embodiment" in this specification do not necessarily all refer to the same embodiment.
[0018] Since various modifications can be made to the embodiments and various forms can be taken, some embodiments are illustrated in the drawings and described in detail. However, it should be understood that the embodiments are not limited to the specific disclosed forms, but include all modifications, equivalents, and alternatives within the spirit and technical scope of the embodiments. The terms used in the specification are used merely to describe the embodiments and are not intended to limit the embodiments.
[0019] The terms used in the embodiments are generally selected as widely used terms as possible, taking into consideration the functions of the embodiments. However, these may vary depending on the intentions of engineers in the technical field to which the embodiments pertain, legal precedents, the emergence of new technologies, etc. In addition, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the relevant section. Therefore, the terms used in the embodiments should be defined based on the meanings of the terms and the overall content of the embodiments, rather than simply the names of the terms.
[0020] Some embodiments of the present disclosure may be illustrated by functional block configurations and various processing steps. Some or all of such functional blocks may be implemented by any number of hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a given function.
[0021] Also, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages, or may be implemented as algorithms executed on one or more processors. Furthermore, the present disclosure may employ conventional techniques for electronic configuration, signal processing, and / or data processing, etc.
[0022] Terms such as "database," "element," "means," and "configuration" can be used broadly and are not limited to mechanical and physical configurations. Furthermore, terms such as "unit" and "module" used in the specification refer to a unit that processes at least one function or operation, and can be realized by hardware or software, or a combination of hardware and software.
[0023] Note that the connecting lines or connecting members between components shown in the drawings are merely exemplary of functional and / or physical connections or circuit connections, and in an actual device, connections between components may be represented by various alternative or additional functional, physical, or circuit connections.
[0024] Furthermore, terms including ordinal numbers such as "first" and "second" used in this specification may be used to describe various components, but the components should not be limited by the terms. The terms are used only to distinguish one component from another.
[0025] In addition, the size and proportions of some components in the drawings may be somewhat exaggerated. Furthermore, components shown in one drawing may not be shown in other drawings.
[0026] Throughout the specification, the term "embodiment" is an arbitrary category for easily describing the invention in the present disclosure, and the respective embodiments do not necessarily have to be mutually exclusive. For example, a configuration disclosed in one embodiment can be applied and / or realized in other embodiments, and can be applied and / or realized with modifications without departing from the scope of the present disclosure.
[0027] Furthermore, the terms used in this disclosure are for the purpose of describing the embodiments and are not intended to limit the embodiments. In this disclosure, the singular includes the plural unless otherwise specified.
[0028]
[0033] Hereinafter, with reference to the accompanying drawings, detailed descriptions will be given of embodiments of the present disclosure so that those skilled in the art can easily implement them. However, the embodiments of the present disclosure may be realized in various different forms and are not limited to the embodiments described in the present disclosure.
[0029] Based on this, the present invention will be described in detail below with reference to the drawings.
[0030] FIG. 1 is a diagram illustrating a neuromorphic chip structure according to an embodiment.
[0031] Referring to Figure 1, a neuromorphic chip 1 is hardware that generates circuits that mimic the morphology of neurons to mimic the functions of the human brain. In other words, the neuromorphic chip 1 refers to a computer chip that mimics the structure of the nervous system.
[0032] Neuromorphic Chip 1 is composed only of the circuits necessary for neural network operations, enabling gains of hundreds of times in terms of power, area, and speed. Unlike conventional computers, the human brain does not consume much power even when processing large amounts of data. Neuromorphic Chip 1 mimics this brain's operating method by configuring a parallel structure connecting neurons and synapses and disconnecting the connections when data is not being processed, thereby saving energy.
[0033] For example, conventional computers with a von Neumann architecture process input data sequentially, making them excellent for executing precisely written programs. However, they have limitations in power consumption and are inefficient in pattern recognition and real-time recognition.
[0034] In contrast, the neuromorphic chip 1 uses analog operations in which data is gradually transformed into various states, rather than digital data such as 0 or 1. In other words, the parallel-configured artificial neurons operate in an event-driven manner without a clock. This allows it to efficiently process irregular characters, sounds, and images that conventional computers found difficult to intuitively recognize. Specifically, data can be processed in parallel by distributing neurons, which are nerve cells, and synapses, which are connecting lines, using silicon transistor circuits and memory elements.
[0035] In one embodiment, when input data such as an image, video, or audio is input to the neuromorphic chip 1 of Fig. 1, the input data is subjected to an internal operation of the neuromorphic chip 1, and predetermined output data is output. The predetermined output data may include a result of audio / image / video recognition based on feature classification of the input data. For example, when an audio signal is input as input data, the predetermined output data may include a result of speech recognition based on binary classification of whether or not the input data contains a preset keyword.
[0036] On the other hand, the data input to the neuromorphic chip 1 is not limited to the above-mentioned images, videos, or audio, but may include data in various forms such as text.
[0037] FIG. 2 is a diagram illustrating the architecture of a fully-connected neural network (FCNN) according to one embodiment.
[0038] Referring to Figure 2, FCNN, a fully connected layer, refers to a convolutional neural network in which all neurons in one layer are connected to all neurons in the next layer. This layer is used to classify data using a matrix flattened in the form of a one-dimensional array.
[0039] In FCNN, each node has a node value, and each neuron has a weight and bias. When moving from one layer to another, the value obtained by multiplying each node by its weight and adding the bias becomes the node value of the next layer. Here, if the calculated value meets certain conditions, the output value after passing through the activation function is input as the node value of the next layer and activated, but if the calculated value does not meet certain conditions, an activation function intervenes to deactivate the node through the activation function.
[0040] Therefore, since the output value differs depending on the type of activation function, it is important to use an appropriate activation function as needed. Typical examples are the ReLU function and the Softmax function.
[0041] Since only one-dimensional array data can be input to FCNN, if the input data is a three-dimensional image array consisting of vertical, horizontal, and channel (color), it must be flattened into one-dimensional data before being input to FCNN. In other words, the spatial information of the image is ignored, which means that features contained in the shape cannot be extracted. Therefore, FCNN is highly useful when data that can be realized as one-dimensional array data, such as audio data, is used as input. According to one embodiment, the neural network described below may be an FCNN.
[0042] FIG. 3 is a block diagram showing the hardware configuration of a neuromorphic device according to an embodiment.
[0043] The neuromorphic device 300 can be realized by various devices such as a personal computer (PC), a server device, a mobile device, and an embedded device, and specific examples thereof include, but are not limited to, smartphones, tablet devices, augmented reality (AR) devices, Internet of Things (IoT) devices, autonomous vehicles, robotics, medical equipment, etc. that perform speech recognition, video recognition, video classification, etc. using a neural network. Furthermore, the neuromorphic device 300 may correspond to a dedicated hardware accelerator (HW accelerator) mounted on such devices, and the neuromorphic device 300 may be, but is not limited to, a hardware accelerator such as an NPU (neural processing unit), a TPU (tensor processing unit), or a neural engine, which is a dedicated module for driving a neural network.
[0044] The neuromorphic device 300 may include a processor and a memory. Only components according to this embodiment are shown in the neuromorphic device 300 shown in Fig. 3, and it will be obvious to a person skilled in the art that the neuromorphic device 300 may further include other general-purpose components in addition to the components shown in Fig. 3.
[0045] According to one embodiment, the neuromorphic device 300 may include an input / output interface 310 .
[0046] In one embodiment, the input / output interface 310 can transmit data from within the neuromorphic device 300 to an external source and receive data from an external source into the neuromorphic device 300. The signals of the input / output interface 310 can be unidirectional or bidirectional, single-ended or differential-mode, or conform to one of several other input / output interface standards.
[0047] According to an embodiment, the input / output interface 310 may further include an audio receiver, such as an audio ADC or DAC, which can digitize audio input from a microphone, may include an amplifier, and can support multiple sampling rates.
[0048] In one embodiment, the input / output interface 310 can receive an audio signal external to the neuromorphic device 300. The input / output interface 310 can also receive an input of a speech recognition result, which is an output value of the neural network 330, and transmit it to the arithmetic circuit 350 or another external unit (not shown).
[0049] According to one embodiment, the input / output interface 310 may further include a general-purpose input / output (GPIO), an integrated interchip sound (I2S), an inter-integrated circuit (I2C), a serial peripheral interface (SPI), a universal asynchronous receiver / transmitter (UART), a pulse width modulation (PWM), etc.
[0050] According to one embodiment, the neuromorphic device 300 may include analog devices 320, 340.
[0051] The analog devices 320 and 340 can provide stable power supply, system monitoring, and supervision functions in the edge AI chip.
[0052] Specifically, the analog devices 320, 340 are devices that operate on parameters that correspond to continuous physical quantities such as voltage, resistance, rotation, pressure, etc. in monitoring control, data collection, and automatic control. The analog devices 320, 340 may include analog display devices and analog input / output devices, and may further include analog-to-digital converters and digital-to-analog converters.
[0053] Additionally, the analog devices 320 and 340 can implement power management and supervision functions for the neuromorphic device 300. The analog devices 320 and 340 further include a low voltage detector (LVD) that can activate a power-on reset (POR) when the digital power supply (VDD) falls below a safe operating level to prevent unstable operation of the neuromorphic device 300.
[0054] According to one embodiment, the neuromorphic device 300 may include an operational circuit 350 .
[0055] The arithmetic circuit 350 may include an MCU (Micro Controller Unit), a DMA (Direct Memory Access), and the like.
[0056] The MCU can play a role in controlling and managing the PIM architecture so that efficient algorithms can be executed. The PIM architecture may be suitable for neural network calculations according to an embodiment of the present invention because of its high efficiency in MAC (Multiply and Accumulate) operations. The PIM architecture will be described later with reference to FIG. 4b.
[0057] DMA allows high-speed data transfer between peripherals and memory, or between memories. Data can be moved quickly with DMA without the CPU being involved, allowing the CPU to perform other tasks. DMA is located inside SRAM and can help accelerate data movement in neural networks.
[0058] The neuromorphic device 300 according to one embodiment may be implemented using an edge AI chip. Edge AI refers to a technology for executing AI algorithms on hardware devices using edge computing based on data generated by a system. AI processing is primarily performed in cloud-based data centers, which require enormous computing capacity and are therefore highly dependent on servers. In contrast, edge AI reduces dependency on the cloud (server) by executing AI algorithm operations locally, thereby reducing communication costs and protecting privacy by preventing sensitive personal information from being transmitted to the cloud. Therefore, configuring the neuromorphic device 300 using an edge AI chip not only reduces costs and improves security, but also realizes a highly responsive system because operations are immediately processed within the same hardware.
[0059] 4a and 4b are illustrative diagrams for comparing the Von Neumann structure and the PIM (Processing-In Memory) structure.
[0060] Referring to Figure 4a, the von Neumann architecture is a computer architecture proposed by John von Neumann, and is a stored-program computer architecture consisting of a typical three-tier structure of a main memory, a central processing unit, and input / output devices.
[0061] The von Neumann architecture has the advantage of greatly improving versatility, since changing a computing device to perform a different task does not require rearranging the hardware (wires, etc.), but rather requires changing only the software (programs). However, because it executes a series of instructions, each of which consists of changing the value of a specific memory location, it poses a serious problem in the design of high-speed computers, known as the von Neumann bottleneck.
[0062] To solve the von Neumann bottleneck phenomenon, alternatives that have been proposed include the Harvard architecture, which divides memory into a section that stores instructions and a section that stores data; the PIM architecture, which not only stores data but also performs data calculations in memory; and neuromorphic computing, which is an integrated circuit in the form of an artificial neural network that mimics the brain structure of higher animals, in which numerous units that combine calculation and memory functions are connected in parallel in a mesh, and each unit then operates in an event-driven manner.
[0063] Referring to Figure 4b, it can be seen that the PIM structure consists of a processor and a memory with computing capabilities.
[0064] Unlike the conventional von Neumann architecture, where all data in memory is transferred to the processor for calculation, the PIM architecture performs calculations in memory when a processor instruction is received and only the resulting data is sent to the processor, eliminating the need to transfer large amounts of data, effectively resolving the von Neumann bottleneck phenomenon mentioned above, and also offering the advantage of significantly reducing power consumption.
[0065] Returning again to FIG. 3, a neuromorphic device 300 according to one embodiment may include a neural network 330 .
[0066] In one embodiment, the neural network 330 can output output data by extracting features related to input data using multiple layers. For example, the neural network 330 can receive an audio signal as input and output a speech recognition result. As described above, the neural network 330 can filter or classify the features of the input audio signal to determine the type of the audio signal, such as whether the audio signal contains a specific keyword, and output the determination result as the speech recognition result.
[0067] Here, the neural network 330 may be configured as an orthogonal matrix, with inputs being row-wise and outputs being generated column-wise, but the configuration of the neural network 330 is not limited to this.
[0068] In one embodiment, the neural network 330 may be a model trained based on predetermined training data. As described above, each layer constituting the neural network 330 may have weights and / or bias values. Since the number of nodes in one layer is equal to the size of the input tensor, there may be hundreds or thousands of nodes in each layer. When multiple such layers are stacked, the number of weights connected to the nodes in the next layer becomes infinitely large, making it difficult to set all the weights and / or bias values. Therefore, through training, it is possible to find the layer (i.e., the weights and / or biases) that is most optimized for the desired classification for a huge amount of data.
[0069] For example, as a learning method for the neural network 330, deep learning can detect and set the most suitable weights in the neural network by combining three methods: loss function or cost function, optimization, and back propagation.
[0070] The neural network 330 can perform calculations using only on-chip memory without using external memory. For example, the neural network 330 can perform calculations for each layer on a PIM basis using only on-chip memory without using external memory (e.g., off-chip memory), thereby performing calculations without memory updates while processing an audio signal. Specifically, the neural network 330 can perform PIM-based calculations in which each memory cell is directly connected to a processor.
[0071] However, PIM requires high memory bandwidth, but memory is sensitive to high temperatures, which can limit the power consumption of a computing device. Therefore, producing a high-performance PIM-structured chip may require new hardware architectures, which can increase manufacturing costs. Therefore, a PIM structure may be advantageous for performing relatively small or simple calculations. To overcome the drawbacks of such PIM structures, neural network 330 may be configured with multi-bit memory. For example, neural network 330 may be configured with memory capable of realizing 7 bits and 128 analog memory states. By configuring neural network 330 with a large capacity, it can process large amounts of data with low power consumption and high performance, even over long periods of use, unlike typical PIM chips that exhibit issues such as heat generation and performance degradation.
[0072] 5a and 5b are diagrams illustrating a method of operation of a neural network according to one embodiment.
[0073] 5a, the neural network may include multiple cores, each of which may be implemented with resistive crossbar memory arrays (RCA). Specifically, each core may include multiple presynaptic neurons 510, multiple postsynaptic neurons 520, and synapses 530 providing respective connections between the multiple presynaptic neurons 510 and the multiple postsynaptic neurons 520.
[0074] In one embodiment, the core of the neural network includes four presynaptic neurons 510, four post-synaptic neurons 520, and 16 synapses 530, although these numbers can be varied. If the number of presynaptic neurons 510 is N (where N is a natural number greater than or equal to 2) and the number of post-synaptic neurons 520 is M (where M is a natural number greater than or equal to 2 and may be the same as or different from N), N*M synapses 530 may be arranged in a matrix.
[0075] Specifically, it is possible to provide wiring 512 connected to each of the plurality of presynaptic neurons 510 and extending in a first direction (e.g., horizontal direction), and wiring 522 connected to each of the plurality of postsynaptic neurons 520 and extending in a second direction (e.g., vertical direction) intersecting the first direction. Hereinafter, for convenience of explanation, the wiring 512 extending in the first direction will be referred to as a row line, and the wiring 522 extending in the second direction will be referred to as a column line. A plurality of synapses 530 are arranged at each intersection of the row wiring 512 and the column wiring 522, and can connect corresponding row wirings 512 and corresponding column wirings 522 to each other.
[0076] The presynaptic neuron 510 generates a signal, e.g., a signal corresponding to specific data, and transmits it to the row wiring 512. The post-synaptic neuron 520 receives and processes a synaptic signal from the synapse element 530 via the column wiring 522. The pre-synaptic neuron 510 may correspond to an axon, and the post-synaptic neuron 520 may correspond to a neuron. However, whether a neuron is a pre-synaptic neuron or a post-synaptic neuron may be determined based on its relative relationship with other neurons. For example, if the pre-synaptic neuron 510 receives a synaptic signal in relation to other neurons, it may function as a post-synaptic neuron. Similarly, if the post-synaptic neuron 520 transmits a signal in relation to other neurons, it may function as a pre-synaptic neuron. The pre-synaptic neuron 510 and the post-synaptic neuron 520 may be implemented using various circuits, such as CMOS.
[0077] The connection between the presynaptic neuron 510 and the postsynaptic neuron 520 can be made via a synapse 530. Here, the synapse 530 is an element whose electrical conductance or weight changes in response to an electrical pulse, such as a voltage or current, applied across it.
[0078] The core of the neural network may be configured using a multi-level memory such as ReRAM (Resistive RAM) in which the synapses 530 are configured with variable resistors, FeRAM (Ferroelectric RAM) in which the synapses 530 are configured with ferroelectrics, PRAM (Phase-change RAM) in which the synapses 530 are configured with chalcogenide glass that changes from an amorphous state to a crystalline state when heat is applied, MRAM (Magnetic RAM) in which the synapses 530 are configured with magnetic elements, or NAND / NOR flash memory.
[0079] The synapse 530 may include, for example, a variable resistance element, which is an element that can be switched between different resistance states depending on the voltage or current applied across it, and may have a single-layer structure or a multi-layer structure including various materials that can have multiple resistance states, such as transition metal oxides, metal oxides such as perovskite-based materials, phase-change materials such as chalcogenide-based materials, ferroelectric materials, and ferromagnetic materials.
[0080] The core synapse 530 can be realized to have various characteristics that are distinct from variable resistance elements in memory, such as exhibiting analog behavior in which the conductivity changes gradually depending on the number of input electrical pulses, without an abrupt change in resistance between the set and reset operations, because the characteristics required for variable resistance elements in memory differ from the characteristics required for synapses 530 in the core of a neural network.
[0081] Specifically, the neuromorphic chip can apply varying voltages to the core synapses 530 through analog devices, thereby gradually changing the resistance or weight of the synapses 530.
[0082] The operation of the above-described neural network will be described below with reference to Fig. 5b. For convenience of explanation, the row wires 512 may be referred to as a first row wire 512A, a second row wire 512B, a third row wire 512C, and a fourth row wire 512D from the top, and the column wires 522 may be referred to as a first column wire 522A, a second column wire 522B, a third column wire 522C, and a fourth column wire 522D from the left.
[0083] Referring to FIG. 5b, initially, all of the synapses 530 may be in a relatively low conductance state, i.e., a high resistance state. If at least some of the synapses 530 are in a low resistance state, an initialization operation may be required to set them to a high resistance state. Each of the synapses 530 may have a predetermined threshold required for a change in resistance and / or conductance. More specifically, when a voltage or current smaller than the predetermined threshold is applied across each synapse 530, the conductance of the synapse 530 does not change, whereas when a voltage or current larger than the predetermined threshold is applied across the synapse 530, the conductance of the synapse 530 may change.
[0084] In that state, an input signal corresponding to the particular data can be input to the row wires 512 corresponding to the output of the presynaptic circuit 510 to operate to output the particular data as a result on a particular column wire 522. For example, the input signal may be manifested by the application of an electrical pulse to each of the row wires 512, and the column wires 522 may be driven with an appropriate voltage or current for output.
[0085] As another example, the column wire 522 that outputs specific data may not be specified. In this case, by measuring the current flowing through each column wire 522 while applying an electrical pulse corresponding to the specific data to the row wires 512, the column wire 522 that first reaches a predetermined threshold current, for example, the third column wire 522C, may be the column wire 522 that outputs the specific data.
[0086] By using the method described above, different data can be output to different column wirings 522, respectively.
[0087] 6a and 6b are diagrams for comparing vector-matrix multiplication according to one embodiment with the operations performed in a neural network.
[0088] 6a and 6b are diagrams for comparing vector-matrix multiplication according to one embodiment with the operations performed in a neural network.
[0089] 6a, the convolution operation between the input data and the kernel may be performed using vector-matrix multiplication. For example, pixel data of the input data may be represented by matrix X 610, and kernel values may be represented by matrix W 611. Pixel data of the output data may be represented by matrix Y 612, which is the result of the multiplication operation between matrix X 610 and matrix W 611.
[0090] Referring to Figure 6b, a vector multiplication operation may be performed using a neural network core. Compared to Figure 6a, pixel data of input data may be received as input values of the core, and the input values may be voltages 620. Furthermore, kernel values may be stored in synapses, i.e., memory cells, of the core, and the kernel values stored in the memory cells may be conductances 621. Therefore, the output value of the core may be represented by currents 622, which are the result of the multiplication operation between voltages 620 and conductances 621.
[0091] FIG. 7 is a diagram for explaining an example in which a convolution operation is performed in a neural network according to an embodiment.
[0092] The neural network may receive an audio signal 710, and the neural network core 700 may be implemented using resistive crossbar memory arrays (RCA). The audio signal 710 may be converted from the time domain to the frequency domain by a Fourier transform layer implemented in the neural network core 700, converted into a mel-spectrogram by a mel generation layer implemented in the neural network core 700, and a convolution operation may be performed on pixel data corresponding to the mel-spectrogram. Therefore, in the following description of FIG. 7, pixel data of the audio signal 710 may refer to pixel data of the converted frequency domain or pixel data of the generated mel-spectrogram.
[0093] In one embodiment, when the core 700 is a matrix of size N×M (N and M are natural numbers greater than or equal to 2), the number of pixel data of the audio signal 710 may be less than or equal to the number of columns M of the core 700. The pixel data of the audio signal 710 may be parameters in floating-point format or fixed-point format.
[0094] The neural network may receive pixel data in the form of a digital signal and convert the received pixel data into a voltage in the form of an analog signal using a digital-to-analog converter (DAC) 720. The pixel data of the audio signal 710 may have various bit resolution values, such as 1-bit, 4-bit, or 8-bit resolution. In one embodiment, the neural network may convert the pixel data into a voltage using the DAC 720 and then receive the voltage as an input value 701 of the core 700.
[0095] Furthermore, a learned kernel value may be stored in the core 700 of the neural network. The kernel value may be stored in a memory cell of the core, and the kernel value stored in the memory cell may be a conductance 702. Here, the neural network can calculate an output value by performing a vector multiplication operation between an input value 701, for example, an input voltage, and a conductance 702, and the output value can be expressed as a current 703. In other words, the neural network can output a result value that is the same as the result of a convolution operation between an audio signal and a kernel using the core 700.
[0096] Because the output value 703, e.g., output current, output from the core 700 is an analog signal, the neural network can use an ADC (Analog Digital Converter) 730 to use the output current 703 as input data for another core. The neural network can use the ADC 730 to convert the analog output current 703 into a digital signal. In one embodiment, the neural network can use the ADC 730 to convert the output current 703 into a digital signal with the same bit resolution as the pixel data of the audio signal 710. For example, if the pixel data of the audio signal 710 has 1-bit resolution, the neural network can use the ADC 730 to convert the output current 703 into a digital signal with 1-bit resolution.
[0097] The neural network can use an activation unit 740 to apply an activation function to the digital signal converted by the ADC 730. The activation function can be a Sigmoid function, a Tanh function, or a ReLU (Rectified Linear Unit) function, but is not limited to these.
[0098] The digital signal to which the activation function has been applied can be used as an input value of another core 750. When the digital signal to which the activation function has been applied is used as an input value of another core 750, the above-described process can be similarly applied to the other core 750.
[0099] On the other hand, the core 700 and the other core 750 are not physically separated, but may refer to the respective cores 700, 750 in which the variable resistance element values of the synapses are changed according to the weight and / or bias values of each core 700, 750.
[0100] FIG. 8 is an example diagram of a neural network implementation according to one embodiment.
[0101] 8, the neural network may include a Fourier transform layer 810 that converts a received audio signal from a time domain to a frequency domain to generate a frequency signal, a mel generation layer 820 that generates a mel spectrogram from the frequency signal, one or more hidden layers 830 and 840 that classify features of the audio signal based on the mel spectrogram, and an output layer 850 that outputs a speech recognition result. The order of the layers may be the same as, but is not limited to, that shown in FIG. 8.
[0102] In one embodiment, the Fourier transform layer 810 may perform a Fourier transform on a time-domain audio signal to generate a frequency signal in a frequency domain. As described above, the audio signal may be data digitized from audio input from a microphone to an input / output interface (or an audio receiving unit included in the input / output interface), and thus may have a time domain. Therefore, the Fourier transform layer 810 may convert the audio signal into frequency-domain data by a Fourier operation so that the data is suitable for input to the other layers 820 to 850 for classification. For example, the Fourier transform layer 810 may generate a frequency signal by a DFT (Discrete Fourier Transform), which is a Fourier transform of a discontinuous discrete function.
[0103] In one embodiment, the neural network further includes a digital-to-analog converter (DAC) for converting a digital audio signal into an analog signal, and the received audio signal may be converted into an analog signal by the digital-to-analog converter and then input into the Fourier transform layer 810.
[0104] In one embodiment, the Mel generation layer 820 can receive frequency signal input from the Fourier transform layer 810 and generate a Mel spectrogram corresponding to the audio signal.
[0105] A spectrogram is a graphical representation of the spectrum of an audio signal. The x-axis of a spectrogram represents time, and the y-axis represents frequency. The frequency values per time are represented by colors according to the magnitude of the values. Since the result of performing a Fourier transform on an audio signal is a complex value, a spectrogram containing only magnitude information can be generated by taking the absolute value of the complex value, eliminating the phase information.
[0106] On the other hand, a Mel spectrogram is a spectrogram whose frequency intervals are rearranged to the Mel scale. The human hearing organ is more sensitive in the low frequency band than in the high frequency band, and the Mel scale reflects this characteristic and shows the relationship between physical frequencies and the frequencies that humans actually perceive. A Mel spectrogram can be generated by applying a Mel scale-based filter bank to a spectrogram.
[0107] In one embodiment, one or more hidden layers 830, 840 may receive an input of a mel spectrogram corresponding to an audio signal from the mel generation layer 820 and classify the audio features of the audio signal or the mel spectrogram.
[0108] One or more hidden layers 830, 840 use tensors corresponding to the input mel spectrogram as presynaptic neurons and reflect weights to the presynaptic neurons to classify the mel spectrogram according to a predetermined criterion. Therefore, the more hidden layers there are, the more precisely the input data can be processed, improving the algorithm's performance.
[0109] In one embodiment, one or more hidden layers 830, 840 may be located between the mel generation layer 820 and the output layer 850. Thus, the one or more hidden layers 830, 840 may receive mel spectrograms from the mel generation layer 820, classify the features, and output the classification results for transmission to the output layer 850.
[0110] In one embodiment, the output layer 850 may receive voice features or the results of classifying the voice features from the mel generation layer 820 or one or more hidden layers 830 and 840 and output a voice recognition result. For example, the output layer 850 may output a voice recognition result by determining whether the results of classifying the voice features from one or more hidden layers 830 and 840 correspond to one of a plurality of preset keywords. Specifically, the neuromorphic device may store a plurality of preset keywords. Here, the hidden layers 830 and 840 classify the features of the audio signal, and the output layer 850 may receive the classification results and determine whether the results correspond to one of a plurality of preset keywords.
[0111] For example, the keywords may be commands consisting of a predetermined number of syllables, such as "turn on the lights," "turn off the lights," "brighten the lights," "dim the lights," "block blue light," or "warm lights." The input / output interface receives an audio signal containing background noise and natural language, and classifies the audio signal's characteristics using the Fourier transform layer 810, the Mel generation layer 820, and one or more hidden layers 830 and 840. The output layer 850 can then determine whether the received audio signal corresponds to one of the preset keywords based on the classification results. That is, if the received audio signal corresponds to one of the preset keywords, the neuromorphic device outputs the keyword as a speech recognition result via the output layer 850. If the received audio signal does not correspond to one of the preset keywords, the neuromorphic device outputs a null value or receives a next audio signal without outputting a value, and repeats the calculations performed by the layers. In one embodiment, the keywords may be set to approximately 20 or less.
[0112] In one embodiment, if an audio signal corresponds to a negative signal similar to at least one of a plurality of preset keywords, the output layer 850 outputs a null value as the speech recognition result, or receives the next audio signal without outputting a null value and repeats the calculations by each layer described above. A negative signal may refer to a keyword that is similar to one of a plurality of preset keywords and is likely to be recognized as a similar keyword by the neural network. For example, if the preset keyword is "auto care," "watch" may be a negative signal. Therefore, the negative signal may be preset to prevent the neural network from erroneously recognizing a negative signal similar to a keyword as a keyword.
[0113] In one embodiment, the neural network further includes an analog-to-digital converter (ADC) that converts the speech recognition result corresponding to an analog signal into a digital signal, and the feature classification result or speech recognition result output from the output layer may be converted into a digital signal by the analog-to-digital converter and then output via the input / output interface.
[0114] In one embodiment, if the neural network determines that the audio signal does not correspond to any of a plurality of preset keywords based on the mel spectrogram output from the mel generation layer 820, the neural network does not input the mel spectrogram to the hidden layers 830 and 840 and outputs a null value as the speech recognition result, or does not output the mel spectrogram and receives the next audio signal to repeat the calculations in each layer. That is, when the neural network receives an audio signal, it performs calculations on the data using a default value up to the Fourier transform layer 810 and the mel generation layer 820. If the neural network determines that the audio signal is not a human voice or does not correspond to any of a plurality of preset keywords based on the mel spectrogram, it may not send the data to the hidden layers 830 and 840 and / or the output layer 850. Since it is not necessary to perform calculations in the hidden layers 830 and 840 and / or the output layer 850 for all audio signals, the amount of calculations is reduced, thereby improving the calculation efficiency of the neuromorphic device and preventing excessive heat generation in the case of a PIM structure.
[0115] Specifically, when the neural network receives an audio signal and uses an inference model trained with predetermined training data to determine whether the received input signal is included in the trained model data, the neural network can suspend calculation and switch to a low-power mode (sleep mode). In other words, if the input signal is merely noise, including everyday noise in the surrounding environment, unnecessary power consumption can be minimized by not performing calculations by the hidden layers 830 and 840. On the other hand, if the received input signal is included in the trained model data, the neural network can perform classification and output a speech recognition result by the next hidden layers 830 and 840 and the output layer 850, and then switch to a low-power mode (sleep mode).
[0116] On the other hand, each layer 810 to 850 of the neural network is not physically separated, but may represent a core in which the variable resistance element value of the synapse is changed according to the weight of each layer. Referring again to FIG. 3 for explanation, each layer 810 to 850 of the neural network may be realized by the neural network 330.
[0117] For example, when an audio signal is received via the input / output interface 310, the audio signal may be converted into an analog signal by the digital-to-analog converter 320 and input to the neural network 330, in which variable resistor values of synapses are set corresponding to the weights of the Fourier transform layer 810, to generate a frequency signal in which the time band of the audio signal is converted into a frequency band. The frequency signal may then be converted into a digital signal by the analog-to-digital converter 340 and input to the arithmetic circuit 350. Alternatively, the frequency signal, in which the output of the arithmetic circuit 350 is converted into an analog signal by the digital-to-analog converter 320, may be input to the neural network 330, in which variable resistor values of synapses are set corresponding to the weights of the mel generation layer 820, to generate a mel spectrogram of the frequency signal. The mel spectrogram may then be converted into a digital signal by the analog-to-digital converter 340 and input to the arithmetic circuit 350. Similarly, the process of classifying the features of an audio signal based on the neural network 330 in which the variable resistance element values of the synapses are set corresponding to the weights of one or more hidden layers 830, 840 and the output layer 850, and outputting the speech recognition results overlaps with the process described above, and therefore will not be described here.
[0118] In one embodiment, the neural network may be trained using predetermined training data. For example, the predetermined training data may be audio data generated by combining a plurality of predetermined keywords and a noise set. In the field of speech recognition, training data used to train a neural network is primarily data containing a mixture of a plurality of predetermined keywords and noise, such as everyday noise, and is then repeatedly trained to build a neural network with good recognition performance. However, building a neural network with good keyword recognition performance using keyword speech in a noisy environment as training data requires a large amount of training data and a training process. In contrast, first training a neural network using a plurality of predetermined keyword speeches as training data and then training the neural network using a data set generated by combining the plurality of predetermined keyword speeches and noise data requires a smaller amount of training data than the above-mentioned training method, which is more efficient and reduces the cost of generating training data.
[0119] In one embodiment, one or more hidden layers 830, 840 may be configured with weights adjusted based on predetermined training data. The calculations of one or more hidden layers 830, 840 are performed using the resistance values of the variable resistance elements in the synapses of the neural network as weights. Here, each weight may be adjusted based on the predetermined training data, i.e., audio data generated by combining a plurality of preset keyword voices and noise sets.
[0120] FIG. 9 is an exemplary diagram illustrating an implementation of a crossbar array circuit according to one embodiment.
[0121] 9, a crossbar array circuit 900 that realizes a neural network may include multiple subcircuits 910, 920, and 930. The subcircuits 910, 920, and 930 are circuits formed by combining multiple cores that make up the crossbar array circuit 900, and synaptic weights may be set so that each of the subcircuits 910, 920, and 930 corresponds to a respective layer that makes up the neural network.
[0122] In one embodiment, the crossbar array circuit 900 may include a first subcircuit 910, a second subcircuit 920, and a third subcircuit 930, where the first subcircuit 910 may implement a Fourier transform layer, the second subcircuit 920 may implement a Mel generation layer, and the third subcircuit 930 may implement an output layer.
[0123] In one embodiment, weights corresponding to the Fourier transform layer may be stored in the synapses of the first subcircuit 910. For example, the first subcircuit 910 may receive a time-domain audio signal as input data via a presynaptic neuron, and send a synaptic signal through a synapse in which weights corresponding to the Fourier transform layer are stored as output data via a post-synaptic neuron, thereby generating a frequency signal.
[0124] In one embodiment, weights corresponding to the Mel generation layer may be stored in synapses of the second subcircuit 920. For example, the second subcircuit 920 can generate a Mel spectrogram by receiving a frequency signal as input data via a presynaptic neuron and transmitting a synaptic signal through a synapse in which weights corresponding to the Mel generation layer are stored as output data via a post-synaptic neuron.
[0125] In one embodiment, weights corresponding to the output layer may be stored in synapses of the third subcircuit 930. For example, the third subcircuit 930 may receive a mel spectrogram or a feature of the mel spectrogram as input data via a presynaptic neuron, and transmit a synaptic signal that has passed through a synapse in which weights corresponding to the output layer are stored as output data via a post-synaptic neuron, thereby outputting a speech recognition result.
[0126] Similarly, the crossbar array circuit 900 may further include another sub-circuit (not shown) that realizes a hidden layer in addition to the first to third sub-circuits 910, 920, and 930. The other sub-circuit (not shown) that realizes the hidden layer may be a circuit in which weights corresponding to the hidden layer are stored in synapses. For example, the sub-circuit (not shown) may receive a mel spectrogram as input data via a pre-synaptic neuron, and transmit a synaptic signal that has passed through a synapse in which weights corresponding to the hidden layer are stored as output data via a post-synaptic neuron, thereby classifying the audio features of the mel spectrogram.
[0127] Meanwhile, input data for each subcircuit 910, 920, 930 can be received via row or column wiring connected to the subcircuit implementing the previous layer, rather than via presynaptic neurons. Similarly, output data for each subcircuit 910, 920, 930 can be transmitted via row or column wiring connected to the subcircuit implementing the next layer, rather than via post-synaptic neurons. The manner in which each subcircuit 910, 920, 930 transmits and receives input and output data is not limited thereto. Also, while FIG. 9 shows the crossbar array circuit 900 including three subcircuits 910, 920, 930, the number of subcircuits is not limited thereto, and the areas, synapse configurations, and locations of the subcircuits 910, 920, 930 in the crossbar array circuit 900 are not limited thereto.
[0128] The crossbar array circuit 900 according to the embodiment described above is included in on-chip memory and can implement a neural network using PIM-based calculations. Such a neuromorphic device not only reduces the chip volume but also provides high performance relative to power consumption, making it possible to output speech recognition results for a relatively small number of keywords very efficiently.
[0129] FIG. 10 is a flowchart of a method of operating a neuromorphic device according to one embodiment.
[0130] In step 1010, a neuromorphic device (hereinafter referred to as "device") may receive an audio signal.
[0131] In step 1020, the device can input the audio signal to a neural network trained based on predetermined training data and output a speech recognition result.
[0132] In one embodiment, the neural network may include a Fourier transform layer that converts an audio signal from a time domain to a frequency domain to generate a frequency signal, a mel generation layer that generates a mel spectrogram from the frequency signal, and an output layer that outputs a speech recognition result based on speech features of the mel spectrogram.
[0133] In one embodiment, the neural network may further include one or more hidden layers that classify audio features in a mel spectrogram, where the one or more hidden layers may be located between the mel generation layer and the output layer.
[0134] In one embodiment, the output layer can output a speech recognition result by determining whether the audio signal corresponds to any one of a plurality of pre-defined keywords as a result of the classification of the speech features by the hidden layer.
[0135] In one embodiment, the predetermined training data is audio data generated using a combination of a plurality of preset keywords and noise sets, and one or more hidden layers may be configured with weights adjusted based on the predetermined training data.
[0136] In one embodiment, the neural network may output a null value as a speech recognition result if the audio signal corresponds to a negative signal similar to at least one of a plurality of predefined keywords.
[0137] In one embodiment, the device may be one in which a memory cell (or on-chip memory) is directly connected to a processor, and performs PIM (Processing-in-Memory) based operations.
[0138] In one embodiment, the device may convert an audio signal corresponding to a digital signal into an analog signal and input it to a Fourier transform layer, and convert the voice recognition result corresponding to the analog signal output from the output layer into a digital signal.
[0139] In one embodiment, in response to determining, based on the mel spectrogram, that the audio signal does not correspond to any of a plurality of preset keywords, the device may not input the mel spectrogram to the hidden layer.
[0140] 11 is a block diagram of a neuromorphic device according to one embodiment of the present invention. In the following description, explanations that overlap with the explanations regarding the above-mentioned FIGS.
[0141] The processor serves to control the overall functions for running the neuromorphic device 1100. For example, the processor 1110 controls the neuromorphic device 1100 overall by executing a program stored in the memory 1120 in the neuromorphic device 1100. The processor 1110 can be realized by a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), or the like provided in the neuromorphic device 1100, but is not limited to these.
[0142] The memory 1120 is hardware that stores various data processed within the neuromorphic device 1100. For example, the memory can store data that has been processed by the neuromorphic device 1100 and data to be processed. The memory 1120 can also store applications, drivers, etc. that are run by the neuromorphic device 1100. The memory includes random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray or other optical disk storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.
[0143] Unlike the neural network 1 shown in Figure 1, an actual neural network driven by the neuromorphic device 1100 may be realized with a more complex architecture, which may require the processor 1110 to perform operations with a significantly larger operation count, reaching hundreds of millions to tens of billions, and may also dramatically increase the frequency with which the processor 1110 accesses the memory 1130 for operations.
[0144] Therefore, the memory 1120 may be an on-chip memory. The neuromorphic device 1100 according to one embodiment includes the memory 1120 only in the form of an on-chip memory, and can perform calculations without accessing the external memory 1130. For example, the memory 1120 may be an SRAM implemented in the form of an on-chip memory. In this case, unlike the above, the type of memory typically used for the external memory 1130, such as a DRAM, a ROM, a HDD, or an SSD, may not be used for the memory 1120.
[0145] The processor 1110 can read / write neural network data, such as voice data (audio signals), from the memory 1120 and execute the neural network using the read / written data. When the neural network is executed, the processor 1110 can repeatedly perform operations on the audio signals to generate data related to the output. That is, the processor 1110 can receive the audio signals at predetermined intervals and repeatedly perform the operations related to the layers described above at each predetermined interval.
[0146] In one embodiment, the processor 1110 is directly connected to the memory 1120 and can perform PIM-based operations to drive the neural network, where the memory 1120 may include a crossbar array circuit that receives instructions from the processor 1110 and performs in-memory operations.
[0147] In one embodiment, the neuromorphic device 1100 may be a server. The server may be implemented as a computer device or multiple computer devices that communicate over a network to provide instructions, code, files, content, services, etc. The server may receive data necessary for speech recognition and perform speech recognition based on the received data.
[0148] Meanwhile, embodiments of the present invention may be realized in the form of a computer program executable by various components on a computer, and such a computer program may be recorded on a computer-readable medium, including magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROMs, RAMs, and flash memories.
[0149] On the other hand, the computer program may be one specially designed and constructed for the present invention, or it may be one that is well known and available to those skilled in the art of computer software. Examples of computer programs include not only machine language code such as that produced by a compiler, but also high-level language code that is executed by a computer using an interpreter, etc.
[0150] According to one embodiment, methods according to various embodiments of the present disclosure can be provided in a computer program product. The computer program product can be traded as a commodity between a seller and a buyer. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., a compact disc read only memory (CD-ROM)) or can be distributed through an application store (e.g., the Play Store). TM) or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored on or temporarily generated by a machine-readable storage medium, such as the memory of a manufacturer's server, an application store server, or an intermediary server.
[0151] Unless explicitly stated or stated to the contrary, steps constituting a method according to the present invention may be performed in any suitable order. The present invention is not necessarily limited to the order of steps described above. The use of all examples or exemplary terms in the present invention is merely for the purpose of explaining the present invention in detail, and the scope of the present invention is not limited by the examples or exemplary terms unless limited by the claims. Furthermore, those skilled in the art will understand that various modifications, combinations, and variations can be made within the scope of the claims or their equivalents, depending on design conditions and factors.
[0152] Therefore, the concept of the present invention should not be limited to the above-described embodiments, and not only the scope of the appended claims, but also all scopes equivalent to or modified equivalently from the scope of the claims, can be said to fall within the scope of the concept of the present invention.
Claims
[Claim 1] In neuromorphic devices that realize neural networks, a memory having at least one program stored therein; an on-chip memory including a crossbar array circuit; at least one processor that executes the at least one program to drive the neural network; The at least one processor Receives audio signals, inputting the audio signal into the neural network trained based on predetermined training data, and outputting a speech recognition result; The neural network A neuromorphic device comprising: a Fourier transform layer that converts the audio signal from a time domain to a frequency domain to generate a frequency signal; a mel generation layer that generates a mel spectrogram from the frequency signal; and an output layer that outputs the speech recognition result based on speech features of the mel spectrogram.