A visual channel coding method based on spiking neural network

By using linear delay time domain coding method in pulsed neural networks, the data point intensity value is mapped into the simulation time window, which solves the problem of mapping spatial domain information into the time domain in pulsed neural networks, and efficient information coding and low-latency processing are achieved.

CN114638286BActive Publication Date: 2025-05-09FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210183074.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-05-09
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

The existing pulse neural networks are difficult to effectively map spatial domain information to the time domain while fully retaining information characteristics, and there are problems of high computational complexity and large delays.

Method used

The linear delay time domain encoding method is adopted to map the data point intensity value approximately linearly into the simulation time window, and information is transmitted through pulse distribution moments to realize time domain encoding, and support parallel pipeline processing to reduce system delay.

Benefits of technology

It improves the distinction of data with high characteristic values, reduces the length of the time window, reduces the system delay, and achieves low power consumption operation, which conforms to the biological neural impulse distribution law.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114638286B_ABST
    Figure CN114638286B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of brain-like technology, and specifically is a visual channel coding method based on a pulse neural network. The method of the present invention includes: using linear delay time domain coding to approximately linearly map the data point intensity value into a simulation time window, wherein the information is transmitted in the pulse neural network with the pulse emission time as the carrier. Taking the image classification task as an example, the higher the intensity value of the image pixel point, the earlier the pulse emission time, and each input pixel point generates and only generates one pulse in a simulation time window. The present invention can improve the discrimination of data with high eigenvalues ​​and improve the coding accuracy; reduce the time window length without losing effective information, reduce system delay, and can operate with low power consumption; the pulse emission time is negatively correlated with the information intensity, which conforms to the law of nerve impulse emission in organisms and has biological explainability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of brain-like technology, and specifically relates to a brain-like visual channel encoding method. Background Art

[0002] In living organisms, neurons use action potentials or pulses as a common language to talk to each other and perform calculations for learning and decision-making. Pulses are in the form of nerve impulses, which are either (digital) time-point events with or without time in the time domain, which allows long-distance neural communication and low-energy neural computation. However, external internal stimuli such as light, sound, pain, and hunger are mainly analog continuous variables in time. These analog spatiotemporal information needs to be converted into corresponding pulse trains by specialized sensory neurons, also called afferent neurons, according to one or more neural coding algorithms. The pulse-coded information is then transmitted to the central nervous system for processing. This process is called sensory transduction.

[0003] The Spiking Neural Network (SNN) is a neural network constructed by simulating the working mechanism of neurons in the cerebral cortex of mammals, and is known as the third generation of artificial neural networks. SNN has fully learned the above-mentioned way of biological neural information transmission. Unlike the mode of sending information in each round in traditional artificial neural networks, SNN only transmits information to the next level at some special moments, that is, when the membrane potential of the neuron reaches the threshold voltage. In addition, the way of information transmission in SNN is also different from that in traditional ANN: the latter often uses the strength value of the signal as the information carrier, while in SNN, the information is carried on the sum of a series of impulse functions, the so-called spike train.

[0004] The pulse data in SNN belongs to time domain information. The data of existing ANN networks can be understood as spatial domain information. The recognition information of common classification tasks, such as license plate recognition, handwritten digit recognition, and face recognition, also inputs spatial domain information. How to effectively map spatial domain information to time domain while fully retaining information characteristics, reduce energy consumption and generate timing pulses that conform to neurobiological principles is the key to the development of pulse neural networks.

[0005] The construction process of the spiking neural network mainly includes three parts: constructing a spiking neuron model; constructing a neural pulse sequence and performing the encoding process; and training the spiking neural network. Different neural information encoding methods express the input information into different pulse sequences. Neural information encoding mainly includes information feature extraction and pulse sequence encoding. The main pulse sequence encoding includes pulse frequency encoding (ratecoding) and time coding (temporalcoding). The mainstream spiking neural network uses frequency encoding (rate encoding). Neurons will adjust the pulse emission rate according to different external stimuli, such as the size of the image pixel value, that is, rate coding records information with the frequency of pulse emission within a certain period of time. Its encoding steps are complex and the information validity is low. It only records the number of pulses in the encoding window and lacks expression for information that changes over time. In most neural systems, effective information lies in the precise moment when the action potential is generated, so temporal encoding (temporal coding) is proposed. At present, temporal coding includes latency coding, phase coding, rank-order coding, population coding, etc. Among them, the idea embodied in delay coding is that the characteristic of neuron pulse emission is that the form of pulses is fixed, and there are only differences in quantity and time. The stronger the stimulus, the earlier the pulse is emitted. For temporal coding, each input will first pass through an input neural layer for encoding and processing before being passed into the spiking neural network for classification. In the layer, each input neuron will generate an independent Poisson pulse train based on the preset pulse emission frequency. However, by encoding individual elements in the input data set, only small data sets can be reliably learned by the network based on this encoding method. When using large data sets, time domain coding will become a very critical issue. If the input variables in the receptive field are still processed independently one by one, the amount of calculation will increase exponentially, which is considerable and will cause huge delays. Summary of the invention

[0006] The purpose of the present invention is to provide a visual channel coding method based on a pulse neural network to solve the above problems.

[0007] The present invention provides a visual channel encoding method based on a pulse neural network, including: using a linear delay time domain encoder to approximately linearly map the data point intensity value to a simulation time window, wherein information is transmitted in the pulse neural network with the pulse emission time as a carrier. Taking the image classification task as an example, the higher the image pixel intensity value, the earlier the pulse emission time, and each input pixel generates and only generates one pulse in a simulation time window. The present invention helps to improve the discrimination of data with high eigenvalues, which is beneficial to improving the encoding accuracy; information with high eigenvalues ​​emits pulses first, which can reduce the time window length without losing effective information, and at the same time uses a parallel pipeline to process input encoding, greatly reducing system delay; the form of single pulse emission in the time window helps the system to run with low power consumption; the pulse emission time is negatively correlated with the information intensity, which conforms to the law of biological nerve impulse emission and has biological interpretability. In addition, the present invention also provides an RRAM-based in-memory computing system based on this encoding method.

[0008] The specific steps of the present invention are:

[0009] (1) External input data to be processed (taking images as an example);

[0010] (2) Data processing: Process the image and extract information to obtain the grayscale value of each pixel in the image and digitize the image into a grayscale value matrix;

[0011] (3) Data normalization: normalize the image gray value matrix;

[0012] (4) Linear delay time domain coding;

[0013] (5) Pulse firing: Input neurons fire pulses encoded by linear delay sequences;

[0014] (6) Training of spiking neural networks: Training spiking neural networks based on linear delay temporal coding and adjusting parameters to obtain a suitable network model;

[0015] (7) Testing of the spiking neural network: For the test set data, the linear delay time series encoding is performed to obtain the test set information in the form of a time series, and the input information is deployed in the spiking neural network model obtained in the above step to perform classification processing;

[0016] (8) Classifier testing: Get the classification results of the test set.

[0017] Optionally, for linear delay coding, the specific coding method is as follows:

[0018] For the normalized image, assuming its size is (C, H, W), the normalized gray value of the element at the position (Ci, Hi, Wi) is Ii, then the pulse emission time corresponding to the pixel is:

[0019] ti=[(T–1)(1–Ii)]

[0020] Where [·] is the Gaussian rounding function, and T is the total simulation time window length. Using this linear delay timing encoder, the intensity value of the pixel point can be approximately linearly mapped to the simulation time window. The higher the intensity value, the earlier the pulse is emitted, and each input pixel point generates and only generates one pulse.

[0021] The present invention proposes a visual channel encoding method based on a pulse neural network to solve the complexity of the encoding of the pulse neural network. By using linear delay time domain coding, the intensity value of the pixel point can be approximately linearly mapped to the simulation time window. The higher the pixel intensity value, the earlier the pulse emission time, and each input pixel point generates and only generates one pulse in a simulation time window. The above method helps to improve the discrimination of effective data, supports parallel processing and pipeline processing, and can realize a low-delay, low-computational complexity pulse timing coding system with biological interpretability.

[0022] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are used to provide further understanding of the embodiments of the present application and constitute a part of the specification. Together with the following specific implementation methods, they are used to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application.

[0024] Figure 1 The present invention is a preferred embodiment of the present invention - a flow chart of a visual channel encoding method based on a pulse neural network.

[0025] Figure 2 A preferred embodiment of the present invention is a data flow diagram of a visual channel encoding method based on a pulse neural network.

[0026] Figure 3 A schematic diagram of the pulse emission of input layer neurons based on a visual channel coding method is a preferred embodiment of the present invention.

[0027] Figure 4 A schematic diagram of a RRAM in-memory computing system based on a visual channel coding method is a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be further described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0029] It should be noted that the technical solutions between the various embodiments of the present application can be combined with each other, but it must be based on the fact that ordinary technicians in the field can implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0030] Figure 1 This is a flow chart of a preferred embodiment of the present invention - a visual channel coding method based on a pulse neural network. Specifically, it includes the following steps:

[0031] Step 1: external input processing data, here we take image as an example;

[0032] Step 2: Process the image to obtain the gray value of each pixel in the image and digitize the image into a gray value matrix;

[0033] Step 3, normalizing the image gray value matrix;

[0034] Step 4: For the normalized image, assuming that its size is (C, H, W), the normalized gray value of the element at the position (Ci, Hi, Wi) is Ii, then the pulse emission time corresponding to the element is:

[0035] ti=[(T–1)(1–Ii)]

[0036] Where [·] is the Gaussian rounding function, and T is the total simulation time window length. Using this linear delay timing encoder, the intensity value of the pixel point can be approximately linearly mapped to the simulation time window. The higher the intensity value, the earlier the pulse is emitted, and each input pixel point generates and only generates one pulse.

[0037] Step 5, the input neuron emits pulses encoded with linear delay timing;

[0038] Step 6, training a spiking neural network based on linear delay temporal coding;

[0039] Step 7, for network testing, the test set is linearly delayed and time-series encoded using steps 1-5 to obtain test set information in the form of time series, and the test set graph is classified based on the spiking neural network using the neural network weights obtained in step 6;

[0040] Step 8, obtain the test set classification results; verify the network function.

[0041] Figure 2 A preferred embodiment of the present application - a data flow diagram of a visual channel encoding method based on a pulse neural network.

[0042] Among them, 201 is an image of a dataset to be classified input from the outside, 202 is a pixel in an image sample, including pixel location information and color information, 203 is the image pixel processing process, including taking the grayscale value of the pixel and normalizing the grayscale value, 204 is the grayscale value of the extracted pixel, 205 is a linear delay encoder, 206 is an input neuron in the input layer of this pulse neural network, 207 is the pulse emitted by the input neuron after being encoded by the linear delay encoder, 208 is the start time of the time axis, and 209 is the time interval between the start time and the pulse generation time, that is, the pulse emission time. In a time window, after being encoded by the linear delay time domain encoder, the information of the input image is encoded into time domain information, and the input neuron emits pulses. A single time window only generates a single pulse, and its information is determined by the pulse emission time, which is independent of the pulse amplitude and pulse width.

[0043] Figure 3This is a schematic diagram of the pulse emission of input layer neurons based on a visual channel coding method, which is a preferred embodiment of the present invention. 301, 302, 303, and 304 are four pixels on an input image (due to space limitations, this schematic diagram only takes these four points as an example). The depth of their colors represents the depth of the color of the pixel in the image. The darker the color, the greater the grayscale value. For image recognition tasks with a white background, the larger the grayscale value, the more important the pixel is. 305 is a linear delay encoder, which performs a time-series encoding operation on the grayscale value of the input pixel to obtain the corresponding pulse emission time. 306, 307, 308, and 309 are the input layer neurons of the pulse neural network corresponding to the four pixels, and 310, 311, 312, and 313 are the pulse emission of the corresponding input layer neurons. It can be seen from the figure that the pulse corresponding to 313 occurs the earliest, because the grayscale value of the corresponding pixel 304 is the highest. Due to the use of a negatively correlated linear delay encoder, the higher the grayscale value of the pixel, the earlier the pulse emission time. Grayscale value is negatively correlated with pulse emission time, so the input neurons corresponding to pixels containing important information will release neural pulse information earlier, which is consistent with biological principles. Since invalid information is released late, the time window cutoff time can be selected through experiments for pipeline processing or parallel processing, so that invalid information can be eliminated without losing valid information, greatly reducing system delay. Furthermore, it can be seen from the figure that within a simulation time window shown in the figure, each pixel will only generate one pulse, and the low activation state system is conducive to the low-power hardware implementation of the encoder.

[0044] Figure 4This is a schematic diagram of an embodiment of the present invention - an RRAM in-memory computing system based on visual channel coding. Among them, 401 is the input data, which can be a processed picture, voice, etc. 402 is the proposed linear delay coding circuit, which can perform timing coding operations on the input data to obtain the moment of corresponding pulse emission and generate pulses with equal pulse width and amplitude. 403 is an RRAM-based storage and computing integrated array, which is composed of 404 units. 404 is a 1T1R (1transistor 1RRAM) structure composed of an RRAM and a transistor in series, the transistor is used for gating, and the memristor is used for storing weights and calculations. 405 is the main controller of the system, which is used for row selection, column selection, module control, etc. The operation mode of the system is that the controller controls the system, and the input data enters the linear delay coding circuit after processing, and is encoded into a voltage signal with different pulse emission time but the same pulse width and amplitude, and is placed in the selected row as the input voltage signal. The 1T1R unit uses the high and low conductivity value of RRAM as the form of storage weight. Through the voltage-current formula and Kirchhoff's law, the product of input and weight is obtained in the form of current and input into the pulse neuron 406. The pulse neuron 406 accumulates the current nonlinearly according to the intensity of the input current, which is reflected in the membrane potential of the pulse neural network. Whether the pulse can be emitted is determined according to whether the membrane potential reaches the threshold voltage. Due to the use of a negatively correlated linear delay coding circuit, the stronger the data intensity and the earlier the pulse emission time, the earlier the corresponding level signal is input into the RRAM array, and the pulse neuron 406 will also emit pulses earlier (if the signal strength is sufficient), which can realize the multi-layer expansion of the pulse neural network.

[0045] In summary, the present invention provides a visual channel encoding method based on a pulse neural network. The present invention adopts linear delay time domain coding, which can approximately linearly map the intensity value of the pixel point to the simulation time window. The higher the pixel intensity value, the earlier the pulse emission time, and each input pixel generates and only generates one pulse in a simulation time window. The invention helps to improve the discrimination of effective data, supports parallel processing and pipeline processing, and can realize a low-latency, low-computational complexity pulse timing coding system with biological interpretability. In addition, the present invention provides an RRAM-based in-memory computing system based on this encoding method, which is used to illustrate the application prospects of this application in in-memory computing.

[0046] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A visual channel coding method based on a pulse neural network, characterized in that: A linear delay time domain encoder is used to approximately linearly map the data point intensity value to the simulation time window, where the information is transmitted in the spiking neural network using the pulse emission time as the carrier; the specific steps are: (1) External input of image data to be processed; (2) Data processing: Process the image and extract information to obtain the grayscale value of each pixel in the image, and digitize the image into a grayscale value matrix; (3) Data normalization: normalize the image gray value matrix; (4) Perform linear delay time domain coding; (5) Pulse emission: The input neuron emits pulses encoded by linear delay sequence; (6) Training of spiking neural networks: Training spiking neural networks based on linear delay temporal coding and adjusting parameters to obtain a suitable network model; (7) Testing of the spiking neural network: For the test set data, after the above-mentioned linear delay time series encoding, the test set information in the form of time series is obtained, and the input information is deployed in the spiking neural network model obtained in the above steps for classification processing; (8) Classifier testing: Get the classification results of the test set.

2. The visual channel coding method according to claim 1, characterized in that: The linear delay time domain coding is performed, and the specific coding method is as follows: For a normalized image, assuming its size is (C, H, W), the normalized grayscale value of the element at the position (Ci, Hi, Wi) is Ii, then the pulse emission time corresponding to the pixel point of the image is: ti = [(T – 1)(1 – Ii)] Where [·] is a Gaussian rounding function, and T is the total simulation time window length. Using this linear delay timing encoder, the intensity value of the pixel point is approximately linearly mapped to the simulation time window. The higher the intensity value, the earlier the pulse is emitted, and each input pixel point generates and only generates one pulse.

Citation Information

Patent Citations

  • Hardware friendly pulse neural network model based on STDP non-supervised learning algorithm

    CN107092959A

  • SAR image ship target identification method based on pulse neural network

    CN113111758A