A neuromorphic visual target recognition method and device
Patent Information
- Application Number
- CN202311300700.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-09
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-09
AI Technical Summary
现有的技术中,Gabor滤波器可以实现方向特征的提取,然而缺乏对时空相关特征的提取
[0039]本申请所提供的一种神经形态视觉目标识别方法及装置,利用调和平均时间的加权时间表面的多尺度以及层次模型来实现对时空脉冲事件进行时空特征提取,充分利用时间在不同空间尺度上的信息。调和平均时间的加权时间表面表示以输入事件为中心的时空邻域内的综合活动情况,并一步考虑过去事件流的最近事件簇来进一步提取时空特征。调和平均时间的加权时间表面的多尺度层次模型考虑了最近过去事件时间和过去事件簇时间的调和平均时间对当前传入事件的不同影响权重大小,并通过多尺度层次模型综合考虑了不同空间尺度下的特征提取以及层次间的复杂特征的提取,进而为分类器提供更加具有区分度的特征。Tempotron神经元的单层脉冲神经网络对特征事件进行分类,通过采用群体编码的多数投票方案提高分类效果,并使用多尺度特征加强脉冲发放性,进而有效解决神经形态视觉的目标识别与分类问题。
Smart Images

Figure CN117370858B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of target recognition technology in neuromorphic vision, and particularly relates to a method and apparatus for target recognition in neuromorphic vision. Background Technology
[0002] With the rapid development of computer vision, traditional image sensors can no longer meet the demands of high-speed, dynamic scenes. Traditional image sensors use fixed frame rates and exposure times when acquiring images, making it impossible to accurately capture rapidly changing scenes. To address this issue, event cameras have emerged. An event camera is a new type of image sensor that, unlike traditional frame sensors, can capture changes in a scene at the pixel level and output a spatiotemporal pulse event stream in real time, rather than continuous frame images. This event stream contains the coordinates, time, and polarity information of pixels that have changed in the scene. This allows event cameras to capture rapidly changing scenes on a microsecond-level timescale, providing high temporal resolution image data with characteristics of high dynamic range, high temporal resolution, low power consumption, and low information redundancy.
[0003] Spiking Neural Networks (SNNs) are brain-inspired computational models that simulate the working principle of neurons in biological nervous systems. Unlike traditional neural networks, SNNs transmit and process information by simulating the spiking firing patterns of neurons, making them naturally suited for event-driven algorithms.
[0004] In recent years, event-driven spiking neural network (SNN) algorithms have been proposed. Inspired by the multi-layered feedforward organization of the brain's visual cortex, Orchard et al. proposed an event-driven classification SNN model called HFirst, which uses Gabor filters for event-driven convolutional operations to extract features and employs a statistical classification method. Xiao et al. proposed an event-driven hierarchical model that uses Gabor filters to extract features and an event-driven Tempotron learning algorithm for classification. Furthermore, Lagorce et al. proposed a model called HOTS, which introduces the concept of a time surface to represent past dynamic information surrounding an event and uses a clustering algorithm to learn a time surface prototype template for feature extraction.
[0005] Utilizing precise temporal information from event streams to extract effective spatiotemporal features is one of the main tasks in neuromorphic visual target recognition and classification based on event processing. Existing technologies, such as Gabor filters, can extract directional features but lack the ability to extract spatiotemporally relevant features. Furthermore, HOTS only considers the most recent past time of pixels with the same polarity during the generation of the time surface, making it overly sensitive to instantaneous changes in the scene. Summary of the Invention
[0006] The purpose of this application is to provide a neuromorphic visual target recognition method and apparatus to effectively extract multi-scale spatiotemporal features and improve the accuracy of neuromorphic visual target recognition.
[0007] To achieve the above objectives, the present invention provides the following solution:
[0008] A neuromorphic visual target recognition method includes:
[0009] Acquire the spatiotemporal pulse event stream asynchronously output by the event camera, and determine the harmonic weighted time surface of each spatiotemporal pulse event in the spatiotemporal pulse event stream at three different spatial scales;
[0010] Based on the harmonic weighted time surface of each spatiotemporal impulse event at three different spatial scales, K1 first-level harmonic weighted time surface prototypes at any spatial scale are learned.
[0011] Each spatiotemporal pulse event in the spatiotemporal pulse event stream is matched with the harmonic weighted time surface at three different spatial scales and K1 first-layer harmonic weighted time surface prototypes at the corresponding spatial scales. The first-layer harmonic weighted time surface prototype with the highest similarity is found. The number corresponding to the first-layer harmonic weighted time surface prototype with the highest similarity is used as the first-layer polarity of the spatiotemporal pulse event. The first-layer polarity is used to replace the original light intensity change polarity in the spatiotemporal pulse event to generate the spatiotemporal feature event stream corresponding to the spatiotemporal pulse event stream.
[0012] Determine the harmonic weighted time surface of each spatiotemporal feature event in the spatiotemporal feature event stream at the fourth spatial scale, and learn K2 second-layer harmonic weighted time surface prototypes based on the harmonic weighted time surface of each spatiotemporal feature event at the fourth spatial scale.
[0013] The event stream of spatiotemporal pulse events to be identified is obtained from the event camera. The harmonic weighted time surface of each spatiotemporal pulse event at three different spatial scales is determined. It is matched with K1 prototypes of the first-layer harmonic weighted time surface at the corresponding spatial scale. The number corresponding to the prototype of the first-layer harmonic weighted time surface with the highest similarity is used as the first-layer polarity. The first-layer polarity is used to replace the original light intensity change polarity in the spatiotemporal pulse event to be identified, so as to obtain the first-layer spatiotemporal feature event stream corresponding to the three different spatial scales.
[0014] Determine the harmonic weighted time surface of each spatiotemporal feature event in the spatiotemporal feature event stream corresponding to three different spatial scales at the fourth spatial scale, and match it with K2 second-layer harmonic weighted time surface prototypes. The number corresponding to the second-layer harmonic weighted time surface prototype with the highest similarity is used as the second-layer polarity.
[0015] The original polarity in the spatiotemporal pulse event is replaced by the first-layer polarity and the second-layer polarity corresponding to each spatiotemporal pulse event in the spatiotemporal pulse event stream to be identified, so as to obtain the spatiotemporal feature event streams of the second-layer output corresponding to three different spatial scales.
[0016] Pooling is performed on the three spatiotemporal feature event streams output from the second layer. Pulse coding is then performed on the three pooled spatiotemporal feature event streams to obtain three pulse sequences, which are then input into a single-layer spiking neural network with population coding Tempotron neurons for classification.
[0017] Furthermore, the spatiotemporal pulse event stream is described using an address event representation protocol as follows:
[0018]
[0019] Where e i Let x represent the i-th spatiotemporal pulse event in the spatiotemporal pulse event stream. i and y i Let t represent the pixel coordinates of the i-th spatiotemporal impulse event. i p represents the timestamp of the i-th spatiotemporal impulse event. i This indicates the polarity of the light intensity change at the pixel corresponding to the i-th spatiotemporal pulse event, where +1 / -1 represent the "ON event" of increased light intensity and the "OFF event" of decreased light intensity, respectively. This represents the number of spatiotemporal pulse events in the spatiotemporal pulse event stream.
[0020] Further, determining the harmonic weighted time surface of each spatiotemporal pulse event in the spatiotemporal pulse event stream at three different spatial scales includes:
[0021] The spatial scale is determined by the neighborhood radius, R m Let m be the neighborhood radius of the m-th spatial scale, and let the spatial scale be (2R). m +1)×(2R m +1);
[0022] Using t(e) i ,R m (b, x, y) represents the i-th spatiotemporal impulse event e i The time of b most recent past events for pixels with x and y offsets within the receptive field at the m-th spatial scale, where x, y ∈ [-R]. m,R m ];
[0023] Using formula Determine the i-th spatiotemporal impulse event e i The harmonic average relative past time of b most recent past events for pixels with x and y offsets within the receptive field at the m-th spatial scale. Here, relative past time refers to t. i -t j That is, Δt i,j That is, the difference between the time of the current event and the time of past events, which is the harmonic average of b relative past times;
[0024] Using formula Determine the i-th spatiotemporal impulse event e i Within the receptive field of the m-th spatial scale, the harmonic average of b most recent past events relative to the past time is the time surface point of the exponential kernel, where τ is the time constant of the exponential kernel.
[0025] use Determine the i-th spatiotemporal impulse event e i The relative past time of the most recent event for a pixel with x and y offsets within the receptive field at the m-th spatial scale. Where t i -t last t represents the difference between the time of the current event and the time of the most recent past event. last It is the i-th spatiotemporal impulse event e i The time of the most recent past event for a pixel with x and y offsets within the receptive field of the m-th spatial scale;
[0026] Using formula Determine the i-th spatiotemporal impulse event e i The time surface point of the relative past time of the most recent past event of a pixel with respect to x and y offsets within the receptive field of the m-th spatial scale;
[0027] Using formula Determine the i-th spatiotemporal impulse event e i The harmonic weighted temporal surface point of the pixels with x and y offsets within the receptive field of the m-th spatial scale, where γ is the spatiotemporal constant;
[0028] use Represents the i-th spatiotemporal impulse event e i For the harmonic weighted time surface at the m-th spatial scale, using the formula:
[0029]
[0030] Determine the spatiotemporal pulse event ei Harmonic weighted time surface Each harmonic weighted time surface point within the receptive field of the m-th spatial scale and utilize The harmonic weighted time surface is normalized to obtain the harmonic weighted time surface.
[0031] Furthermore, based on the harmonic weighted time surfaces of each spatiotemporal impulse event at three different spatial scales, K1 prototypes of the first-layer harmonic weighted time surfaces at any spatial scale are learned, including:
[0032] For any spatial scale, the harmonic weighted time surface of the first K1 spatiotemporal pulse events in the spatiotemporal pulse event stream is selected as K1 initial first-layer harmonic weighted time surface prototypes at this spatial scale, and the harmonic weighted time surface of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream is used to update the K1 first-layer harmonic weighted time surface prototypes at this spatial scale, so as to obtain K1 first-layer harmonic weighted time surface prototypes at any spatial scale.
[0033] Further, the step of updating the first-level harmonic weighted time surface prototypes at the K1 local spatial scales using the harmonic weighted time table of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream includes:
[0034] Calculate the Euclidean distance between the harmonic weighted time surface of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream and the K1 prototypes of the first-layer harmonic weighted time surface at the local spatial scale, and take the prototype of the first-layer harmonic weighted time surface with the shortest Euclidean distance as the prototype of the first-layer harmonic weighted time surface with the highest similarity.
[0035] Then, the first-layer harmonic weighted time surface prototype is updated according to the following formula:
[0036]
[0037] Among them, C k (R m b) indicates that the updated neighborhood radius is R. m The k-th harmonic weighted time surface prototype of the past b events, C′ k (R m b) indicates that the neighborhood radius obtained from the previous update is R. m The k-th harmonic weighted time surface prototype of the past b events, where α is the update amplitude parameter and β is the current spatiotemporal impulse event e. i Harmonic weighted time surface at the m-th spatial scale The cosine distance between it and the first-layer harmonic weighted time surface prototype with the highest similarity.
[0038] This application also proposes a neuromorphic visual target recognition device, including a processor and a memory storing a plurality of computer instructions, which, when executed by the processor, implement the steps of the above method.
[0039] This application provides a neuromorphic visual target recognition method and apparatus that utilizes a multi-scale and hierarchical model of a weighted time surface based on harmonic mean time to extract spatiotemporal features from spatiotemporal impulse events, fully leveraging information from time at different spatial scales. The weighted time surface based on harmonic mean time represents the comprehensive activity within the spatiotemporal neighborhood centered on the input event, and further considers the most recent event clusters from past event flows to extract spatiotemporal features. The multi-scale hierarchical model of the weighted time surface based on harmonic mean time considers the different weights of the harmonic mean time of the most recent past event and the time of past event clusters on the current input event. It comprehensively considers feature extraction at different spatial scales and the extraction of complex features between levels, thus providing the classifier with more discriminative features. A single-layer spiking neural network of Tempotron neurons classifies feature events, improves classification performance by employing a majority voting scheme with population coding, and enhances impulse firing by using multi-scale features, thereby effectively solving the target recognition and classification problem in neuromorphic vision. Attached Figure Description
[0040] Figure 1 This is a flowchart of the neuromorphic visual target recognition method of this application.
[0041] Figure 2 This is a schematic diagram of the neuromorphic visual target recognition method of this application. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0043] This application provides a neuromorphic visual target recognition method and system, which effectively achieves feature fusion based on event processing and improves the accuracy of target recognition and classification in neuromorphic vision. To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] An event camera is a biomimetic visual perception sensor that employs a novel paradigm for acquiring visual information. Unlike traditional visual sensors that acquire images synchronously at a fixed frame rate, the event camera asynchronously encodes visual information about the external scene, forming a continuous stream of spatiotemporal pulse events, and uses an address event representation protocol to express these events. Each pixel in the event camera responds independently and asynchronously to changes in light intensity in the visual scene. Once the light intensity change of a pixel exceeds a preset threshold, the event camera immediately outputs a precise spatiotemporal pulse event. The event camera offers several advantages, such as low power consumption, low information redundancy, and high dynamic range, making it a promising candidate for widespread applications in the field of visual perception.
[0045] One embodiment of this application, such as Figure 1 , Figure 2 As shown, a neuromorphic visual target recognition method is provided, including:
[0046] Step S101: Obtain the spatiotemporal pulse event stream asynchronously output by the event camera, and determine the harmonic weighted time surface of each spatiotemporal pulse event in the spatiotemporal pulse event stream at three different spatial scales.
[0047] In this embodiment, the spatiotemporal pulse event stream asynchronously output by the event camera includes multiple spatiotemporal pulse events, which are described using the address event representation protocol.
[0048] Specifically, using The spatiotemporal pulse event stream is described, where e i Let x represent the i-th spatiotemporal pulse event in the spatiotemporal pulse event stream. i and y i Represents the i-th spatiotemporal impulse event e i pixel coordinates, t i Represents the i-th spatiotemporal impulse event e i timestamp, p i This indicates the polarity of the light intensity change of the pixel corresponding to the i-th spatiotemporal pulse event. +1 / -1 represent the "ON event" of increased light intensity and the "OFF event" of decreased light intensity, respectively. This represents the number of spatiotemporal pulse events in the spatiotemporal pulse event stream.
[0049] The spatial scale is determined by the neighborhood radius, R m Let R represent the radius of the m-th neighborhood. m Confirmed (2R) m +1)×(2R mThe spatial scale of +1) is the size of the harmonic weighted time surface. In this embodiment, m takes values ranging from (1, 2, 3, 4). In a specific embodiment, R1 = 1, R2 = 2, R3 = 3, R4 = 4 is used for the neuromorphic digital dataset, and R1 = 2, R2 = 3, R3 = 4, R4 = 5 is used for the neuromorphic action dataset. R1, R2, and R3 belong to the first layer, and R4 belongs to the second layer.
[0050] In one specific embodiment, the harmonic weighted time surface of each spatiotemporal pulse event in the spatiotemporal pulse event stream at three different spatial scales is determined, including:
[0051] Using the formula t(e i ,R m ,b,x,y)=the top b values of{t j |x j =x i +x,y j =y i +y,p j =p i ,t j <t i} Determine the i-th spatiotemporal pulse event e in the spatiotemporal pulse event stream i The pixel points (x, y) with respect to the x and y offsets within the receptive field at the m-th spatial scale i +x,y i A local memory array of x, y, and b; wherein the local memory array records the times of the b most recent past events of the pixel point at the offset; where x, y ∈ [-R] m ,R m ] and R m Indicates the neighborhood radius;
[0052] Using formula Determine the i-th spatiotemporal pulse event e in the spatiotemporal pulse event stream. i Within the receptive field at the m-th spatial scale, the pixel points with x and y offsets (x i +x,y i +y)b harmonic average relative past time of the most recent past events; the relative past time refers to t i -t j That is, Δt i,j That is, the difference between the time of the current event and the time of the past event, wherein the harmonic average is adjusted relative to the past time by the offset of the pixels b relative to the past time;
[0053] Using formula Determine the i-th spatiotemporal pulse event e in the spatiotemporal pulse event stream. iThe pixel points (x, y) with respect to the x and y offsets within the receptive field at the m-th spatial scale i +x,y i +y) is the harmonic mean relative to the past time time surface point; the harmonic mean relative to the past time time time surface point represents the overall influence of the past b events on the current event of the pixel; where τ is the time constant of the exponential kernel;
[0054] Using formula Determine the i-th spatiotemporal pulse event e in the spatiotemporal pulse event stream. i The pixel points (x, y) with respect to the x and y offsets within the receptive field at the m-th spatial scale i +x,y i +y) the relative past time of the most recent event; where t i -t last t represents the difference between the time of the current event and the time of the most recent past event. last It is the i-th spatiotemporal pulse event e in the spatiotemporal pulse event stream. i The pixel points (x, y) with respect to the x and y offsets within the receptive field at the m-th spatial scale i +x,y i +y) the time of the most recent past event, i.e., t last =max{t j |x j =x i +x,y j =y i +y,p j =p i ,t j ≤t i And x,y∈[-R m ,R m ] and R m Indicates the neighborhood radius;
[0055] Using formula Determine the i-th spatiotemporal pulse event e in the spatiotemporal pulse event stream. i The pixel points (x, y) with respect to the x and y offsets within the receptive field at the m-th spatial scale i +x,y i The time surface point of the most recent past event relative to the past time of the pixel (+y); the time surface point of the most recent past event relative to the past time of the pixel represents the magnitude of the influence of the most recent past event on the current event; where τ is the time constant of the exponential kernel;
[0056] Using formula Determine the i-th spatiotemporal pulse event e in the spatiotemporal pulse event stream. iWithin the receptive field at the m-th spatial scale, the pixel points with x and y offsets (x i +x,y i The harmonic weighted time surface point (τ) represents the weighted influence of past events on the current event of the pixel; where τ is the time constant of the exponential kernel; the spatiotemporal constant γ determines the weight of the recent event time and the past event time when describing the characteristics of the event flow, and the larger γ is, the more important the time characteristics of the recent event are.
[0057] This embodiment adopts Represents the i-th spatiotemporal pulse event e in the spatiotemporal pulse event stream. i Considering the harmonic weighted time surface of the first b events at the m-th spatial scale, using the formula:
[0058]
[0059] Determine the spatiotemporal pulse event e i Harmonic weighted time surface Each harmonic weighted time surface point within the receptive field of the m-th spatial scale and utilize The harmonic weighted time surface is normalized to obtain the harmonic weighted time surface.
[0060] Wherein, the harmonic weighted time surface is a (2R) m +1)×(2R m A two-dimensional array (+1) describes the i-th spatiotemporal impulse event e in the spatiotemporal impulse event stream. i Its spatial neighborhood radius is R m The information on past events within the receptive field provides dynamic spatiotemporal contextual information about the events surrounding them. In this embodiment, the harmonic weighted time surfaces of each event in the event stream at three different spatial scales are as follows:
[0061] Step S102: Based on the harmonic weighted time surfaces of each spatiotemporal impulse event at three different spatial scales, learn K1 first-layer harmonic weighted time surface prototypes at any spatial scale.
[0062] This step is used to obtain K1 first-layer harmonic weighted time surface prototypes at any spatial scale.
[0063] In a specific embodiment, for any spatial scale, the harmonic weighted time surface of the first K1 spatiotemporal pulse events in the spatiotemporal pulse event stream is selected as K1 initial first-layer harmonic weighted time surface prototypes at this spatial scale, and the harmonic weighted time surface of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream is used to update the K1 first-layer harmonic weighted time surface prototypes at this spatial scale, so as to obtain K1 first-layer harmonic weighted time surface prototypes at any spatial scale.
[0064] According to the K-means clustering algorithm, this application first selects the harmonic weighted time surface of the first K1 spatiotemporal pulse events in the spatiotemporal pulse event stream as K1 initial harmonic weighted time surface prototypes; the K1 harmonic weighted time surface prototypes are the initial cluster centers; where K1 is the number of cluster centers and the number of harmonic weighted time surface prototypes.
[0065] For example, using the first K1 spatiotemporal impulse events, the K1 initial first-layer harmonic weighted time surface prototypes at the m-th spatial scale are obtained as follows:
[0066]
[0067]
[0068] ........
[0070]
[0071] Then, unsupervised online incremental clustering updates are performed on the K1 first-level harmonic weighted time surface prototypes using the harmonic weighted time table of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream. This embodiment utilizes the harmonic weighted time table of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream to update the harmonic weighted time surface prototypes, thereby learning more discriminative harmonic weighted time surface prototypes. Here, the remaining spatiotemporal pulse events refer to the other spatiotemporal pulse events in the input spatiotemporal pulse event stream besides the first K1 spatiotemporal pulse events.
[0072] In a specific embodiment, updating the first-layer harmonic weighted time surface prototypes at the K1 local spatial scales using the harmonic weighted time table of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream includes:
[0073] Calculate the Euclidean distance between the harmonic weighted time surface of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream and the K1 prototypes of the first-layer harmonic weighted time surface at the local spatial scale, and take the prototype of the first-layer harmonic weighted time surface with the shortest Euclidean distance as the prototype of the first-layer harmonic weighted time surface with the highest similarity.
[0074] Then, the first-layer harmonic weighted time surface prototype is updated according to the following formula:
[0075]
[0076] Among them, C k (R m b) indicates that the updated neighborhood radius is R. m The k-th harmonic weighted time surface prototype of the past b events, C′ k (R m b) indicates that the neighborhood radius obtained from the previous update is R. m The k-th harmonic weighted time surface prototype of the past b events, where α is the update amplitude parameter and β is the current spatiotemporal impulse event e. i Harmonic weighted time surface at the m-th spatial scale The cosine distance between it and the first-layer harmonic weighted time surface prototype with the highest similarity.
[0077] Specifically, this embodiment calculates the harmonic weighted time surface of the remaining spatiotemporal impulse events in the spatiotemporal impulse event stream. Calculate the Euclidean distance between each of the K1 first-layer harmonic weighted time surface prototypes, find the harmonic weighted time surface prototype with the shortest Euclidean distance (i.e., the highest similarity), and then update it.
[0078] During the first update, C′ k (R m b) is the initial first-layer harmonic weighted time surface prototype. During subsequent updates, C′ k (R m b) indicates that the neighborhood radius obtained from the previous update is R. m The k-th harmonic weighted time surface prototype of the past b events.
[0079] α represents the update magnitude of the cluster centers. In a specific embodiment, α = 0.01 / (1+p) k (R m b) / 20000), p k (R m b) indicates that it has been assigned to the same harmonic weighted time surface prototype C k (R m The number of events in b).
[0080] β is the cosine distance between the harmonic weighted time surface of the current event and its most similar prototype.
[0081] Thus, K1 first-layer harmonic weighted time surface prototypes at any spatial scale are learned, which are more discriminative. In this embodiment, the harmonic weighted time surface prototypes are a set of discriminative harmonic weighted time surfaces learned from the spatiotemporal impulse event stream, and they are completely consistent with the harmonic weighted time surfaces in terms of expression.
[0082] Step S103: Match the harmonic weighted time surface of each spatiotemporal pulse event in the spatiotemporal pulse event stream at three different spatial scales with K1 first-layer harmonic weighted time surface prototypes at the corresponding spatial scales, find the first-layer harmonic weighted time surface prototype with the highest similarity, use the number corresponding to the first-layer harmonic weighted time surface prototype with the highest similarity as the first-layer polarity of the spatiotemporal pulse event, replace the original light intensity change polarity in the spatiotemporal pulse event with the first-layer polarity, and generate the spatiotemporal feature event stream corresponding to the spatiotemporal pulse event stream.
[0083] In this step, the harmonic weighted time surface prototype remains fixed and is not updated for any new incoming spatiotemporal impulse events. Instead, it is only necessary to find the harmonic weighted time surface prototype that is closest to the harmonic weighted time surface of the new incoming spatiotemporal impulse event based on the Euclidean distance, and modify the polarity p of the incoming spatiotemporal impulse event. i The prototype number of the harmonic weighted time surface with the highest similarity is selected.
[0084] The incoming spatiotemporal impulse event is matched using prototypes to find the harmonic weighted time surface prototype with the highest similarity. Its corresponding index k is stored as an extracted feature in the event polarity. The incoming spatiotemporal impulse event polarity p... i Corresponding to an increase in light intensity of +1 and a decrease in light intensity of -1, the output spatiotemporal characteristic event polarity p′ i The value range is [1, k1], and the structure maintains the same address event expression form as the input spatiotemporal impulse event, i.e., feature i =(x i ,y i ,t i ,p′ i ). Where p′ i The extracted features are recorded and are numerically equal to the number of the closest harmonic weighted time surface prototype found for the i-th spatiotemporal pulse event in the spatiotemporal pulse event stream.
[0085] For example, for a spatiotemporal impulse event e i Harmonic weighted time surface at the m-th spatial scale The first-level harmonic weighted time surface prototype matched is C k (R m b), whose number is k, then its polarity p iReplace it with the number k to obtain the spatiotemporal feature event feat i =(x i ,y i ,t i ,p′ i ).
[0086] That is, for each non-noise spatiotemporal impulse event in the incoming spatiotemporal impulse event stream, features are extracted using the harmonic weighted time surface prototype at three scales, and the spatiotemporal feature event stream F at three scales is output. L1 , where F L1 ={F m |m=1,2,3},F L1 It refers to the set of three spatiotemporal feature event streams output by the first layer, F m It is the spatiotemporal characteristic event stream output at the m-th spatial scale. m represents the m-th spatial scale, and its structure is the same as that of the spatiotemporal impulse event stream E. After extracting the prototype matching features of each non-noise spatiotemporal impulse event in the spatiotemporal impulse event stream at three different scales, the spatiotemporal feature event streams at three different spatial scales are output.
[0087] Step S104: Determine the harmonic weighted time surface of each spatiotemporal feature event in the spatiotemporal feature event stream at the fourth spatial scale, and learn K2 second-layer harmonic weighted time surface prototypes based on the harmonic weighted time surface of each spatiotemporal feature event at the fourth spatial scale.
[0088] This step is used to obtain K2 second-layer harmonic weighted time surface prototypes. First, after the previous step, the original light intensity change polarity in the spatiotemporal pulse event is replaced with the first-layer polarity to obtain the spatiotemporal feature event feat corresponding to each spatiotemporal pulse event. i =(x i ,y i ,t i ,p′ i Then, the harmonic weighted time surface at the fourth spatial scale is determined. R4 is the fourth scale, which is different from the first three scales. For example, the first scale is represented by 3*3, the second scale by 5*5, the third scale by 7*7, and the fourth scale by 9*9.
[0089] The determination of the harmonic weighted event surface at the fourth spatial scale is described in step S101 and will not be repeated here.
[0090] Then, using the same method as in step S102, K2 second-layer harmonic weighted time surface prototypes can be obtained, which will not be described in detail here. It should be noted that for any spatial scale, the number of K1 first-layer harmonic weighted time surface prototypes and the number of K2 second-layer harmonic weighted time surface prototypes in this step can be set separately, for example, K1 is 4, K2 is 8, or other settings are made. This application does not impose any restrictions on this.
[0091] In a specific embodiment, the step of learning K2 second-layer harmonic weighted time surface prototypes based on the harmonic weighted time surface of each spatiotemporal feature event at the fourth spatial scale includes:
[0092] The fourth spatial scale down-weighted time surface of the first K2 spatiotemporal feature events is selected as the K2 initial second-layer harmonic weighted time surface prototypes, and the fourth spatial scale down-weighted time surface of the remaining spatiotemporal impulse events is used to update the K2 second-layer harmonic weighted time surface prototypes.
[0093] It should be noted that, in order to extract more effective spatiotemporal features, when calculating the harmonic weighted time surface, if the difference between the current input event time and the most recent past event time of most pixels in the receptive field is greater than the time constant τ of the exponential kernel, the spatiotemporal pulse event is regarded as noise and discarded. This allows for simultaneous noise reduction while extracting spatiotemporal features, making the extracted spatiotemporal features more robust.
[0094] It is easy to understand that the first layer takes a spatiotemporal impulse event stream as input and outputs three spatiotemporal characteristic event streams at three spatial scales. Since the structural representation of the spatiotemporal characteristic event stream output by the first layer is completely consistent with the spatiotemporal impulse event stream, the output of the first layer can be used as the input of the second layer to obtain K2 second-layer harmonic weighted time surface prototypes.
[0095] Unlike the first layer, the input of the second layer is a spatiotemporal feature event stream. The polarity of the feature event represents the spatiotemporal feature of the event, and its value ranges from 1 to K2. The input of the first layer is the original spatiotemporal pulse event. The polarity of the spatiotemporal pulse event can only be +1 and -1, which represent the increase in light intensity and the decrease in light intensity, respectively.
[0096] The three spatiotemporal feature event streams output from all spatiotemporal impulse event streams serve as input to the second layer. This repeats the harmonic weighted time surface prototype learning operation from the previous layer, learning a harmonic weighted time surface prototype at a fourth spatial scale. The first layer learns harmonic weighted time surface prototypes at three different scales from the input spatiotemporal impulse event streams. The second layer, using the three spatiotemporal feature event streams at different scales extracted from the first layer's features, collectively learns a harmonic weighted time surface prototype at a larger spatial scale.
[0097] Step S105: Obtain the spatiotemporal pulse event stream to be identified output by the event camera, determine the harmonic weighted time surface of each spatiotemporal pulse event at three different spatial scales, match it with K1 first-layer harmonic weighted time surface prototypes at the corresponding spatial scales, take the number corresponding to the first-layer harmonic weighted time surface prototype with the highest similarity as the first-layer polarity, replace the original light intensity change polarity in the spatiotemporal pulse event to be identified with the first-layer polarity, and obtain the first-layer spatiotemporal feature event stream corresponding to the three different spatial scales.
[0098] In the preceding steps, a first-layer harmonic weighted time surface prototype and a second-layer harmonic weighted time surface prototype were obtained, which enabled target recognition of the spatiotemporal pulse event stream output by the camera of the event to be identified.
[0099] During identification, referring to the previous steps, the harmonic weighted time surface of each spatiotemporal impulse event at three different spatial scales is first determined. Then, the K1 first-layer harmonic weighted time surface prototypes obtained earlier are matched, that is, the Euclidean distance is calculated, and the first-layer harmonic weighted time surface prototype with the shortest Euclidean distance, that is, the highest similarity, is found. Then, the number of the first-layer harmonic weighted time surface prototype is used to obtain the first-layer polarity, and the polarity in the original spatiotemporal impulse event is replaced to obtain the corresponding first-layer spatiotemporal feature event.
[0100] It should be noted that, since the operation is performed at three different spatial scales, the first layer outputs a spatiotemporal feature event stream F at the three spatial scales. L1 , where F L1 ={F m |m=1,2,3}.
[0101] Step S106: Determine the harmonic weighted time surface of each spatiotemporal feature event in the spatiotemporal feature event stream corresponding to three different spatial scales at the fourth spatial scale, and match it with K2 second-layer harmonic weighted time surface prototypes respectively, and take the number corresponding to the second-layer harmonic weighted time surface prototype with the highest similarity as the second-layer polarity.
[0102] This step first involves determining the harmonic weighted time surface of each spatiotemporal feature event in the spatiotemporal feature event stream to be identified, using the same method as in the previous steps.
[0103] Then, the harmonic weighted time surface of each spatiotemporal feature event at the fourth spatial scale is matched with K2 second-layer harmonic weighted time surface prototypes to find the second-layer harmonic weighted time surface prototype with the highest similarity. The number corresponding to the second-layer harmonic weighted time surface prototype with the highest similarity is used as the second-layer polarity.
[0104] The spatiotemporal feature events in the spatiotemporal feature event stream at each spatial scale of the second layer input. i =(x i ,y i ,t i ,p′ i Prototype matching is performed on the harmonic weighted time surface to find the prototype of the harmonic weighted time surface with the highest similarity, and its number is p″. i Recording, thus obtaining the spatiotemporal features p″ extracted in the second layer. i That is, the second layer of polarity.
[0105] It should be noted that different spatial scales will yield corresponding second-layer polarities, denoted by p″. mi express.
[0106] Step S107: Replace the original polarity in the spatiotemporal pulse event with the first-layer polarity and the second-layer polarity corresponding to each spatiotemporal pulse event in the spatiotemporal pulse event stream to be identified, so as to obtain the spatiotemporal feature event streams of the second-layer output corresponding to three different spatial scales.
[0107] In this embodiment, the spatiotemporal feature events output by the second layer are based on the spatiotemporal feature events of the first layer, with an additional feature storage variable p″. i That is, the second-layer spatiotemporal feature event feat′ i =(x i ,y i ,t i ,p′ i ,p″ i ), where p″ i p′ represents the second-layer polarity, specifically the number of the most similar harmonic weighted time surface prototype found for the i-th feature event in the second-layer incoming feature event stream. i This represents the first layer of polarity.
[0108] Since the first layer outputs a spatiotemporal feature event stream with three spatial scales, the second layer prototype matching also outputs a second layer spatiotemporal feature event stream with three scales, denoted as F. L2 , where F L2 ={F′ m |m=1,2,3},F L2 It refers to the set of three spatiotemporal feature event streams output by the second layer, F′ m It is the spatiotemporal feature event stream output after feature extraction in the second layer at the m-th spatial scale.
[0109] Step S108: Perform pooling operation on the three spatiotemporal feature event streams output by the second layer, and perform pulse coding on the three pooled spatiotemporal feature event streams to obtain three pulse sequences, which are then input into a single-layer spiking neural network of population coding Tempotron neurons for classification.
[0110] This embodiment reduces the computational load of the learning classification layer through pooling operations and incorporates a refractory period to reduce the number of impulses, thereby accelerating the overall model's learning and classification speed. The pooling window is 4×4. Pooling operations are used to adjust the physical coordinate values of spatiotemporal feature events. After pooling, the window region enters a refractory period, and spatiotemporal impulse events within the same window region will not be emitted. This controls the number of events input to the spiking neural network, reducing the number of feature events in the next stage while maintaining basic spatiotemporal information, thus lowering the computational load.
[0111] In this embodiment, the pulse coding process involves encoding the spatiotemporal feature event stream F′. m spatiotemporal characteristic events feat′ m,i =(x i ,y i ,t i ,p′ m,i ,p″ m,i ) is converted to the structure (addr m,i ,t i The process of a pulse (t). Where, t i The timestamps of characteristic events remain unchanged, mainly by keeping x i and y i Two physical coordinate characteristic variables and p′ m,i and p″ m,i The two extracted spatiotemporal feature variables are converted into a single variable, addr, using pulse coding. m,i This variable corresponds to the index of the input neuron.
[0112] The pulse coding process maps four feature variables to one variable. This requires finding the maximum value of the four feature variables and then determining the range of the addr variable. Since there are three spatiotemporal feature event streams output from feature extraction, the range of the addr variable is [1, size(1) + size(2) + size(3)], where size(m) represents the maximum value that the spatiotemporal feature event stream output from the second layer feature extraction can map at the m-th spatial scale of the first layer. The range of the addr variable is also the index range of the input neuron. The maximum values of the four variables correspond to x: size_x, y: size_y, and p′: K, respectively. m, p″:K4, where size_x and size_y refer to the maximum values that the physical coordinates x and y can take, respectively, which is the input resolution size of the feature event stream, K m K represents the number of clusters of the harmonic weighted time surface prototype at the m-th spatial scale of the first layer, and K4 represents the number of clusters of the harmonic weighted time surface prototype at the fourth spatial scale of the second layer.
[0113] Then, the maximum value that the spatiotemporal feature event stream output by the second feature extraction layer can map at the m-th spatial scale is size(m) = size_x × size_y × K. m ×K4. Since there are three spatiotemporal feature event streams output by feature extraction, the range of input neuron indices that can be mapped by the spatiotemporal feature event stream output by the second layer of feature extraction at the first spatial scale is [1, size(1)], the range of input neuron indices that can be mapped by the spatiotemporal feature event stream output by the second layer of feature extraction at the second spatial scale is [size(1)+1, size(1)+size(2)], and the range of input neuron indices that can be mapped by the spatiotemporal feature event stream output by the second layer of feature extraction at the third spatial scale is [size(1)+size(2)+1, size(1)+size(2)+size(3)]. In particular, size(0) = 0.
[0114] Therefore, for the i-th spatiotemporal feature event feat′ in the spatiotemporal feature event stream output after the second layer of feature extraction at the m-th spatial scale... m,i =(x i ,y i ,t i ,p′ m,i ,p″ m,i The pulse coding process is as follows:
[0115] addr m,i =size(m-1)+(p″) m,i -1)×K m ×size_x×size_y+(p′ m,i -1)×size_x×size_y+(y i -1)×size_x+x i ;
[0116] feat′ m,i =(x i ,y i ,t i ,p′ m,i ,p″ m,i )→feat″ m,i =(addr m,i ,ti );
[0117] Among them, feat′ m,i Let t be the i-th spatiotemporal feature event in the spatiotemporal feature event stream output after feature extraction in the second layer at the m-th spatial scale. i Let x be the timestamp of the i-th spatiotemporal feature event. i and y i These are the two physical coordinate addresses of the i-th spatiotemporal characteristic event, p′ m,i and p″ m,i The two variables, addr, store the features in the i-th spatiotemporal feature event of the spatiotemporal feature event stream output after feature extraction in the second layer at the m-th spatial scale. m,i This refers to the index position of the input neuron corresponding to the i-th spatiotemporal feature event in the spatiotemporal feature event stream output after feature extraction in the second layer at the m-th spatial scale.
[0118] The three spatiotemporal feature event streams are processed by pulse coding to output a pulse sequence, that is, the spatiotemporal feature event stream F′ output after feature extraction at the m-th spatial scale in the second layer. m = After pulse encoding, the output pulse sequence is... Among them, feat″ m,i F′ is the spatiotemporal feature event stream output after feature extraction at the second layer, representing the spatiotemporal feature event stream at the m-th spatial scale. m The index position of the input neuron corresponding to the i-th spatiotemporal feature event.
[0119] Spiking neural networks (SNNs) are a brain-inspired computational model. Unlike traditional neural networks, SNNs transmit and process information by simulating the pulse firing of neurons, making them naturally suited for event-driven algorithms. Since the events output by an event camera are spatiotemporal pulse signals, after spatiotemporal feature extraction, the feature events, based on the address event representation of the input events, add a storage variable, thus also possessing the characteristics of spatiotemporal pulse signals, making them suitable for the learning and classification of SNN models.
[0120] Tempotron is a supervised learning algorithm for spiking neural networks. A single Tempotron neuron can effectively perform binary classification tasks by simply marking the neuron as firing or not firing. However, for multi-class classification tasks, a combination of multiple Tempotron neurons is required.
[0121] The algorithm selects the Leaky Integrate-and-Fire (LIF) model as the neuron model, meaning that each time a pulse is input to a presynaptic neuron, a postsynaptic potential is generated. The neuronal membrane potential is the weighted sum of the postsynaptic potentials generated by all presynaptic neurons:
[0122]
[0123] Among them, w i t represents the synaptic weight value of the i-th presynaptic neuron. i V represents the timestamp of the output pulse of the i-th presynaptic neuron. rest This represents the resting potential of a neuron. When the weighted sum of postsynaptic potentials V(t) exceeds a preset threshold V0... thr At this time, the neuron generates a pulse and resets V(t) to its resting potential V. rest K represents the postsynaptic potential and function:
[0124]
[0125] Where, τ m It is the membrane potential delay constant, τ s It is the synaptic current delay constant, and is generally chosen as τ. m =4τ s V0 is used for normalization, making the maximum value of the kernel function 1.
[0126] In this embodiment, the resting potential V rest If the value is 0, then the formula for calculating the neuron membrane potential is as follows:
[0127]
[0128] This embodiment employs an event-driven membrane potential update process for LIF neurons. Since it's based on an event-driven model, the accumulated membrane potential for the i-th pulse event input is as follows:
[0129] V(t i ) = V m (t i )-V s (t i )
[0130] in,
[0131]
[0132] Wherein, V(t) i ) refers to the accumulated membrane potential after processing the i-th pulse event, Δt = t i -t i-1V refers to the difference in timestamps between the current pulse event and the previous pulse event. m and V s The initial values are all 0. w represents the connection weights between the input neuron and the Tempotron neuron, which can be determined using the input neuron's position index addr. m,i Find the connection weights at the corresponding positions.
[0133] The Tempotron learning rule adjusts connection weights by determining whether a neuron fires a pulse. Assume that firing and not firing pulses represent the positive and negative classes, respectively. When the input pulse sequence belongs to the positive class, the neuron should fire; if it doesn't fire, the corresponding connection weight should be increased to make the membrane potential greater than a threshold and fire a pulse. Conversely, when the input pulse sequence belongs to the negative class, the neuron should not fire; if it does fire, the connection weight should be decreased to make the membrane potential less than a threshold and not fire a pulse. The Tempotron learning rule calculates the peak membrane potential V of the neuron. max With pulse firing threshold V thr The difference is used to update the connection weights. The variable calculation formula for the connection weight value is as follows:
[0134]
[0135] Where, Δw i t represents the update amount of the weight of the i-th connection. max λ represents the moment when the neuron's membrane potential reaches its peak, and λ represents the learning rate.
[0136] In this application, a spiking neural network composed of Tempotron neurons is used as the classifier in the final stage of the system model. The input spiking sequence is an event stream including two variables: timestamp and weight address index, encoded from the spatiotemporal feature extraction output of the previous stage. Considering that individual neurons are easily disturbed, this application adopts a swarm encoding method, associating multiple Tempotron neurons with each class. The output signal is expressed through the joint activity of multiple neurons, thereby enhancing the expressive power of the neurons and their robustness against noise.
[0137] Population coding refers to a classification task of N classes, where each class is associated with M tempotron neurons. Therefore, the output layer has N groups of neurons and N×M tempotron neurons. Normally, for an N-class classification task, the output layer can only have N tempotron neurons, and each tempotron neuron has only two states: firing or not firing. The class corresponding to the firing tempotron neuron is the prediction result. However, the network cannot handle situations where multiple tempotron neurons fire. Prediction under population coding, on the other hand, becomes the prediction based on the tempotron neuron in the corresponding output neuron group firing the most (i.e., the one with the most firings). During the learning phase of the spiking neural network, only the tempotron neurons in the corresponding class group need to fire, while the tempotron neurons in other groups do not need to fire. In the classification prediction phase of the spiking neural network, a majority voting method is used; that is, the group of tempotron neurons that fires the most pulses is predicted as the corresponding class, thus achieving the recognition of the target object. In this embodiment, an input pulse event stream will generate three pulse sequences at three spatial scales, which will be output to the network for classification. Therefore, the total voting result of these three pulse sequences is the prediction result of the input pulse event stream. The spiking neural network in this model uses an event-driven approach for simulation, resulting in higher recognition efficiency.
[0138] This application presents a neuromorphic visual target recognition method that fully utilizes the precise temporal information of spatiotemporal pulse events in an event-driven manner. The spiking neural network (SNN) event-driven algorithm extracts and retains as much precise temporal information as possible when processing spatiotemporal pulse events output by an event camera. The event camera asynchronously encodes visual information as a spatiotemporal pulse event stream, with each pixel immediately outputting a pulse when light intensity changes, thus enabling real-time response to changes in the visual scene. The SNN event-driven algorithm effectively utilizes these spatiotemporal pulse events without converting them into traditional fixed-frame-rate images. In contrast, traditional frame-driven algorithms require converting the event stream into fixed-frame-rate images for preprocessing and classification. During this process, a large amount of event information is compressed into the fixed-frame-rate image, resulting in the loss of the original event's precise temporal information. This frame-to-frame preprocessing leads to information loss and redundancy, potentially reducing the algorithm's accuracy and efficiency.
[0139] This application utilizes an event-driven approach to directly process spatiotemporal pulse events, enabling better utilization of the asynchronous operation of the event camera and preserving the precise temporal information of the original events as much as possible. This allows the algorithm to adapt more flexibly and efficiently to complex visual scenes, achieving better performance and accuracy in various visual perception tasks. Therefore, the event-driven algorithm based on spiking neural networks has significant advantages in processing spatiotemporal pulse events output by event cameras.
[0140] The technical solution of this application achieves high recognition accuracy on neuromorphic action datasets: it can effectively extract spatiotemporal related features from spatiotemporal pulse event streams and effectively extract the dynamic activity of past events within the receptive field of each incoming event, thereby capturing scene dynamic information and achieving high recognition accuracy on neuromorphic action datasets. Table 1 compares the recognition accuracy of this application with existing work on neuromorphic action datasets.
[0141] Table 1
[0142]
[0143] Comparative Example 1: Synaptic Plasticity Dynamics for Deep Continuous Local Learning (DECOLLE), published in Frontiers Neuroscience, 2020, 14:424, authors: Kaiser J, Mostafa H and Neftci E.
[0144] Comparative Example 2: An event-driven categorization model for AER image sensors using multispike encoding and learning [J], published in IEEE transaction-s on neural networks and learning systems, 2019, 31(9): 3649-3657, author: Xiao R, Tang H, Ma Y, etal.
[0145] Comparative Example 3: Effective AER object classification using segmented probability-maximization learning in spiking neural networks [C], published in Proceedings of the AAAI conference on artificial intelligence, 2020, 34(02):1308-1315, authors: Liu Q, Ruan H, Xing D, et al.
[0146] Comparative Example 4: Event-based Action Recognition Using Motion Information and Spiking Neural Networks [C], published in IJCAI, 2021:1743-1749, authors: Liu Q, Xing D, Tang H, et al.
[0147] This application achieves high recognition accuracy on neuromorphic digital datasets: Neuromorphic digital datasets do not have high requirements for extracting spatiotemporal related features from spatiotemporal impulse event streams, but have higher requirements for extracting edge direction features. The high recognition accuracy of this method on neuromorphic digital datasets indicates that the harmonic weighted temporal surface prototype templates learned in clustering also possess a certain degree of scene direction feature differentiation. Table 2 compares the recognition accuracy of this application with existing works on neuromorphic digital datasets:
[0148] Table 2
[0149]
[0150] Comparative Example 1: Feedforward categorization on AER motion events using cortex-like features in a spiking neural network[J], published in IEEE transactions on neural networks and learning systems, 2014, 26(9):1963-1978, authors: Zhao B, Ding R, Chen S, et al.
[0151] Comparative Example 2: HFirst: A temporal approach to object recognition[J], published in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015, 37(10): 2028-2040, authors: Orchard G, Meyer C, Etienne-Cummings R, et al.
[0152] Comparative Example 3: Hots: a hierarchy of event-based time-surfaces for pattern recognition[J], published in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2016, 39(7): 1346-1359, authors: Lagorce X, Orchard G, Galluppi F, et al.
[0153] Comparative Example 4: Unsupervised aer object recognition based on multiscalespatio-temporal features and spiking neurons[J], published in IEEE Transactions on Neural Networks and Learning Systems, 2020, 31(12): 5300-5311, authors: Liu Q, Pan G, Ruan H, et al.
[0154] Comparative Example 5: An event-driven categorization model for AER image sensor using multispike encoding and learning[J], published in IEEE Transactions on Neural Networks and Learning Systems, 2019, 31(9):3649-3657, authors: Xiao R, Tang H, Ma Y, et al.
[0155] Comparative Example 6: An event-driven object recognition model using activated connected domain detection [C], published in IEEE Symposium Series on Computational Intelligence (SSCI), 2020:3049-3056, authors: Tang T, Jiang R, Yan R, et al.
[0156] Comparative Example 7: Effective AER object classification using segmented probability-maximization learning in spiking neural networks [C], published in Proceedings of the AAAI conference on artificial intelligence. 2020, 34(02): 1308-1315, authors: Liu Q, Ruan H, Xing D, et al.
[0157] In another embodiment of this application, a neuromorphic visual target recognition device is also provided, including a processor and a memory storing a plurality of computer instructions, which, when executed by the processor, implement the steps of the above method.
[0158] Specific limitations regarding the neuromorphic visual target recognition device can be found in the limitations of the neuromorphic visual target recognition method described above, and will not be repeated here. The aforementioned neuromorphic visual target recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. It can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0159] The memory and processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The memory stores a computer program that can run on the processor, which implements the network topology layout method in this embodiment of the invention by running the computer program stored in the memory.
[0160] The memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory stores the program, and the processor executes the program upon receiving an execution instruction.
[0161] The processor may be an integrated circuit chip with data processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.
[0162] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A neuromorphic visual target recognition method, characterized in that, The neuromorphic visual target recognition method includes: Acquire the spatiotemporal pulse event stream asynchronously output by the event camera, and determine the harmonic weighted time surface of each spatiotemporal pulse event in the spatiotemporal pulse event stream at three different spatial scales; Based on the harmonic weighted time surface of each spatiotemporal impulse event at three different spatial scales, K1 first-level harmonic weighted time surface prototypes at any spatial scale are learned. Each spatiotemporal pulse event in the spatiotemporal pulse event stream is matched with the harmonic weighted time surface at three different spatial scales and K1 first-layer harmonic weighted time surface prototypes at the corresponding spatial scales. The first-layer harmonic weighted time surface prototype with the highest similarity is found. The number corresponding to the first-layer harmonic weighted time surface prototype with the highest similarity is used as the first-layer polarity of the spatiotemporal pulse event. The first-layer polarity is used to replace the original light intensity change polarity in the spatiotemporal pulse event to generate the spatiotemporal feature event stream corresponding to the spatiotemporal pulse event stream. Determine the harmonic weighted time surface of each spatiotemporal feature event in the spatiotemporal feature event stream at the fourth spatial scale, and learn K2 second-layer harmonic weighted time surface prototypes based on the harmonic weighted time surface of each spatiotemporal feature event at the fourth spatial scale. The event stream of spatiotemporal pulse events to be identified is obtained from the event camera. The harmonic weighted time surface of each spatiotemporal pulse event at three different spatial scales is determined. It is matched with K1 prototypes of the first-layer harmonic weighted time surface at the corresponding spatial scale. The number corresponding to the prototype of the first-layer harmonic weighted time surface with the highest similarity is used as the first-layer polarity. The first-layer polarity is used to replace the original light intensity change polarity in the spatiotemporal pulse event to be identified, so as to obtain the first-layer spatiotemporal feature event stream corresponding to the three different spatial scales. Determine the harmonic weighted time surface of each spatiotemporal feature event in the spatiotemporal feature event stream corresponding to three different spatial scales at the fourth spatial scale, and match it with K2 second-layer harmonic weighted time surface prototypes. The number corresponding to the second-layer harmonic weighted time surface prototype with the highest similarity is used as the second-layer polarity. The original polarity in the spatiotemporal pulse event is replaced by the first-layer polarity and the second-layer polarity corresponding to each spatiotemporal pulse event in the spatiotemporal pulse event stream to be identified, so as to obtain the spatiotemporal feature event streams of the second-layer output corresponding to three different spatial scales. Pooling is performed on the three spatiotemporal feature event streams output from the second layer. Pulse coding is then performed on the three pooled spatiotemporal feature event streams to obtain three pulse sequences, which are then input into a single-layer spiking neural network with population coding Tempotron neurons for classification.
2. The neuromorphic visual target recognition method as described in claim 1, characterized in that, The spatiotemporal pulse event stream is described using the address event representation protocol as follows: Where e i Let x represent the i-th spatiotemporal pulse event in the spatiotemporal pulse event stream. i and y i Let t represent the pixel coordinates of the i-th spatiotemporal impulse event. i p represents the timestamp of the i-th spatiotemporal impulse event. i This indicates the polarity of the light intensity change at the pixel corresponding to the i-th spatiotemporal pulse event, where +1 / -1 represent the "ON event" of increased light intensity and the "OFF event" of decreased light intensity, respectively. This represents the number of spatiotemporal pulse events in the spatiotemporal pulse event stream.
3. The neuromorphic visual target recognition method as described in claim 2, characterized in that, The step of determining the harmonic weighted time surface of each spatiotemporal pulse event in the spatiotemporal pulse event stream at three different spatial scales includes: The spatial scale is determined by the neighborhood radius, R m Let m be the neighborhood radius of the m-th spatial scale, and let the spatial scale be (2R). m +1)×(2R m +1); Using t(e) i ,R m (b, x, y) represents the i-th spatiotemporal impulse event e i The time of b most recent past events for pixels with x and y offsets within the receptive field at the m-th spatial scale, where x, y ∈ [-R]. m ,R m ]; Using formula Determine the i-th spatiotemporal impulse event e i The harmonic mean relative past time of b most recent past events for pixels with x and y offsets within the receptive field at the m-th spatial scale; where relative past time refers to t. i -t j That is, Δt i,j That is, the difference between the time of the current event and the time of past events, which is the harmonic average of b relative past times; Using formula Determine the i-th spatiotemporal impulse event e i Within the receptive field of the m-th spatial scale, the harmonic average of b most recent past events relative to the past time is the time surface point of the exponential kernel, where τ is the time constant of the exponential kernel. use Determine the i-th spatiotemporal impulse event e i The relative past time of the most recent event for a pixel with x and y offsets within the receptive field at the m-th spatial scale; where t i -t last t represents the difference between the time of the current event and the time of the most recent past event. last It is the i-th spatiotemporal impulse event e i The time of the most recent past event for a pixel with x and y offsets within the receptive field of the m-th spatial scale; Using formula Determine the i-th spatiotemporal impulse event e i The time surface point of the relative past time of the most recent past event of a pixel with respect to x and y offsets within the receptive field of the m-th spatial scale; Using formula Determine the i-th spatiotemporal impulse event e i The harmonic weighted temporal surface point of the pixels with x and y offsets within the receptive field of the m-th spatial scale, where γ is the spatiotemporal constant; use Represents the i-th spatiotemporal impulse event e i For the harmonic weighted time surface at the m-th spatial scale, using the formula: Determine the spatiotemporal pulse event e i Harmonic weighted time surface Each harmonic weighted time surface point within the receptive field of the m-th spatial scale and utilize The harmonic weighted time surface is normalized to obtain the harmonic weighted time surface.
4. The neuromorphic visual target recognition method as described in claim 2, characterized in that, Based on the harmonic weighted time surfaces of each spatiotemporal impulse event at three different spatial scales, K1 prototypes of the first-level harmonic weighted time surfaces at any spatial scale are learned, including: For any spatial scale, the harmonic weighted time surface of the first K1 spatiotemporal pulse events in the spatiotemporal pulse event stream is selected as K1 initial first-layer harmonic weighted time surface prototypes at this spatial scale, and the harmonic weighted time surface of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream is used to update the K1 first-layer harmonic weighted time surface prototypes at this spatial scale, so as to obtain K1 first-layer harmonic weighted time surface prototypes at any spatial scale.
5. The neuromorphic visual target recognition method as described in claim 4, characterized in that, The process of updating the first-level harmonic weighted time surface prototypes at the K1 local spatial scales using the harmonic weighted time table of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream includes: Calculate the Euclidean distance between the harmonic weighted time surface of the remaining spatiotemporal pulse events in the spatiotemporal pulse event stream and the K1 prototypes of the first-layer harmonic weighted time surface at the local spatial scale, and take the prototype of the first-layer harmonic weighted time surface with the shortest Euclidean distance as the prototype of the first-layer harmonic weighted time surface with the highest similarity. Then, the first-layer harmonic weighted time surface prototype is updated according to the following formula: Among them, C k (R m b) indicates that the updated neighborhood radius is R. m The k-th harmonic weighted time surface prototype of the past b events, C′ k (R m b) indicates that the neighborhood radius obtained from the previous update is R. m The k-th harmonic weighted time surface prototype of the past b events, where α is the update amplitude parameter and β is the current spatiotemporal impulse event e. i Harmonic weighted time surface at the m-th spatial scale The cosine distance between it and the first-layer harmonic weighted time surface prototype with the highest similarity.
6. A neuromorphic visual target recognition device, comprising a processor and a memory storing a plurality of computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 5.