Processing unit, related method and data center
By introducing processing units of sensing layer, spatiotemporal computing core array, connection layer and decoding layer in deep learning, the problem of high computational volume in the prior art when processing spatiotemporal-related tasks is solved, and efficient spatiotemporal event recognition is achieved.
Patent Information
- Application Number
- CN202010812969.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-08-13
AI Technical Summary
Existing deep learning technologies have a lot of computation when processing tasks related to time and space, making it difficult to effectively handle energy or computing resources when energy or computing resources are limited.
A processing unit is proposed, including a sensing layer, a spatiotemporal computing core array, a connection layer and a decoding layer. By encoding the space-time event signal to be identified into a pulse sequence, and using the space-time computing core array to be used to double accumulation of time and space, generate space-time accumulation result pulses, and finally classify events through the decoding layer.
This method greatly reduces the amount of calculations for processing space-time related tasks, improves processing efficiency, and is versatile, and is suitable for various application scenarios.
Smart Images

Figure CN114077856B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of deep learning, and more particularly, to a processing unit, a related method, and a data center. Background Art
[0002] Today, deep learning has had a significant impact on a wide range of industrial applications. However, deep learning involves a large amount of computation when dealing with spatio-temporal related tasks. Spatio-temporal related tasks refer to tasks that determine and classify the behavior of an object based on the comprehensive performance of the object in terms of time and space. For example, based on the path traveled by a car over a period of time after the current time, determining the change in the driving direction of the car. This is different from static face recognition, etc. Face recognition only needs to recognize a face based on a static picture and does not need to consider the changes in the face in space and time.
[0003] Current mainstream deep learning is not ideal when dealing with spatio-temporal related tasks (such as event recognition) under limited energy or computing resources. Therefore, spatio-temporal neural networks have emerged. Currently, three main types of spatio-temporal neural networks (SNNs) have been proposed: 1) SNNs based on backpropagation (BP); 2) SNNs based on convolutional neural networks (CNNs); 3) SNNs based on biological learning behaviors. For BP-based SNNs, although the inference accuracy is very high, the BP mechanism determines that the computational amount is large. For CNN-based SNNs, since CNNs are not suitable for recognizing postures and events, the computational amount is still large. As for SNNs based on biological learning behaviors, they utilize the principles of neural plasticity or spike-timing-dependent plasticity (STDP) learning, and have excellent training performance, but their application scope in industry is limited. Summary of the Invention
[0004] In view of this, the present disclosure aims to propose a processing unit and a related method thereof, which can reduce the computational amount when dealing with spatio-temporal related tasks and have generality.
[0005] To achieve this purpose, according to one aspect of the present disclosure, the present disclosure provides a processing unit, including:
[0006] A sensing layer, including a plurality of sensing units, respectively configured to encode a spatio-temporal event signal to be recognized into a pulse sequence, where the pulse sequence reflects the time element and the space element of the spatio-temporal event to be recognized;
[0007] A spatio-temporal computing core array, including a plurality of spatio-temporal computing cores, respectively connected to at least a part of the plurality of sensing units, and generating a spatio-temporal cumulative result pulse according to the pulse sequence output by the at least a part of the sensing units, where the spatio-temporal cumulative result pulse reflects the spatio-temporal cumulative result of the pulse sequence;
[0008] A connection layer;
[0009] The decoding layer includes a plurality of decoding units, each decoding unit corresponding to an event classification respectively. The decoding units are connected to a corresponding part of the plurality of spatio-temporal computing cores through the connection layer, and generate output values according to the spatio-temporal cumulative result pulses output by the corresponding part of the spatio-temporal computing cores, and use the event classification corresponding to the decoding unit with the largest output value as the recognized event classification.
[0010] Optionally, the spatio-temporal computing core array includes spatial neurons and temporal neurons. Among them, the spatial neurons are connected to the at least part of the sensing units, generate a spatial cumulative result pulse sequence according to the pulse sequence output by the at least part of the sensing units, and the spatial cumulative result pulse sequence reflects the spatial cumulative result of the pulse sequence; the temporal neurons generate the spatio-temporal cumulative result pulses according to the spatial cumulative result pulse sequence generated by the spatial neurons, and the spatio-temporal cumulative result pulses reflect the temporal cumulative result of the spatial cumulative result pulse sequence.
[0011] Optionally, the connection relationship from the decoding unit to the spatio-temporal computing core in the connection layer is pre-trained in the following way:
[0012] Connect the plurality of decoding units to a predetermined number of spatio-temporal computing cores respectively;
[0013] Input a set of spatio-temporal event signal samples each containing at least one spatio-temporal event signal sample of each event classification into the processing unit. For each spatio-temporal event signal sample, determine the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to each decoding unit. If the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of this spatio-temporal event signal sample is less than the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to other decoding units, adjust the connection relationship between the decoding unit corresponding to the event classification of this spatio-temporal event signal sample and the spatio-temporal computing cores, and the connection relationship between this other decoding unit and the spatio-temporal computing cores, so that the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of this spatio-temporal event signal sample is greater than the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to this other decoding unit.
[0014] Optionally, the step of connecting the plurality of decoding units to a predetermined number of spatio-temporal computing cores respectively includes:
[0015] For each decoding unit, select a plurality of spatio-temporal event signal samples corresponding to the event classification of this decoding unit and input them into the processing unit respectively, and determine the number of times of outputting valid pulses for the plurality of spatio-temporal event signal samples for each spatio-temporal computing core;
[0016] Connect the predetermined number of spatio-temporal computing cores with the highest to lowest number of output valid pulses to the decoding unit.
[0017] Optionally, encoding the spatio-temporal event signal to be recognized into a pulse sequence includes:
[0018] Taking the spatio-temporal event signal to be recognized as a function of position and time, and for a specific position, determining the difference between the function value at the specific position and the current time and the function value at the time point of the previous cycle at the specific position and the current time;
[0019] According to the comparison result between the difference and a predetermined difference threshold, determining whether the pulse at a specific position is a valid pulse or an invalid pulse;
[0020] Combining the determined pulses at each position into the pulse sequence.
[0021] Optionally, generating a spatial accumulation result pulse sequence according to the pulse sequence output by at least a part of the sensing units includes:
[0022] Determining the number of pulses output by at least a part of the sensing units as valid pulses at a specific time point;
[0023] If the number is greater than a first threshold, outputting a valid pulse at the specific time point, otherwise outputting an invalid pulse;
[0024] Combining the pulses output at each time point into the spatial accumulation result pulse sequence.
[0025] Optionally, generating the spatio-temporal accumulation result pulse according to the spatial accumulation result pulse sequence generated by the spatial neuron includes:
[0026] Generating an integral within a time window for the spatial accumulation result pulse sequence generated by the spatial neuron;
[0027] If the integral value is greater than a second threshold, outputting a valid spatio-temporal accumulation result pulse, otherwise outputting an invalid spatio-temporal accumulation result pulse.
[0028] Optionally, generating an output value according to the spatio-temporal accumulation result pulse output by the corresponding part of the spatio-temporal computing cores includes:
[0029] Taking the number of spatio-temporal accumulation result pulses output by the corresponding part of the spatio-temporal computing cores as valid pulses as the output value.
[0030] According to one aspect of the present disclosure, there is provided a spatio-temporal event recognition method, including:
[0031] Encoding the spatio-temporal event signal to be recognized into a pulse sequence, where the pulse sequence reflects the time element and the space element of the spatio-temporal event to be recognized;
[0032] Generate a spatio-temporal cumulative result pulse according to at least a part of the pulse sequence, where the spatio-temporal cumulative result pulse reflects the spatio-temporal cumulative result of the at least a part of the pulse sequence;
[0033] Let multiple decoding units respectively generate output values according to the spatio-temporal cumulative result pulses output by a part of the spatio-temporal calculation kernels connected to the decoding unit, and classify the event corresponding to the decoding unit with the largest output value as the recognized event classification.
[0034] Optionally, the generating a spatio-temporal cumulative result pulse according to at least a part of the pulse sequence includes:
[0035] Generate a spatial cumulative result pulse sequence according to at least a part of the pulse sequence, where the spatial cumulative result pulse sequence reflects the spatial cumulative result of the at least a part of the pulse sequence;
[0036] Generate the spatio-temporal cumulative result pulse according to the spatial cumulative result pulse sequence, where the spatio-temporal cumulative result pulse reflects the temporal cumulative result of the spatial cumulative result pulse sequence.
[0037] Optionally, the encoding the spatio-temporal event signal to be recognized into a pulse sequence includes:
[0038] Regarding the spatio-temporal event signal to be recognized as a function of position and time, for a specific position, determine the difference between the function value at the specific position and the current time and the function value at the time point of the previous cycle of the specific position and the current time;
[0039] According to the comparison result between the difference and a predetermined difference threshold, determine whether the pulse at the specific position is a valid pulse or an invalid pulse;
[0040] Synthesize the determined pulses at each position into the pulse sequence.
[0041] Optionally, the generating a spatial cumulative result pulse sequence according to at least a part of the pulse sequence includes:
[0042] Determine the number of pulses determined to be valid pulses according to at least a part of the pulse sequence at a specific time point;
[0043] If the number is greater than a first threshold, output a valid pulse at the specific time point, otherwise output an invalid pulse;
[0044] Synthesize the pulses output at each time point into the spatial cumulative result pulse sequence.
[0045] Optionally, the generating the spatio-temporal cumulative result pulse according to the spatial cumulative result pulse sequence includes:
[0046] Integrate the spatial accumulation result pulse sequence over a time window;
[0047] If the integral value is greater than a second threshold, output a valid spatio-temporal accumulation result pulse; otherwise, output an invalid spatio-temporal accumulation result pulse.
[0048] Optionally, generating an output value based on the spatio-temporal accumulation result pulses output by a part of the spatio-temporal calculation cores connected to the decoding unit includes: using the number of valid pulses among the spatio-temporal accumulation result pulses output by a part of the spatio-temporal calculation cores connected to the decoding unit as the output value.
[0049] According to one aspect of the present disclosure, there is provided a configuration method for a processing unit, the processing unit including a sensing layer, a spatio-temporal calculation core array, a connection layer, and a decoding layer, the spatio-temporal calculation core array including a plurality of spatio-temporal calculation cores, the decoding layer including a plurality of decoding units, the connection layer being configured to connect the plurality of decoding units to the plurality of spatio-temporal calculation cores, the method including:
[0050] Connect the plurality of decoding units to a predetermined number of spatio-temporal calculation cores among the plurality of spatio-temporal calculation cores respectively;
[0051] Input a set of spatio-temporal event signal samples each including at least one spatio-temporal event signal sample of each event classification into the processing unit. For each spatio-temporal event signal sample, determine the number of valid pulses output by the predetermined number of spatio-temporal calculation cores connected to each decoding unit. If the number of valid pulses output by the predetermined number of spatio-temporal calculation cores connected to the decoding unit corresponding to the event classification of the spatio-temporal event signal sample is less than the number of valid pulses output by the predetermined number of spatio-temporal calculation cores connected to other decoding units, adjust the connection relationship between the decoding unit corresponding to the event classification of the spatio-temporal event signal sample and the spatio-temporal calculation cores, and the connection relationship between the other decoding unit and the spatio-temporal calculation cores, such that the number of valid pulses output by the predetermined number of spatio-temporal calculation cores connected to the decoding unit corresponding to the event classification of the spatio-temporal event signal sample is greater than the number of valid pulses output by the predetermined number of spatio-temporal calculation cores connected to the other decoding units.
[0052] Optionally, the connecting the plurality of decoding units to a predetermined number of spatio-temporal calculation cores among the plurality of spatio-temporal calculation cores respectively includes:
[0053] For each decoding unit, select a plurality of spatio-temporal event signal samples corresponding to the event classification of the decoding unit and input them into the processing unit respectively, and determine the number of times of outputting valid pulses for each spatio-temporal calculation core with respect to the plurality of spatio-temporal event signal samples;
[0054] Connect the spatio-temporal calculation cores with the top predetermined number of times of outputting valid pulses in descending order to the decoding unit.
[0055] According to one aspect of the present disclosure, a data center is further provided, which includes a plurality of servers. Computer-readable codes are distributedly stored on the plurality of servers. When the computer-readable codes are executed by a processor on a corresponding server, the above-mentioned spatio-temporal event recognition method is implemented.
[0056] The embodiment of the present disclosure proposes a unique processing unit structure, which includes a sensing layer, a spatio-temporal computing core array, a connection layer, and a decoding layer. The sensing layer encodes the spatio-temporal event signal to be recognized into a pulse sequence, which reflects both the time element and the space element of the spatio-temporal event to be recognized. The spatio-temporal computing core array performs double accumulation of time and space on the pulse sequence output by the sensing layer to generate a spatio-temporal accumulation result pulse, which reflects the characteristics accumulated in time and space of the spatio-temporal event to be recognized. The connection layer connects the decoding units in the decoding layer to the spatio-temporal computing cores in the spatio-temporal computing core array. This connection relationship is pre-trained. Then, each decoding unit in the decoding layer generates an output value based on the spatio-temporal accumulation result pulse output by the spatio-temporal computing core it is connected to, and event classification is performed based on the output value. In the whole process, only by accumulating pulses, that is, performing addition and integration operations, different from the existing BP and CNN-based methods, the amount of calculation required for processing spatio-temporal related tasks is greatly reduced. At the same time, there is no limitation on the application scope. This neural network structure is particularly suitable for intelligent applications based on edge and terminal devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Through the description of the embodiments of the present disclosure with reference to the following drawings, the above-mentioned and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0058] Figure 1 is the overall architecture diagram of the data center applied in the embodiment of the present disclosure;
[0059] Figure 2 is the structural diagram of the processing unit according to an embodiment of the present disclosure;
[0060] Figure 3A -I is a schematic diagram of the application interface when the processing unit is configured and actually used, where Figure 3A -G shows the interface during configuration, Figure 3H -I shows the interface during actual use.
[0061] Figure 4 shows a schematic diagram of three spatio-temporal events abstracted from the real world.
[0062] Figure 5 is Figure 2Schematic diagram of the pulse sequences output by each node of the processing unit shown
[0063] Figure 6A -Figure F shows an example process diagram of the connection relationship between each spatio-temporal computing kernel and each decoding unit in the training connection layer
[0064] Figure 7 Flowchart showing a spatio-temporal event recognition method according to an embodiment of the present disclosure
[0065] Figure 8 Flowchart showing a processing unit configuration method according to an embodiment of the present disclosure Detailed implementation manners
[0066] The present disclosure is described below based on embodiments, but the present disclosure is not limited to these embodiments. In the following detailed description of the present disclosure, some specific details are described in detail. Those skilled in the art can fully understand the present disclosure without the description of these details. In order to avoid obscuring the essence of the present disclosure, well-known methods, processes, and flows are not described in detail. Additionally, the drawings are not necessarily drawn to scale
[0067] Glossary of Terms
[0068] The following terms are used herein
[0069] Processing unit: A unit for performing traditional data processing, and for cases where the efficiency is not high in some special-purpose fields (such as processing images, performing various operations of deep learning models, etc.), a unit designed to improve the data processing speed in these special-purpose fields, which includes a central processing unit (CPU) for performing traditional processing, a graphics processing unit (GPU) dedicated to performing image processing, a neural network processing unit (NPU) dedicated to the arithmetic processing of deep learning models, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and other various units
[0070] Spatio-temporal neural network: The world is essentially structured. It includes components that interact with each other in space and time, forming a spatio-temporal combination. The interaction between humans and the environment in the real world is spatio-temporal. For example, when cooking, humans interact with multiple objects in both space and time. Similarly, a person's body (arms, legs, etc.) has separate functions but cooperate with each other in actual actions. Therefore, a high-order spatio-temporal structure is quite important for many applications. A spatio-temporal neural network is a deep learning model infused with high-order information, enabling the deep learning model to fully learn both the time element and the space element in an event
[0071] Spatio-temporal event: It refers to the behavioral manifestation of an object in the real world in time and space. Any object in the real world can be a target, such as a vehicle and a human hand. A target is located at a position in space at a time point on the time axis and may be located at other positions in space at other time points. The change in position at different time points forms the behavioral manifestation in time and space, that is, a spatio-temporal event. For example, if a vehicle changes its direction from north to west as time passes, then traveling from north to west is a spatio-temporal event of the vehicle. A human hand makes various gestures, and the relative positions of the parts of the hand change at each time point for each gesture. Therefore, a gesture is a behavioral manifestation of the hand in time and space and is a spatio-temporal event.
[0072] Spatio-temporal event signal: A signal that represents an abstract spatio-temporal event. Spatio-temporal events are abstract and must be represented as spatio-temporal event signals for data processing. For example, for the spatio-temporal event of a vehicle traveling from north to west, each pixel in the trajectory photo of the vehicle taken at each sampling time point is the spatio-temporal event signal transformed from this spatio-temporal event. For the spatio-temporal event of a human gesture, each pixel in the photo of the human gesture taken at each sampling time point is the spatio-temporal event signal transformed from this spatio-temporal event.
[0073] Encoding: The act of converting a signal into a code according to a predetermined rule. In the embodiments of the present disclosure, the spatio-temporal event signal is converted into a pulse sequence according to the encoding rule, and this pulse sequence reflects the time and space elements of the spatio-temporal event.
[0074] Pulse sequence: A sequence composed of binary pulses, that is, a sequence composed of a series of pulses of "1" or "0", where "1" is a valid pulse and "0" is an invalid pulse.
[0075] Time and space elements: Since a spatio-temporal event refers to the behavioral manifestation of an object in the real world in time and space, therefore, a spatio-temporal event has manifestations in both time and space. The manifestation in time is called a time element, and the manifestation in space is called a space element. For the spatio-temporal event of a vehicle traveling from north to west, in the trajectory photo of the vehicle taken at a sampling time point, some adjacent pixels in certain areas will exhibit certain common characteristics, that is, the space element. Among the trajectory photos taken at multiple sampling time points, the pixels at the same position in the photos also have common characteristics between adjacent photos (or frames), that is, the time element.
[0076] Spatio-temporal computing core array: An array composed of spatio-temporal computing cores. In the array, there are several spatio-temporal computing cores in each row and several spatio-temporal computing cores in each column.
[0077] Space-time calculation core: A unit that processes the pulse sequence of space-time events in terms of time and space to obtain the cumulative result of the pulse sequence in time and space. Since the pulse sequence reflecting the time and space elements of space-time events is too long, in order to extract effective information and avoid occupying too much storage space, it is necessary to accumulate the pulse sequence in time and space, so as to obtain some characteristics common to the pulse sequence in time and space. The unit that executes this process is called the space-time calculation core. The space-time calculation core includes space neurons and time neurons.
[0078] Space neuron: A unit that accumulates the pulse sequence reflecting the time and space elements of the space-time event in space. Pulses generated at the same time point and adjacent spatial positions often exhibit some similar characteristics. Therefore, accumulating these pulse sequences in space can reflect these similar characteristics. One way of processing by space neurons can be to directly accumulate the pulses generated at the same time point and adjacent spatial positions, and the other is to accumulate first. If the accumulation reaches a predetermined threshold, an effective pulse "1" is generated, otherwise an invalid pulse "0" is generated. The latter is beneficial to saving storage space, and the comparison with the predetermined threshold can also generally reflect the more or less of the effective pulses generated near a certain point in space, and will not cause too much damage to the amount of information.
[0079] Space cumulative result pulse sequence: The pulse sequence generated by the space neuron and serving as the cumulative result of the space neuron.
[0080] Time neuron: A unit that accumulates the space cumulative result pulse sequence generated by the space neuron in time. Pulses generated at the same spatial position at different time points often exhibit some similar characteristics. Therefore, accumulating these pulse sequences in time can reflect these similar characteristics. One way of processing by time neurons can be to directly accumulate the space cumulative result pulse sequence generated by the space neuron, and the other is to accumulate first. If the accumulation reaches a predetermined threshold, an effective pulse "1" is generated, otherwise an invalid pulse "0" is generated. The latter is beneficial to saving storage space, and the comparison with the predetermined threshold can also generally reflect the more or less of the effective pulses generated near a certain time point, and will not cause too much damage to the amount of information.
[0081] Space-time cumulative result pulse: The pulse generated by the time neuron and serving as the cumulative result of the time neuron.
[0082] Event classification: The category to which the above spatio-temporal event is assigned. Usually, tasks related to spatio-temporal processed by spatio-temporal neural networks are tasks of classifying events based on the comprehensive performance of events in terms of time and space. Therefore, the processing result of a spatio-temporal neural network usually is to classify an event into a category. For example, based on the trajectory photos of each captured car, the driving event of the car is determined to be driving from north to west.
[0083] Connection layer: A layer that connects the decoding units of the decoding layer to a corresponding part of the spatio-temporal computing cores. By configuring the connection layer, the connection relationship between the decoding units and the spatio-temporal computing cores can be configured.
[0084] Application Scenarios of Embodiments of the Present Disclosure
[0085] The present disclosure can be applied to urban brains, autonomous driving, speech-vision fusion systems, smart homes, chip-based Internet servers (ISPs), network interconnection protocols (IPs), etc. Below, taking urban brains and autonomous driving as examples, the scenarios to which the embodiments of the present disclosure are applied will be described. Then, the scenarios of speech-vision fusion systems and smart homes will be briefly introduced.
[0086] An urban brain is a digital interface created for urban life. Citizens rely on it to feel the pulse of the city, experience the temperature of the city, and enjoy urban services. Urban managers use it to allocate public resources, make scientific decisions, and improve governance efficiency. The core of the urban brain is the data center. Each edge device accesses the data center through the Internet of Things for centralized data processing. Below, mainly taking the cameras at each intersection of the city to regularly capture pictures of each car and upload them to the data center for processing to obtain the classification of the driving events of each car (driving from north to west, driving from south to east, etc.) as an example, an explanation will be given.
[0087] A data center is a specific network of collaborative devices used to transfer, accelerate, display, compute, and store data information on the Internet network infrastructure. In future development, data centers will also become assets for enterprise competition. With the widespread application of data centers, artificial intelligence and the like are increasingly applied to data centers. As an important technology of artificial intelligence, deep learning has been widely applied to big data analysis operations in data centers.
[0088] In a traditional large data center, the network structure is usually as Figure 1 shown, that is, a hierarchical inter-networking model. This model includes the following parts:
[0089] Server 140: Each server 140 is a processing and storage entity of the data center, and the processing and storage of a large amount of data in the data center are completed by these servers 140.
[0090] Access Switch 130: The access switch 130 is used to enable the server 140 to access the switch in the data center. One access switch 130 connects multiple servers 140. The access switches 130 are usually located at the top of the rack, so they are also called Top of Rack switches, and they are physically connected to the servers.
[0091] Aggregation Switch 120: Each aggregation switch 120 connects multiple access switches 130 and provides other services at the same time, such as firewall, intrusion detection, network analysis, etc.
[0092] Core Switch 110: The core switch 110 provides high-speed forwarding for the packets entering and leaving the data center and provides connectivity for the aggregation switches 120. The network of the entire data center is divided into an L3 layer routing network and an L2 layer routing network. The core switch 110 usually provides a flexible L3 layer routing network for the network of the entire data center.
[0093] Normally, the aggregation switch 120 is the demarcation point between the L2 and L3 layer routing networks. Below the aggregation switch 120 is the L2 network, and above is the L3 network. Each group of aggregation switches manages a Point Of Delivery (POD), and each POD is an independent VLAN network. When a server migrates within a POD, it does not need to modify the IP address and default gateway because one POD corresponds to one L2 broadcast domain.
[0094] The Spanning Tree Protocol (STP) is usually used between the aggregation switch 120 and the access switch 130. STP makes only one aggregation layer switch 120 available for a VLAN network, and other aggregation switches 120 are used only when a failure occurs (the dotted lines in the above figure). That is to say, at the level of the aggregation switch 120, horizontal expansion cannot be achieved because even if multiple aggregation switches 120 are added, only one is still working.
[0095] The cameras at each intersection in the city regularly take pictures of each car and upload them to Figure 1 a server 140 in the data center. This server 140, together with other servers 140 in the data center, processes the pictures. After being processed by the processing units in these servers 140, the driving event classifications of each car (driving from north to west, driving from south to east, etc.) are obtained.
[0096] Figure 3A The pictures in the upper left corner of -G show the images of cars driving from east to west, from east to north, from west to east, from west to north, from north to west, from north to east, and from north to south taken by the cameras at the city intersections.Figure 3A The output of the decoding layer on the right side of -G represents the classification result of the driving event of the vehicle determined by the processing unit of the embodiment of the present disclosure (the determined event classification is represented by a slash). Figure 3A - Other pictures in -G will be discussed in detail in the specific description of the embodiment of the present disclosure below. It can be seen from this that according to the vehicle driving images captured by the cameras at urban intersections, the driving direction of the vehicle can be determined by the processing unit. And these are difficult for the existing deep learning technology to achieve. The existing deep learning cannot recognize spatio-temporal related tasks, or can recognize them but with a large amount of calculation. In addition, Figure 3H Shows the event classification (driving from north to east) determined by the processing unit after inputting an actual image of a vehicle driving into the processing unit configured in the embodiment of the present disclosure. Figure 3I Shows the recognized event classification determined by the spatio-temporal neural network after inputting an actual image of multiple vehicles driving in different directions into the processing unit configured in the embodiment of the present disclosure. Among them, there is a 1% probability that there will be a vehicle driving from east to west in the image, a 10% probability that there will be a vehicle driving from east to north in the image, a 100% probability that there will be a vehicle driving from west to east in the image, a 66% probability that there will be a vehicle driving from west to north in the image, a 100% probability that there will be a vehicle driving from north to west in the image, an 88% probability that there will be a vehicle driving from north to east in the image, and a 100% probability that there will be a vehicle driving from north to south in the image.
[0097] Driverless is a further deepening of the above-mentioned urban brain application. As Figure 3I shown, after the server in the data center can determine the probability of vehicles driving in various driving directions at various locations in the city according to the images captured at various locations in the city, it can send instruction signals to the driverless vehicles at various locations, indicating which directions are safe for the driverless vehicles to drive, so as to achieve the purpose of ensuring safety while being driverless.
[0098] The language-vision fusion system is such a system that it can not only respond according to the user's voice, but also combine the user's actions and gestures to make a response. For example, the language-vision robot not only responds to the human voice, but also responds according to the human's actions and gestures. For example, when a person makes a "shh" gesture, the robot stops speaking. When a person makes a gesture of hooking the finger inward, the robot moves towards the user. However, actions and gestures are a kind of spatio-temporal behavior, and recognizing them is a spatio-temporal related task. The existing deep learning cannot recognize these tasks, or can recognize them but with a large amount of calculation. Using the processing unit of the embodiment of the present disclosure, these actions and gestures can be efficiently recognized, and combined with the human voice, the robot can make a correct response.
[0099] Embodiments of the present disclosure can also be applied to the scenario of smart home. The user can control smart home devices not by voice, but by gestures. For example, when the user makes a gesture of spreading the hands from the center to both sides, the curtain is controlled to open. When the user makes a gesture of pressing a remote control, the TV is automatically turned on. Existing deep learning cannot recognize these tasks, or can recognize them but with a large amount of computation. By using the processing unit of the embodiments of the present disclosure, these gestures of the user can be efficiently recognized, making it possible to control smart home devices with gestures.
[0100] Background of the Present Disclosure and General Network Structure
[0101] Currently, deep learning is widely used in various fields. However, when dealing with tasks related to time and space, the amount of computation is very large. Since this type of task requires the network to capture dynamic information across both time and space dimensions, the current mainstream deep learning is not ideal in this regard. In particular, for edge inference (deploying a deep learning model to the edge for inference), the gap between the computing performance of the hardware at the edge and the computing performance required for time and space related tasks is widening. To solve this problem, a Spatio-Temporal Neural Network (SNN) is proposed to handle tasks related to time and space.
[0102] A Spatio-Temporal Neural Network (SNN) can utilize timing and event-driven behavior to trigger computations, thereby encoding discrete events in the real world as spike trains, serving as inputs along the time frame, and adopting a brain-like parallel integration process to interpret the encoded spike trains to obtain inference results through various neural decoding methods (such as spike count decoding).
[0103] There are currently three types of SNNs: 1) SNNs based on backpropagation (BP); 2) SNNs based on convolutional neural networks (CNNs); 3) SNNs based on biological learning behaviors. For BP-based SNNs, a spatio-temporal backpropagation (STBP) algorithm for dynamic N-MNIST dataset classification has been proposed. This algorithm successfully combines the time domain and time domain kernels and achieves an inference accuracy of 98.78% under a fully connected architecture. However, the disadvantage is that due to its BP-based nature, the computational complexity is high. For CNN-based SNNs, spiking neurons are configured as a convolutional neural network (CNN) for pose recognition. By utilizing event-driven sensor inputs, the computational complexity of the system is reduced to a certain extent, resulting in a power consumption of 178.8 mW and an accuracy of 96.49% on customized hardware. However, due to the large number of convolutional operations involved, this computational complexity is still high. SNNs based on biological learning behaviors utilize the principles of neural plasticity or spike-timing-dependent plasticity (STDP) learning. For example, an SNN based on the mammalian olfactory bulb circuit has been developed for online odor classification. A neural plasticity rule has been proposed that uses five cycles to learn the odor spike timing. Since this network is designed specifically for signal classification, the developed network has excellent training performance, but its application scope in industry is limited.
[0104] Embodiments of the present disclosure propose a processing unit that can both reduce computational complexity and have generality. As Figure 2 shown, it includes a sensing layer 141, a spatio-temporal computing kernel array 144, a connection layer 145, and a decoding layer 148. Such a structure actually constitutes a new spatio-temporal neural network. The sensing layer 141 includes a plurality of sensing units 147. The decoding layer 148 includes a plurality of decoding units 146. The spatio-temporal computing kernel array 144 includes spatial neurons 143 and temporal neurons 144.
[0105] Each sensing unit 147 of the sensing layer 141 encodes the spatio-temporal event signal to be recognized into a pulse sequence, which reflects both the time element and the space element of the spatio-temporal event to be recognized. The spatio-temporal computing core array 144 performs double accumulation of time and space on the pulse sequence output by the sensing layer 141 to generate a spatio-temporal accumulation result pulse, which reflects the characteristics accumulated by the spatio-temporal event to be recognized in time and space. The connection layer 145 connects the decoding unit 146 in the decoding layer 148 to the spatio-temporal computing core in the spatio-temporal computing core array 144. This connection relationship is pre-trained. Then, each decoding unit 146 of the decoding layer 148 generates an output value based on the spatio-temporal accumulation result pulse output by the connected spatio-temporal computing core, and performs event classification based on the output value. In the whole process, only by accumulating pulses, that is, performing addition and integration operations, which is different from the existing BP and CNN-based methods, greatly reduces the required computational amount. At the same time, there is no restriction on the application scope. This neural network structure is particularly conducive to deployment to edge and terminal devices.
[0106] A spatio-temporal event refers to the behavior performance of an object in the real world in time and space. Any object in the real world can be regarded as an object, such as a vehicle and a human hand. An object is located at a position in space at a time point on the time axis, and may be located at other positions in space at other time points. The change of position at different time points forms the behavior performance in time and space, that is, a spatio-temporal event. Figure 4 Three spatio-temporal events of a car induced from the real world are shown, where event 1 represents the car driving from north to west, event 2 represents the car driving from east to south, and event 3 represents the car driving from north to east, where F n represents the position of the car at the current sampling time point, F n-1 represents the position of the car at the penultimate sampling time point before the current sampling time point, F n-2 represents the position of the car at the second penultimate sampling time point before the current sampling time point, F n-3 represents the position of the car at the third penultimate sampling time point before the current sampling time point. Events 1-3 are three events induced from the images taken by a camera at a certain intersection in the city at regular intervals. In addition, for a speech-vision fusion system, a robot recognizes the change of the position of a human hand at different time points from the consecutive pictures of the user taken, so as to recognize the actions and gestures made by the user, and these actions and gestures are spatio-temporal events. For a smart home, an indoor camera recognizes the change of the position of a human hand at different time points from the consecutive pictures of the user taken, so as to recognize the gesture, that is, a spatio-temporal event, and reacts to the spatio-temporal event.
[0107] The spatio-temporal event to be recognized is a spatio-temporal event for which its event classification is to be determined. Generally, a spatio-temporal neural network is used to comprehensively represent an event in terms of time and space to classify the event into a category, i.e., event classification. For example, based on the trajectory photos of each captured car, the driving event of the car is determined to be driving from north to west. Based on the change in the position of the human hand recognized in consecutive pictures of the captured user at different time points, the user gesture is determined to be "Follow me".
[0108] A spatio-temporal event signal is a signal that represents an abstract spatio-temporal event. Spatio-temporal events are abstract and must be represented as spatio-temporal event signals for data processing. For example, for the spatio-temporal event of a car driving from north to west, the pixels in the trajectory photo of the car taken at each sampling time point are the spatio-temporal event signal transformed from this spatio-temporal event. For example, each frame of the captured photo is 480×320 pixels, 1 frame is taken every 1 second, and there are 15 consecutive seconds with 15 frames taken. Then these 480×320×15 pixel values are the spatio-temporal event signal.
[0109] When the spatio-temporal event signal to be recognized is the pixel values of an image collected at consecutive sampling points, the number of sensing units included in the sensing layer can be equal to the number of pixels in the image (per frame) collected at each sampling point. In this way, each sensing unit can correspond to a pixel position within a frame and specifically process the pixel values at that pixel position within the frame. Taking the above spatio-temporal event signal of 480×320×15 pixel values as an example, since each frame has 480×320 = 153,600 pixels, there are 153,600 sensing units, and each sensing unit corresponds to a pixel position within a frame.
[0110] At this time, the spatio-temporal event signal to be recognized can be regarded as a function V(x k ,t m ) of position and time, where x k represents the k-th pixel position of the collected image frame, for example, the k-th pixel position among the 153,600 pixel positions in each frame as described above, and t m represents the m-th sampling time point. V(x k ,t m ) represents the pixel value at the k-th pixel position of the frame collected at the m-th sampling time point. In this way, the spatio-temporal event signal to be recognized becomes a function V(x k ,t m ) of position and time. Then, for a specific position, the function value V(x k and the current time t m at the specific position x k ,t m ) and the function value at the specific position x k and the previous cycle of the current time t mThe function value V(x) at the time point of -Δt k ,t m -Δt) difference ΔV(x k ,t m ), that is, Formula 1:
[0111] ΔV(x k ,t m )=V(x k ,t m )-V(x k ,t m -Δt) Formula 1
[0112] The cycle length Δt may be equal to the interval between sampling time points, such as 1 second mentioned above, or may be an integer multiple of the interval between sampling time points, such as 3 seconds, etc. k and the current time t m The function value V(x k ,t m ) and the specific position x k and the period t before the current time m The function value V(x) at the time point of -Δt k ,t m -Δt) is subtracted to obtain the specific position x k The change in the pixel value at the current frame relative to the frame of the previous cycle. Therefore, if a specific position x in the image frames of several consecutive cycles collected k If the pixel value changes greatly, it means that the position is a key position that needs to be paid attention to, because it changes greatly in the recent consecutive frames and is likely to reflect the characteristics of the spatiotemporal event to be identified.
[0113] Then, according to the comparison result of the difference with the predetermined difference threshold, it is determined whether the pulse at the specific position is a valid pulse or an invalid pulse. For example, the pulse generated at the specific position is determined according to the following formula, that is, Formula 2:
[0114]
[0115] C k That is, the predetermined difference threshold may vary with the pixel position in the acquired image frame, that is, each pixel position in the acquired image frame corresponds to a difference threshold, so that there are as many difference thresholds as there are pixel positions in the image frame. In this case, C k Indicates the predetermined difference threshold corresponding to the k-th pixel position in the frame. [u] + is a sign function. When u is positive, [u] + = 1. On the contrary, [u] += 0. For Equation 2, when the difference ΔV(x k , t m ) is greater than a predetermined difference threshold C k , an impulse at a specific position k at the m-th sampling time point is output as 1, that is, an effective impulse is output, otherwise 0 is output, that is, an invalid impulse. The determined impulses at each position are combined into the impulse sequence. For example, in the order of positions, the corresponding determined impulses are combined into an impulse sequence. That is, in the impulse sequence, the impulse at the k - 1-th position is before the impulse at the k-th position, the impulse at the k - 2-th position is before the impulse at the k - 1-th position, and so on.
[0116] Comparing the difference obtained from Equation 1 with the predetermined difference threshold to determine whether to generate an effective impulse or an invalid impulse at a specific position is to avoid excessive occupation of storage space caused by simply storing the obtained difference. In fact, comparing the difference obtained from Equation 1 with the predetermined difference threshold can reflect whether the pixel value change at a specific position x k of several consecutive periods of image frames collected is large, thus indicating whether this position is a key position that needs attention.
[0117] As Figure 5 shown, for the sensing unit 147, assuming Δt = 1 and C k = 10, for the time point t n-3 , the received V(x k , t m ) for a specific pixel position k is 20, and at the time point t n-3 - 1, the received V(x k , t m - 1) for the specific pixel position k is 8. Therefore, 20 - 8 > 10, and an effective impulse 1 is generated at the time point t n-3 ; for the time point t n-2 , the received V(x k , t m ) for the specific pixel position k is 18, and at the time point t n-2 - 1, the received V(x k , t m - 1) for the specific pixel position k is 18. Therefore, 18 - 18 < 10, and an invalid impulse 0 is generated at the time point t n-2 ; similarly, at the time points t n-1 and t n , invalid impulses 0 are also generated... In this way, the sensing unit 147 outputs an impulse sequence... 1000...
[0118] The spatio-temporal computing core array 142 includes a plurality of spatio-temporal computing cores. Each spatio-temporal computing core includes a spatial neuron 143 and a temporal neuron 144 connected in sequence.
[0119] Each spatial neuron 143 is connected to at least a portion of the sensing units 147. Among them, the outputs of every 4 sensing units 147 are connected to one spatial neuron 143. In the example where there are 153,600 sensing units, there are 153,600 / 4 = 38,400 spatial neurons 143. The at least a portion of the sensing units 147 connected to each spatial neuron 143 are the sensing units 147 that sense the pixels at adjacent pixel positions in the image frame captured by sensing. Figure 2
[0120] Each spatial neuron 143 generates a spatial cumulative result pulse sequence reflecting the spatial cumulative result of the pulse sequence according to the pulse sequence output by at least a portion of the sensing units 147 connected thereto. Pulses generated at the same time point and at adjacent spatial positions often exhibit some similar characteristics. Therefore, accumulating these pulse sequences spatially can reflect these similar characteristics. The spatial cumulative result pulse sequence obtained by the spatial neuron 143 is to reflect these characteristics. One way to obtain the spatial cumulative result pulse sequence is to directly accumulate the pulses generated at the same time point and at adjacent spatial positions, and then convert the cumulative result into a pulse sequence of "1" or "0"; another way is to first accumulate the pulses generated at the same time point and at adjacent spatial positions, and if the accumulation reaches a predetermined threshold, generate a valid pulse "1", otherwise generate an invalid pulse "0". The latter is beneficial to saving storage space, and the comparison with the predetermined threshold can also generally reflect the more or less of the valid pulses generated at the spatial positions near a certain point, without causing too much damage to the information volume. However, generally accumulating and converting into a pulse sequence often occupies more pulse bits to convey similar information, resulting in a high storage space cost.
[0121] In one embodiment of the present disclosure, the number of pulses output by at least a part of the sensing units 147 connected to the spatial neuron 143 at a specific time point is determined as the number of valid pulses, that is, the number of "1"s. Since at least a part of the sensing units 147 connected to the spatial neuron 143 are all sensing units 147 that sense the pixel changes at adjacent positions in the captured image, if most of the pulses generated at these adjacent positions are 1s, that is, valid pulses, it indicates that the regions of these pixels have changed significantly compared to the previous few frames and should be extracted as important regions for analysis. Therefore, it is determined how many of the pulses output by at least a part of the sensing units 147 connected to the spatial neuron 143 are valid pulses. If the number is greater than the first threshold, a valid pulse is output at the specific time point; otherwise, an invalid pulse is output. Then, the pulses output at each time point are synthesized into the spatial cumulative result pulse sequence. The membrane voltage of the spatial neuron 143 has the following formula 3:
[0122]
[0123] where v s represents the membrane voltage of the spatial neuron 143, represents the time derivative of the membrane voltage of the spatial neuron 143, i = 1, 2, 3, …… M / α, M is the total number of sensing units 147, α is the number of sensing units 147, and M / α is the number of sensing units 147 connected to each spatial neuron 143, s i is the pulse output by the i-th sensing unit 147 connected to the spatial neuron 143 at a specific sampling time point. Therefore, it can be seen from formula 3 that the membrane voltage of the spatial neuron 143 is actually the integral result of the pulses output by the sensing units 147 connected to the spatial neuron 143 over time. If the sensing units 147 connected to the spatial neuron 143 output invalid pulses, it has no impact on the integration. Therefore, it is determined how many of the pulses output by at least a part of the sensing units 147 connected to the spatial neuron 143 are valid pulses, and this number is the integral value. If the number is greater than the first threshold, a valid pulse is output at the specific time point; otherwise, an invalid pulse is output.
[0124] As Figure 5 shown, a spatial neuron 143 is connected to 4 sensing units 147 and receives the pulses output by the 4 sensing units 147 at each sampling time point. For example, at the sampling time point t n-3 , among the pulses output by the 4 sensing units 147 received, the 1st, 2nd, and 4th channels are "1", and the accumulated number is 3, which is greater than the first threshold 2. A valid pulse "1" is output at the sampling time point t n-3 ; at the sampling time point t n-2, the second to fourth of the pulses output by the 4 sensing units 147 are "1", the accumulated number is 3, which is greater than the first threshold 2, at the sampling time point t n-2 Output a valid pulse "1"; at the sampling time point t n-1 , the second to fourth of the pulses output by the 4 sensing units 147 are "1", the accumulated number is 3, which is greater than the first threshold 2, at the sampling time point t n-1 Output a valid pulse "1"; at the sampling time point t n , the second to third of the pulses output by the 4 sensing units 147 are "1", the accumulated number is 2, which is not greater than the first threshold 2, at the sampling time point t n-1 Output a valid pulse "0". Therefore, the spatial neuron 143 outputs a spatial accumulation result pulse sequence of...1110....
[0125] Then, the temporal neuron 144 generates the spatio-temporal accumulation result pulse reflecting the temporal accumulation result of the spatial accumulation result pulse sequence according to the spatial accumulation result pulse sequence generated by the spatial neuron 143. Pulses generated at the same spatial position at different time points often exhibit some similar characteristics. Therefore, accumulating these pulse sequences in time can reflect these similar characteristics. One way to generate the spatio-temporal accumulation result pulse is to directly accumulate the spatial accumulation result pulse sequence generated by the spatial neuron, and the other is to accumulate first. If the accumulation reaches a predetermined threshold, a valid pulse "1" is generated, otherwise an invalid pulse "0" is generated. The latter is beneficial to saving storage space, and the comparison with the predetermined threshold can also generally reflect the more or less of the valid pulses generated near a certain time point, and will not cause too much damage to the amount of information.
[0126] In an embodiment of the present disclosure, the temporal neuron 144 performs an integration within a time window on the spatial accumulation result pulse sequence generated by the spatial neuron 143 connected to it. If the integration value is greater than the second threshold, a valid spatio-temporal accumulation result pulse is output, otherwise an invalid spatio-temporal accumulation result pulse is output. If the proportion of valid pulse of the spatial accumulation result pulse generated by the spatial neuron 143 connected to the temporal neuron 144 is very high at several consecutive sampling time points recently, it is very likely that the pixel values in the corresponding area of the image frames captured during this period change significantly and should be extracted as an important area for analysis. By comparing the integration value with the second threshold, it can be seen whether the pixel values in the corresponding area change significantly within a period of time. If it is obvious, a valid pulse is output, otherwise an invalid pulse is output.
[0127] The membrane voltage of the temporal neuron 144 has the following formula 4:
[0128]
[0129] Among them, v t represents the membrane voltage of the temporal neuron 144, represents the time derivative of the membrane voltage of the temporal neuron 144, T is the total number of sampling time points, and β is the length of the time window, represents the pulse generated by the spatial neuron 143, represents the integral of the pulses output by the spatial neuron 143 at each time sampling point within a time window over that time window. If the spatial neuron 143 outputs an invalid pulse, it has no impact on the integral. Therefore, determine how many valid pulses are received from the spatial neuron 143 within that time window, and this number is the integral value. If the said number is greater than the second threshold, output a valid pulse; otherwise, output an invalid pulse.
[0130] As Figure 5 shown, assume the time window is 4, that is, it includes 4 sampling time points. The temporal neuron 144 receives a pulse "1" from the spatial neuron 143 at the sampling time point t n-3 , receives a pulse "0" from the spatial neuron 143 at the sampling time point t n-4 , receives a pulse "1" from the spatial neuron 143 at the sampling time point t n-5 , and receives a pulse "1" from the spatial neuron 143 at the sampling time point t n-6 . Therefore, the integral value of the pulse sequence generated by the spatial neuron 143 within a time window is 3, which is greater than the second threshold, and a valid pulse "1" is output as the spatio-temporal accumulation result pulse.
[0131] Each decoding unit 146 corresponds to an event classification respectively. It is pre-configured in the connection layer 145 to connect each decoding unit 146 to several spatio-temporal computing cores. Each decoding unit 146 generates an output value according to the spatio-temporal accumulation result pulses output by the connected spatio-temporal computing cores. For example, accumulate the number of valid pulses among these spatio-temporal accumulation result pulses as the output value, and take the event classification corresponding to the decoding unit with the largest output value as the recognized event classification. For example, as Figure 3A shown in -I, there are 7 event classifications, namely traveling from east to west, traveling from east to north, traveling from west to east, traveling from west to north, traveling from north to west, traveling from north to east, and traveling from north to south. Correspondingly, there are 7 decoding units 146, and the 7 decoding units 146 are respectively connected to some spatio-temporal computing cores, and accumulate the number of valid pulses among the spatio-temporal accumulation result pulses output by the corresponding spatio-temporal computing cores as the output value. Whichever decoding unit 146 obtains the largest output value is considered to be the event classification recognized by the processing unit of the present disclosure embodiment.
[0132] When decoding the connection relationship between the decoding unit 146 and the spatio-temporal computing core in the pre-configured connection layer 145, considering the fairness of event classification recognition based on the output value of the decoding unit 146, each decoding unit 146 is respectively connected to a predetermined number of spatio-temporal computing cores. For example, each decoding unit 146 is connected to 5 spatio-temporal computing cores. Since event classification actually compares how many spatio-temporal computing cores that output valid pulses among the spatio-temporal computing cores connected to each decoding unit 146, making the number of spatio-temporal computing cores connected to each decoding unit 146 equal can provide a reasonable basis for event classification. The specific method of selecting a predetermined number of spatio-temporal computing cores to be connected to a decoding unit 146 will be described in detail below.
[0133] After respectively connecting each of the decoding units 146 to a predetermined number of spatio-temporal computing cores, a set containing at least one spatio-temporal event signal sample of each event classification is input into the processing unit. The spatio-temporal event signal sample is the spatio-temporal event signal serving as a training sample. Taking the spatio-temporal event signal as the pixels in the trajectory photo of a car taken at each sampling time point as an example, the spatio-temporal event signal sample is the pixels in the sample of the trajectory photo of the car taken at each sampling time point. To comprehensively configure the processing unit of the embodiments of the present disclosure so that it can recognize all event classifications, it is necessary to ensure that there is at least one sample of each event classification in this set. Taking Figure 3A -I as an example, each of the event classifications of traveling from east to west, traveling from east to north, traveling from west to east, traveling from west to north, traveling from north to west, traveling from north to east, and traveling from north to south must ensure that there is at least one. For each spatio-temporal event signal sample, it is input into the sensing layer 141, and the spatio-temporal cumulative result pulse is output through the spatio-temporal computing core. Determine the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to each decoding unit 146. If the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to the decoding unit 146 corresponding to the event classification of this spatio-temporal event signal sample is less than the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to other decoding units 146, adjust the spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of this spatio-temporal event signal sample.
[0134] The following takes Figure 6A -F as an example to illustrate the specific training process of the connection layer 145. Figure 6A The spatio-temporal computing core array 142 is shown, where each circle represents a spatio-temporal computing core. Figure 6A The spatio-temporal computing core array of has 5×4 = 20 spatio-temporal computing cores. Figure 6B Each row of represents 3 decoding units 146, corresponding to event classifications 1-3 respectively, and each column represents the event classification actually determined by the processing unit after inputting the spatio-temporal event signal samples of each event classification. The actually determined time classification is represented by a solid circle. InFigure 6B In it, the first column indicates that when the spatio-temporal event signal sample of event classification 1 is input, the event classification determined by the processing unit is event classification 1; the second column indicates that when the spatio-temporal event signal sample of event classification 2 is input, the event classification determined by the processing unit is event classification 1, but it should be event classification 2 (represented by a hollow solid circle); the third column indicates that when the spatio-temporal event signal sample of event classification 3 is input, the event classification determined by the processing unit is event classification 2, but it should be event classification 3 (represented by a hollow solid circle). Figure 6A It shows the initial connection of each decoding unit 146 to 5 spatio-temporal computing cores. The spatio-temporal computing core connected to decoding unit 1 is represented by a triangle, the spatio-temporal computing core connected to decoding unit 2 is represented by a diamond, and the spatio-temporal computing core connected to decoding unit 3 is represented by a square. Figure 6A In it, the spatio-temporal computing cores connected to decoding unit 1 are (1,3), (2,3), (3,1), (3,2), (4,1), a total of 5 spatio-temporal computing cores; the spatio-temporal computing cores connected to decoding unit 2 are (1,3), (2,3), (3,3), (3,4), (3,5), a total of 5 spatio-temporal computing cores; the spatio-temporal computing cores connected to decoding unit 3 are (2,3), (2,4), (3,3), (4,2), (4,3), a total of 5 spatio-temporal computing cores. The abscissa represents the row number of the spatio-temporal computing core in the spatio-temporal computing core array, and the ordinate represents the column number of the spatio-temporal computing core in the spatio-temporal computing core array. Some spatio-temporal computing cores are connected to multiple decoding units.
[0135] When as Figure 6B As shown, when the spatio-temporal event signal sample of event classification 1 is input, the number of valid pulses output in each of the 5 spatio-temporal computing cores respectively connected to each of the decoding units 1-3 is determined. Assume that 5 valid pulses are output in the 5 spatio-temporal computing cores connected to decoding unit 1, 4 valid pulses are output in the 5 spatio-temporal computing cores connected to decoding unit 2, and 3 valid pulses are output in the 5 spatio-temporal computing cores connected to decoding unit 3. 5 is the maximum number of valid pulses. Therefore, it is determined that the event classification 1 corresponding to decoding unit 1 is the determined event classification, which is exactly the same as the event classification 1 of the input spatio-temporal event signal sample and no adjustment is made.
[0136] After inputting the spatio-temporal event signal samples of event classification 2, determine the number of valid pulses output by each of the five spatio-temporal calculation cores respectively connected to each of the decoding units 1-3. Assume that 5 valid pulses are output by the five spatio-temporal calculation cores connected to decoding unit 1, 4 valid pulses are output by the five spatio-temporal calculation cores connected to decoding unit 2, and 1 valid pulse is output by the five spatio-temporal calculation cores connected to decoding unit 3. 5 is the maximum number of valid pulses. Therefore, it is determined that event classification 1 corresponding to decoding unit 1 is the determined event classification, which is inconsistent with event classification 2 of the input spatio-temporal event signal sample. That is, the number of valid pulses 4 output by the five spatio-temporal calculation cores connected to the decoding unit corresponding to event classification 2 of this spatio-temporal event signal sample is less than the number of valid pulses 5 output by the five spatio-temporal calculation cores connected to decoding unit 1. Adjust the connection relationship between decoding unit 2 corresponding to event classification 2 of this spatio-temporal event signal sample and the spatio-temporal calculation cores, and the connection relationship between this decoding unit 1 and the spatio-temporal calculation cores, so that the number of valid pulses output by the five spatio-temporal calculation cores connected to decoding unit 2 corresponding to event classification 2 of this spatio-temporal event signal sample is greater than the number of valid pulses output by the five spatio-temporal calculation cores connected to decoding unit 1. As Figure 6A shown by the arrow, change the connection of spatio-temporal calculation cores (2,3) to be connected to decoding units 2-3 only instead of being connected to the three decoding units simultaneously, and change spatio-temporal calculation cores (1,2) to be connected to decoding unit 1. The modified result is as Figure 6C shown. In this way, the number of spatio-temporal calculation cores connected to each decoding unit still remains 5. And finally, the number of valid pulses output by the five spatio-temporal calculation cores connected to decoding unit 2 becomes 5, while the number of valid pulses output by the five spatio-temporal calculation cores connected to decoding unit 1 becomes 4. The processing unit determines event classification 2, which is consistent with event classification 2 of this sample, as Figure 6D shown.
[0137] After inputting the spatio-temporal event signal samples of event classification 3, determine the number of valid pulses output by each of the 5 spatio-temporal computing cores respectively connected to each of the decoding units 1-3. Assume that 1 of the 5 spatio-temporal computing cores connected to decoding unit 1 outputs a valid pulse, 4 of the 5 spatio-temporal computing cores connected to decoding unit 2 output valid pulses, and 3 of the 5 spatio-temporal computing cores connected to decoding unit 3 output valid pulses. 4 is the maximum number of valid pulses. Therefore, it is determined that the event classification 2 corresponding to decoding unit 2 is the determined event classification, which is inconsistent with the event classification 3 of the input spatio-temporal event signal sample. That is, the number of valid pulses 3 output by the 5 spatio-temporal computing cores connected to the decoding unit corresponding to the event classification 3 of this spatio-temporal event signal sample is less than the number of valid pulses 4 output by the 5 spatio-temporal computing cores connected to decoding unit 2. Adjust the connection relationship between decoding unit 3 corresponding to the event classification 3 of this spatio-temporal event signal sample and the spatio-temporal computing cores, and the connection relationship between this decoding unit 2 and the spatio-temporal computing cores, so that the number of valid pulses output by the 5 spatio-temporal computing cores connected to decoding unit 3 corresponding to the event classification 3 of this spatio-temporal event signal sample is greater than the number of valid pulses output by the 5 spatio-temporal computing cores connected to decoding unit 2. As Figure 6C shown by the arrow, change the connection of spatio-temporal computing cores (2,3) to decoding units 2-3 to only connect to decoding unit 3, change spatio-temporal computing cores (1,4) to connect to decoding unit 2, change the connection of spatio-temporal computing cores (3,3) to decoding units 2-3 to only connect to decoding unit 3, change spatio-temporal computing cores (4,4) to connect to decoding unit 2. The modified result is as Figure 6E shown. In this way, there are still 5 spatio-temporal computing cores connected to each decoding unit. And finally, the number of valid pulses output by the 5 spatio-temporal computing cores connected to decoding unit 3 becomes 4, while the number of valid pulses output by the 5 spatio-temporal computing cores connected to decoding unit 2 becomes 3. The processing unit determines event classification 3, which is consistent with the event classification 3 of this sample, as Figure 6F shown. The training process is completed. After the training process is completed, input any spatio-temporal event signal to be recognized into the sensing layer 141. The decoding unit generates an output value according to the number of valid pulses in the spatio-temporal cumulative result pulses output by the spatio-temporal computing cores it is connected to. The event classification corresponding to the decoding unit with the largest output value is the recognized event classification.
[0138] Figure 3A-G respectively show the spatio-temporal event signal samples (images of the input vehicle traveling in the above various directions) of the input vehicle traveling from east to west, from east to north, from west to east, from west to north, from north to west, from north to east, and from north to south, the output result of the sensing layer 141, the output result of the spatial neuron 143, the output result of the temporal neuron 144, and the event classification results determined after adjusting the connection relationship between the decoding unit and the spatio-temporal calculation kernel in the 7 event classifications corresponding to the 7 decoding units (the determined event classifications are represented by slashes). Figure 3H Shows the event classification (traveling from north to east) determined by the processing unit after inputting an actual image of a vehicle traveling into the processing unit configured in the embodiment of the present disclosure. Figure 3I Shows the recognized event classifications determined by the processing unit after inputting an actual image of multiple vehicles traveling in different directions into the processing unit configured in the embodiment of the present disclosure.
[0139] Next, it is discussed how to select a predetermined number of spatio-temporal calculation kernels to connect to each decoding unit 146. In one embodiment, for each decoding unit 146, multiple spatio-temporal event signal samples corresponding to the event classification of this decoding unit 146 can be selected and input into the processing unit respectively, and the number of times of outputting valid pulses for the multiple spatio-temporal event signal samples is determined for each spatio-temporal calculation kernel. Then, the spatio-temporal calculation kernels with the top predetermined number of times of outputting valid pulses from high to low are connected to this decoding unit.
[0140] For example, there are 3 event classifications, and there are also 3 corresponding decoding units 146. For each decoding unit 146, multiple spatio-temporal event signal samples corresponding to its corresponding event classification (for example, for a vehicle traveling from north to west, multiple images of such vehicles traveling are taken as multiple samples) are constructed and input into the sensing layer 141. For a sample of this event classification, the spatio-temporal calculation kernels connected to this decoding unit 146 will each generate a spatio-temporal accumulation result pulse, either a valid pulse or an invalid pulse. Then, the number of times of generating valid pulses for each spatio-temporal calculation kernel for these multiple spatio-temporal event signal samples is counted, and this number must be less than or equal to the number of samples. For example, the number of samples is 100, and among these 100 samples, one spatio-temporal calculation kernel generates valid pulses for 70 of these 100 samples, then the number of times of outputting valid pulses is 70. Then, the spatio-temporal calculation kernels with the top predetermined number of times of outputting valid pulses from high to low are connected to this decoding unit. If a spatio-temporal calculation kernel generates more valid pulses for the sample set of the same event classification, it means that its role in distinguishing this event classification is greater, and it should be connected to the corresponding decoding unit more. This embodiment improves the accuracy of configuring the processing unit.
[0141] In the embodiments of the present disclosure, for each event classification, only at least one positive data sample (data sample under this event classification) is used for training, reducing the training cost. In the whole training and inference process, through addition and integration operations, the complexity of the learning algorithm is reduced, thereby reducing the inference calculation amount and the training cost. The embodiments of the present disclosure can maintain the same inference accuracy while changing the network scale, and are particularly suitable for the deployment of spatio-temporal neural networks based on edge and terminal devices.
[0142] As Figure 7 shown, according to an embodiment of the present disclosure, a spatio-temporal event recognition method is provided, including:
[0143] Step 710: Encode the spatio-temporal event signal to be recognized into a pulse sequence reflecting the time and space elements of the spatio-temporal event to be recognized;
[0144] Step 720: Generate a spatio-temporal accumulation result pulse reflecting the spatio-temporal accumulation result of at least a part of the pulse sequence according to at least a part of the pulse sequence;
[0145] Step 730: Let multiple decoding units respectively generate output values according to the spatio-temporal accumulation result pulses output by a part of the spatio-temporal calculation kernels connected to the decoding unit, and use the event classification corresponding to the decoding unit with the largest output value as the recognized event classification.
[0146] Optionally, step 720 includes:
[0147] Generate a spatial accumulation result pulse sequence reflecting the spatial accumulation result of at least a part of the pulse sequence according to at least a part of the pulse sequence;
[0148] Generate the spatio-temporal accumulation result pulse reflecting the time accumulation result of the spatial accumulation result pulse sequence according to the spatial accumulation result pulse sequence.
[0149] Optionally, step 710 includes:
[0150] Regarding the spatio-temporal event signal to be recognized as a function of position and time, for a specific position, determine the difference between the function value at the specific position and the current time and the function value at the time point of the previous cycle of the specific position and the current time;
[0151] According to the comparison result between the difference and a predetermined difference threshold, determine whether the pulse at the specific position is a valid pulse or an invalid pulse;
[0152] Combine the determined pulses at each position into the pulse sequence.
[0153] Optionally, generating a spatial cumulative result pulse sequence that reflects a spatial cumulative result of at least a part of the pulse sequence according to at least a part of the pulse sequence includes:
[0154] Determining the number of valid pulses determined according to at least a part of the pulse sequence at a specific time point;
[0155] If the number is greater than a first threshold, outputting a valid pulse at the specific time point, otherwise outputting an invalid pulse;
[0156] Combining the pulses output at each time point into the spatial cumulative result pulse sequence.
[0157] Optionally, generating the spatio-temporal cumulative result pulse that reflects a temporal cumulative result of the spatial cumulative result pulse sequence according to the spatial cumulative result pulse sequence includes:
[0158] Generating an integral within a time window for the spatial cumulative result pulse sequence;
[0159] If the integral value is greater than a second threshold, outputting a valid spatio-temporal cumulative result pulse, otherwise outputting an invalid spatio-temporal cumulative result pulse.
[0160] Optionally, generating an output value according to the spatio-temporal cumulative result pulses output by a part of the spatio-temporal calculation cores connected to the decoding unit includes: using the number of valid pulses among the spatio-temporal cumulative result pulses output by a part of the spatio-temporal calculation cores connected to the decoding unit as the output value.
[0161] The implementation details of the method embodiment have been described in detail in the above device embodiment. For the sake of saving space, they will not be elaborated here.
[0162] As Figure 8 shown, according to an embodiment of the present disclosure, there is also provided a method for configuring a processing unit. The processing unit includes a sensing layer, a spatio-temporal calculation core array, a connection layer, and a decoding layer. The spatio-temporal calculation core array includes a plurality of spatio-temporal calculation cores, and the decoding layer includes a plurality of decoding units. The connection layer is used to connect the plurality of decoding units to the plurality of spatio-temporal calculation cores. The method includes:
[0163] Step 810: Connecting the plurality of decoding units to a predetermined number of spatio-temporal calculation cores among the plurality of spatio-temporal calculation cores respectively;
[0164] Step 820: Input a set of spatio-temporal event signal samples each containing at least one spatio-temporal event signal sample of each event classification into the processing unit. For each spatio-temporal event signal sample, determine the number of valid pulses output by a predetermined number of spatio-temporal computing cores connected to each decoding unit. If the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of this spatio-temporal event signal sample is less than the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to other decoding units, adjust the spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of this spatio-temporal event signal sample and the spatio-temporal computing cores connected to the other decoding units, so that the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of this spatio-temporal event signal sample is greater than the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to the other decoding units.
[0165] Optionally, step 810 includes:
[0166] For each decoding unit, select multiple spatio-temporal event signal samples of the event classification corresponding to this decoding unit and input them into the processing unit respectively, and determine the number of times of outputting valid pulses for the multiple spatio-temporal event signal samples for each spatio-temporal computing core respectively;
[0167] Connect the top predetermined number of spatio-temporal computing cores with the number of times of outputting valid pulses from high to low to this decoding unit.
[0168] The implementation details of the method embodiment have been described in detail in the above device embodiment. For the sake of saving space, they will not be repeated here.
[0169] Commercial Value of the Present Disclosure
[0170] Experiments prove that the processing time of the processing unit deployed in the embodiment of the present disclosure during actual use is reduced to 60% of the existing spatio-temporal neural network, and the configuration time is reduced to 50% of the existing spatio-temporal neural network, having extremely strong market prospects.
[0171] It should be understood that the embodiments in this specification are all described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
[0172] It should be understood that the specific embodiments of this specification have been described above. Other embodiments are within the scope of the claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0173] It should be understood that an element described herein in the singular or shown as only one in the drawings does not represent that the quantity of the element is limited to one. In addition, modules or elements described or shown herein as separate may be combined into a single module or element, and modules or elements described or shown herein as a single one may be split into multiple modules or elements.
[0174] It should also be understood that the terms and expressions adopted herein are only for description, and one or more embodiments of this specification should not be limited to these terms and expressions. The use of these terms and expressions does not mean excluding any equivalent features of the illustration and description (or parts thereof), and it should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations and substitutions may also exist. Accordingly, the claims should be regarded as covering all such equivalents.
Claims
1. A processing unit, comprising: A sensing layer, including a plurality of sensing units, which are respectively used to encode a spatio-temporal event signal to be recognized into a pulse sequence, and the pulse sequence reflects the time element and the space element of the spatio-temporal event to be recognized; A spatio-temporal computing core array, including a plurality of spatio-temporal computing cores, which are respectively connected to at least a part of the plurality of sensing units, and generate a spatio-temporal cumulative result pulse according to the pulse sequence output by the at least a part of the sensing units, and the spatio-temporal cumulative result pulse reflects the spatio-temporal cumulative result of the pulse sequence; A connection layer ; A decoding layer, including a plurality of decoding units, each decoding unit corresponding to an event classification respectively. The decoding unit is connected to a corresponding part of the plurality of spatio-temporal computing cores through the connection layer, and generates an output value according to the spatio-temporal cumulative result pulse output by the corresponding part of the spatio-temporal computing cores, and takes the event classification corresponding to the decoding unit with the largest output value as the recognized event classification; Among them, the decoding unit is connected to a corresponding part of the plurality of spatio-temporal computing cores through the connection layer, including: connecting the plurality of decoding units to a predetermined number of spatio-temporal computing cores respectively, inputting a set of spatio-temporal event signal samples including at least one spatio-temporal event signal sample of each event classification into the processing unit, and determining the number of effective pulses output by the predetermined number of spatio-temporal computing cores connected to each decoding unit; if the number of effective pulses output by the predetermined number of spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of the spatio-temporal event signal sample is less than the number of effective pulses output by the predetermined number of spatio-temporal computing cores connected to other decoding units, adjust the spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of the spatio-temporal event signal sample.
2. The processing unit according to claim 1, wherein, The spatio-temporal computing core array includes a spatial neuron and a temporal neuron. Among them, the spatial neuron is connected to the at least a part of the sensing units, and generates a spatial cumulative result pulse sequence according to the pulse sequence output by the at least a part of the sensing units, and the spatial cumulative result pulse sequence reflects the spatial cumulative result of the pulse sequence; the temporal neuron generates the spatio-temporal cumulative result pulse according to the spatial cumulative result pulse sequence generated by the spatial neuron, and the spatio-temporal cumulative result pulse reflects the spatio-temporal cumulative result of the pulse sequence.
3. The processing unit according to claim 1, wherein, The connecting the plurality of decoding units to a predetermined number of spatio-temporal computing cores respectively includes: For each decoding unit, select a plurality of spatio-temporal event signal samples corresponding to the event classification of the decoding unit and input them into the processing unit respectively, and determine the number of times of outputting effective pulses for the plurality of spatio-temporal event signal samples by each spatio-temporal computing core; Connect the spatio-temporal computing cores with the top predetermined number of times of outputting effective pulses to this decoding unit from high to low.
4. The processing unit according to claim 1, wherein, The encoding the spatio-temporal event signal to be recognized into a pulse sequence reflecting the time element and the space element of the spatio-temporal event to be recognized includes: Regarding the spatio-temporal event signal to be recognized as a function of position and time, for a specific position, determine the difference between the function value at the specific position and the current time and the function value at the time point of the previous cycle of the specific position and the current time; Based on the comparison result of the difference with a predetermined difference threshold, determine whether the pulse at a specific position is a valid pulse or an invalid pulse; Combine the determined pulses at each position into the pulse sequence.
5. The processing unit according to claim 2, wherein, Generating a spatially cumulative result pulse sequence according to the pulse sequence output by at least a part of the sensing units includes: Determine the number of valid pulses among the pulses output by at least a part of the sensing units at a specific time point; If the number is greater than the first threshold, output a valid pulse at the specific time point, otherwise output an invalid pulse; Combine the pulses output at each time point into the spatially cumulative result pulse sequence.
6. The processing unit according to claim 2, wherein, Generating the spatio-temporal cumulative result pulse according to the spatially cumulative result pulse sequence generated by the spatial neuron includes: Generate an integral within a time window for the spatially cumulative result pulse sequence generated by the spatial neuron; If the integral value is greater than the second threshold, output a valid spatio-temporal cumulative result pulse, otherwise output an invalid spatio-temporal cumulative result pulse.
7. The processing unit according to claim 1, wherein, Generating an output value according to the spatio-temporal cumulative result pulse output by a corresponding part of the spatio-temporal calculation kernels includes: Use the number of valid pulses among the spatio-temporal cumulative result pulses output by the corresponding part of the spatio-temporal calculation kernels as the output value.
8. A spatio-temporal event recognition method, comprising: Encode the spatio-temporal event signal to be recognized into a pulse sequence, and the pulse sequence reflects the time element and space element of the spatio-temporal event to be recognized; Generate a spatio-temporal cumulative result pulse according to at least a part of the pulse sequence, and the spatio-temporal cumulative result pulse reflects the spatio-temporal cumulative result of at least a part of the pulse sequence; Let multiple decoding units respectively generate output values according to the spatio-temporal cumulative result pulses output by a part of the spatio-temporal calculation kernels connected to the decoding unit, and classify the event corresponding to the decoding unit with the largest output value as the recognized event classification; Wherein, the multiple decoding units are connected to a corresponding part of the spatio-temporal calculation kernels in a plurality of spatio-temporal calculation kernels through a connection layer, including: connecting the multiple decoding units to a predetermined number of spatio-temporal calculation kernels respectively, inputting a set of spatio-temporal event signal samples each containing at least one spatio-temporal event signal sample of each event classification into the processing unit, and determining the number of valid pulses output by the predetermined number of spatio-temporal calculation kernels connected to each decoding unit; if the number of valid pulses output by the predetermined number of spatio-temporal calculation kernels connected to the decoding unit corresponding to the event classification of the spatio-temporal event signal sample is less than the number of valid pulses output by the predetermined number of spatio-temporal calculation kernels connected to other decoding units, adjust the spatio-temporal calculation kernels connected to the decoding unit corresponding to the event classification of the spatio-temporal event signal sample.
9. The method according to claim 8, wherein, Generating a spatio-temporal cumulative result pulse according to at least a part of the pulse sequence includes: Generate a spatially cumulative result pulse sequence according to at least a part of the pulse sequence, and the spatially cumulative result pulse sequence reflects the spatial cumulative result of at least a part of the pulse sequence; Generate the spatio-temporal cumulative result pulse according to the spatially cumulative result pulse sequence, and the spatio-temporal cumulative result pulse reflects the time cumulative result of the spatially cumulative result pulse sequence.
10. The method according to claim 8, wherein, Encoding the spatio-temporal event signal to be recognized into a pulse sequence includes: Taking the spatio-temporal event signal to be recognized as a function of position and time, for a specific position, determining the difference between the function value at the specific position and the current time and the function value at the specific position and the time point of the previous cycle of the current time; According to the comparison result between the difference and a predetermined difference threshold, determining whether the pulse at the specific position is a valid pulse or an invalid pulse; Combining the determined pulses at each position into the pulse sequence.
11. The method according to claim 9, wherein, Generating a spatially cumulative result pulse sequence according to at least a part of the pulse sequence, including: Determining the number of pulses determined to be valid pulses according to at least a part of the pulse sequence at a specific time point; If the number is greater than a first threshold, outputting a valid pulse at the specific time point, otherwise outputting an invalid pulse; Combining the pulses output at each time point into the spatially cumulative result pulse sequence.
12. The method according to claim 9, wherein, Generating the spatio-temporal cumulative result pulse according to the spatially cumulative result pulse sequence, including: Generating an integral of the spatially cumulative result pulse sequence over a time window; If the integral value is greater than a second threshold, outputting a valid spatio-temporal cumulative result pulse, otherwise outputting an invalid spatio-temporal cumulative result pulse.
13. The method according to claim 8, wherein, Generating an output value according to the spatio-temporal cumulative result pulse output by a part of the spatio-temporal computing cores connected to the decoding unit, including: Taking the number of spatio-temporal cumulative result pulses output by a part of the spatio-temporal computing cores connected to the decoding unit that are valid pulses as the output value.
14. A processing unit configuration method, the processing unit includes a sensing layer, a spatio-temporal computing core array, a connection layer, and a decoding layer. The spatio-temporal computing core array includes a plurality of spatio-temporal computing cores, the decoding layer includes a plurality of decoding units, and the connection layer is used to connect the plurality of decoding units to the plurality of spatio-temporal computing cores. The method includes: Connecting the multiple decoding units to a predetermined number of spatio-temporal computing cores among the multiple spatio-temporal computing cores respectively; Inputting a set of spatio-temporal event signal samples each containing at least one spatio-temporal event signal sample of each event classification into the processing unit. For each spatio-temporal event signal sample, determining the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to each decoding unit. If the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of the spatio-temporal event signal sample is less than the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to other decoding units, adjusting the connection relationship between the decoding unit corresponding to the event classification of the spatio-temporal event signal sample and the spatio-temporal computing cores and the connection relationship between the other decoding units and the spatio-temporal computing cores, so that the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to the decoding unit corresponding to the event classification of the spatio-temporal event signal sample is greater than the number of valid pulses output by the predetermined number of spatio-temporal computing cores connected to other decoding units.
15. The method according to claim 14, wherein, The connecting the multiple decoding units to a predetermined number of spatio-temporal computing cores among the multiple spatio-temporal computing cores respectively includes: For each decoding unit, selecting multiple spatio-temporal event signal samples corresponding to the event classification of the decoding unit and inputting them into the processing unit respectively, and determining the number of times of outputting valid pulses for each spatio-temporal computing core for the multiple spatio-temporal event signal samples; Connecting the top predetermined number of spatio-temporal computing cores with the highest number of times of outputting valid pulses to the decoding unit.
16. A data center, including a plurality of servers, on which computer-readable code is distributedly stored. When the computer-readable code is executed by a processor on the corresponding server, it implements the spatio-temporal event recognition method according to any one of claims 8-13.
Citation Information
Patent Citations
Method for learning and recognizing image pulse data space-time information based on Spike cube SNN
CN110210563A