Spiking neural network configuration method, image recognition device, and surveillance camera system
By configuring a spiking neural network with specific convolution layers to meet the power and accuracy requirements, the method addresses the challenges of increasing surveillance cameras in video surveillance systems, reducing the need for multiple GPU servers and lowering energy consumption.
Patent Information
- Application Number
- JP2021143528
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-02
- Publication Date
- 2025-05-12
- Estimated Expiration
- 2041-09-02
AI Technical Summary
Conventional video surveillance systems face challenges in reducing power consumption and cost due to the increasing number of surveillance cameras, which requires more GPU servers and leads to higher energy consumption.
A method for constructing a spiking neural network (SNN) that converts a deep neural network (DNN) into basic blocks with convolution layers, where the input is connected to the first convolution layer and the output is a combined input and output from the last convolution layer, configured to satisfy a predetermined number of brain-type devices and recognition accuracy.
This approach allows for the appropriate configuration of a spiking neural network that meets the power consumption and recognition accuracy requirements in surveillance camera systems, reducing the number of brain-type devices needed and lowering energy consumption.
Smart Images

Figure 0007674962000001 
Figure 0007674962000002 
Figure 0007674962000003
Abstract
Description
[Technical field]
[0001] The present invention relates generally to methods for constructing spiking neural networks that mimic brain processing. [Background technology]
[0002] In recent years, the use of AI (Artificial Intelligence), especially deep learning, has progressed in various applications including image recognition applications and voice recognition applications, and social implementation is progressing in video surveillance, chatbots, etc. Video surveillance, which is one application of deep learning, can be used to detect and track specific individuals such as suspicious or missing persons in public places such as stations, airports, and streets, and to detect specific behavior such as inappropriate behavior of workers at work sites.
[0003] In conventional video surveillance systems, inference of images extracted from video output from 3-4 surveillance cameras is processed by a single GPU (Graphics Processing Unit) server using deep learning. In recent years, the number of surveillance cameras has increased, but if the number of surveillance cameras is increased, the number of GPU servers must be increased, which poses an issue of increased costs for introducing GPU servers. In addition, the power consumption of GPUs is higher than that of CPUs (Central Processing Units), etc., which poses an issue of significantly increasing energy consumption.
[0004] In image learning and inference, a neural network model called ResNet (Residual Network), which is a type of deep neural network (DNN), is used because it can improve the accuracy of recognizing complex images (Non-Patent Document 1). In addition, in order to reduce energy consumption, research is being conducted on a method of converting a DNN into a spiking neural network (SNN) that mimics brain processing and performing inference on a low-power brain-like device (Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] “Deep Residual Learning for Image Recognition”, Kaiming He, et.Al., CVPR, 2016, Volume 1, pp 770-78 [Non-Patent Document 2] “Benchmarking Keyword Spotting Efficiency on Neuromorphic Hardware”, P. Blouw, et al., in NICE '19, March 2019. Summary of the Invention [Problem to be solved by the invention]
[0006] A surveillance camera system is an example of a system that recognizes people and objects (people, etc.) in an image. In a surveillance camera system, the allowable power consumption is limited depending on the system configuration and the installation location. Therefore, when implementing a brain-like device that recognizes people, etc. inside a surveillance camera, it is necessary to reduce the number of brain-like devices as much as possible. On the other hand, in the conventional ResNet technology, it is necessary to increase the number of layers of the neural network to improve the image recognition accuracy. In addition, when an SNN converted from a DNN is executed on a brain-like device, the number of layers of the SNN that can be implemented on one brain-like device is limited. Therefore, if the number of layers of the SNN is increased to improve the image recognition accuracy, the number of brain-like devices required increases.
[0007] The present invention has been made in consideration of the above points, and aims to propose a spiking neural network configuration method etc. that can configure a spiking neural network suited to the usage environment. [Means for solving the problem]
[0008] In order to solve such problems, the present invention provides a spiking neural network configuration method for configuring a spiking neural network by converting a deep neural network that has completed learning of a deep neural network model, the deep neural network model being configured by connecting one or more basic blocks each having one or more convolutional layers, an input to the basic block being connected to a first convolutional layer of the basic block, and an output from the basic block being an output obtained by combining an input branched from the input and an output from a last convolutional layer of the basic block, and a processing unit configuring the convolutional layer of the deep neural network model so as to satisfy a predetermined number of brain-type devices that implement the spiking neural network and a predetermined recognition accuracy of inference by the spiking neural network.
[0009] In the above configuration, for example, the convolutional layer is configured to satisfy a predetermined number of brain-type devices and a predetermined inference recognition accuracy, so that a spiking neural network can be configured that satisfies the requirements of the number of brain-type devices permitted in the usage environment and the inference recognition accuracy. Effect of the Invention
[0010] According to the present invention, a spiking neural network can be appropriately configured. [Brief description of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram illustrating an example of an SNN according to a first embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of a configuration of a basic block according to the first embodiment. [Diagram 3] FIG. 2 is a diagram illustrating an example of an SNN configuration process according to the first embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of an SNN configuration process according to the first embodiment. [Diagram 5]FIG. 1 illustrates an example of an image recognition device according to a first embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of image recognition processing according to the first embodiment. [Figure 7] FIG. 1 is a diagram illustrating an example of a surveillance camera system according to a first embodiment. [Figure 8] FIG. 1 is a diagram illustrating an example of a surveillance camera system according to a first embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] (I) First embodiment An embodiment of the present invention will be described in detail below, however, the present invention is not limited to the embodiment.
[0013] In this embodiment, a method for constructing an SNN (a spiking neural network configuration method) that satisfies a predetermined number of brain-type devices and a predetermined recognition accuracy in an image to be inferred is described. More specifically, in this embodiment, a method for realizing a system that recognizes people, etc. in an image obtained from a surveillance camera, etc., that satisfies the power consumption allowed at the installation location of the system and the recognition accuracy required for the system is described.
[0014] According to this embodiment, for example, in a surveillance camera system, an SNN can be configured to satisfy a predetermined number of brain-type devices and a predetermined recognition accuracy in an image of an inference target. As a result, the system can be configured to satisfy the power consumption allowable for the installation location of the surveillance camera system.
[0015] Next, an embodiment of the present invention will be described based on the drawings. The following description and drawings are examples for explaining the present invention, and are omitted and simplified as appropriate for clarity of explanation. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.
[0016] In the following description, the same elements in the drawings are given the same numbers, and the description is omitted as appropriate. When describing elements of the same type without distinction, the common part (part excluding the branch number) of the reference sign including the branch number is used, and when describing elements of the same type with distinction, the reference sign including the branch number may be used. For example, when describing basic blocks without distinction, they may be written as "basic block 102," and when describing individual basic blocks with distinction, they may be written as "basic block 102-1," "basic block 102-2," and so on.
[0017] The designations "first," "second," "third," and the like in this specification are given to identify components and do not necessarily limit the number or order. Furthermore, numbers for identifying components are used in each context, and a number used in one context does not necessarily indicate the same configuration in another context. Furthermore, a component identified by a certain number is not prevented from also having the function of a component identified by another number.
[0018] <Spiking Neural Network Configuration> 1 shows an example of an SNN (SNN101) according to this embodiment. The SNN101 is a spiking neural network that mimics brain processing for image recognition. The SNN101 is configured by one basic block 102 or two or more basic blocks 102 connected in series.
[0019] The SNN 101 is constructed by converting a DNN that has completed learning of a DNN model. A software tool may be used for the conversion. Examples of the software tool include NxSDK from Intel Corporation and Nengo from Applied Brain Research Corporation. Note that the DNN model is constructed by connecting one basic block 102 or two or more basic blocks 102 in series, similar to the SNN 101.
[0020] 2 shows an example of the configuration of the basic block 102. The basic block 102 is composed of a convolution layer group in which one to three convolution layers 201 are connected in series. A first input 211 to the basic block 102 is connected to the first convolution layer 201-1 of the convolution layer group. A second input 212 branched from the first input 211 is combined with a first output 221 from the last convolution layer 201-2 of the convolution layer group, and becomes the output of the basic block 102 as a second output 222.
[0021] <How to configure a spiking neural network> 3 and 4 show an example of a procedure for determining the configuration of the SNN 101 shown in FIG. 1, more specifically, a procedure for determining the number of basic blocks 102 in the SNN 101 and the number of convolution layers 201 in the basic blocks 102 (SNN configuration process). In order to determine these, a predetermined number of brain-type devices and a predetermined recognition accuracy are set based on the requirements of the system that performs image recognition. The SNN 101 is configured based on these two conditions.
[0022] The SNN configuration process shown in FIG. 3 and FIG. 4 is realized by the CPU of the brain-type computer reading out a program stored in the auxiliary storage device to the main storage device and executing it, or by performing inference using an SNN in which a brain-type device of the brain-type computer is implemented. For the brain-type computer, for example, the brain-type computer 510 shown in FIG. 5 can be adopted. The functions of the brain-type computer (for example, a processing unit) may be realized, for example, by the CPU reading out a program stored in the auxiliary storage device to the main storage device and executing it (software), or may be realized by hardware such as a dedicated circuit, or may be realized by a combination of software and hardware. In addition, one function of the brain-type computer may be divided into multiple functions, or multiple functions may be integrated into one function. In addition, some of the functions of the brain-type computer may be provided as a separate function, or may be included in another function. In addition, some of the functions of the brain-type computer may be realized by another computer that can communicate with the brain-type computer.
[0023] When the processing unit starts the SNN construction process (S301), it first performs a DNN model design process (S302).
[0024] FIG. 4 shows an example of a DNN model design process.
[0025] When the processing unit starts the DNN model design process (S401), first, it determines whether or not it is the first process (S402). If it is determined that it is the first process, the processing unit sets the number of connections of the basic block 102 and the number of neurons constituting each convolution layer in the basic block 102 (S403). For example, the number of connections is set according to the complexity (for example, resolution) of the shape of a person or the like to be recognized. The number of neurons is set according to the number of pixels of an input image, for example. For example, the processing unit sets the number of connections by referring to a table in which the number of connections is defined corresponding to the resolution, and sets the number of neurons by referring to a table in which the number of neurons is defined corresponding to the number of pixels.
[0026] Next, the processing unit again determines whether or not it is the first processing (S404). If it is determined that it is the first processing, the processing unit sets the number of convolution layers in the basic block 102 to "3" (S405), and ends the DNN model design process (S406).
[0027] Next, the processing unit sets the learning parameters shown in Fig. 3 (S303). Here, the processing unit sets at least the batch size, the number of epochs, and the size and number of filters, which are weight matrices multiplied by the input to the neurons constituting the convolution layer, as parameters used in learning the DNN model.
[0028] Next, the processing unit inputs a predetermined number of training images for learning, and trains the designed DNN model based on the set parameters (S304).
[0029] The processing unit judges whether or not the value of loss (deviation from the correct image) is equal to or less than a predetermined value as a result of the learning (whether or not the learning is completed) (S305). If the processing unit judges that the value of loss is equal to or less than the predetermined value, the processing unit completes the learning and moves the process to S306, and if the processing unit judges that the value of loss is not equal to or less than the predetermined value, the processing unit changes the values (parameters) in the setting of the learning parameters (S303) and performs learning again (S304).
[0030] After the learning is completed, the processing unit converts the learned DNN into an SNN using the aforementioned software tool (S306).
[0031] Next, the converted SNN is used to perform inference on a plurality of test images. First, the processing unit sets parameters for inference (S307). In inference using the SNN, the brain-type device repeatedly inputs each test image to the SNN n times. Therefore, the processing unit sets at least the number of repetitions n. In addition, the processing unit sets at least the maximum value of the firing rate of the neuron.
[0032] When the processing unit completes these settings, it automatically implements the SNN in the brain-like device using the aforementioned software tool (S308). When the implementation is complete, the processing unit uses the software tool to display the number of brain-like devices in which the SNN is implemented. Then, the processing unit determines whether the number of brain-like devices in which the SNN is implemented is equal to or less than a predetermined number of devices (S309).
[0033] If the processing unit determines that the number of devices does not fall within the predetermined number, it designs the DNN model again (S302 and subsequent steps). The processing unit determines whether or not this is the first processing in the DNN model design processing shown in FIG. 4 (S402). If this is the second or subsequent processing, the processing unit checks whether the number of devices is NG or the recognition accuracy is NG (NG condition) in the processing of FIG. 3 (S411). Here, since the number of devices is NG, in setting the number of connections of the basic blocks (S403), the processing unit reduces the number of connections of the basic blocks by one (reducing the scale of the DNN model) and determines again whether or not this is the first processing (S404). Here, since this is not the first processing, the processing unit ends the DNN model design processing (S406).
[0034] When the processing unit determines that the number of devices is equal to or less than a predetermined number in checking the number of devices (S309), the processing unit inputs a predetermined number of test images to the brain-type device, and the brain-type device performs inference (S310).
[0035] The processing unit determines whether the recognition accuracy (correct answer rate for test images) is equal to or higher than a predetermined recognition accuracy as a result of the inference (S311). If the processing unit determines that the recognition accuracy is equal to or higher than the predetermined recognition accuracy, it ends the SNN construction process (S313). If the processing unit determines that the recognition accuracy is lower than the predetermined recognition accuracy, it determines whether the repetition with different inference parameters has exceeded a predetermined number of times (whether the repetition has been performed a predetermined number of times) (S312).
[0036] If the processing unit determines that the process has not been repeated the predetermined number of times, it changes the inference parameters in the inference parameter setting (S307) and performs the subsequent processes (S308 to S311). On the other hand, if the processing unit determines that the process has been repeated the predetermined number of times (if the result of repeating the inference by changing the inference parameters does not reach a predetermined recognition accuracy or higher), it performs the DNN model design process (S302 and subsequent processes) again. In this case, since the recognition accuracy is NG in the NG condition (S411) of FIG. 4, in the setting of the number of convolution layers in the basic block 102 (S405), the processing unit reduces the number of layers by one and ends the DNN model design process (S406). After that, the processing unit performs the learning parameter setting (S303 and subsequent processes) of FIG. 3.
[0037] According to the above processing, it is possible to configure the SNN 101 so as to satisfy a predetermined number of brain-type devices and a predetermined recognition accuracy for the image of the inference target, which are determined based on the system requirements.
[0038] <Image recognition device configuration> FIG. 5 shows an example of an image recognition device (image recognition device 500) that uses a brain-type device that implements the SNN 101 configured by the above processing.
[0039] The image recognition device 500 includes a brain-type computer 510 and one or more cameras 520 connected to the brain-type computer 510 via a network (not shown). The brain-type computer 510 includes a processing unit 511 (e.g., a CPU), a brain-type device 512, an auxiliary storage device 513, a chipset 514 connecting these, and a main storage device 515. The camera 520 acquires an image of the place where it is installed (captures an image of a subject) and outputs the image to the brain-type computer 510 in a predetermined format. The camera 520 may be connected to the chipset 514 via a network or directly.
[0040] The processing device 511 is a processor that performs data processing related to the overall operational control of the image recognition device 500. The processing device 511 executes software stored in the main memory device 515, extracts images of predetermined frames from the video input from the camera 520, inputs the images to the brain-type device 512 for processing, and recognizes people and the like appearing in the images based on the processing results. The brain-type device 512 is a dedicated device for executing a spiking neural network (SNN) that mimics the brain.
[0041] The auxiliary storage device 513 is a large-capacity non-volatile storage device such as a hard disk device or SSD (Solid State Drive). The auxiliary storage device 513 stores images sent from the camera 520, frame images extracted from the images, processing results output by the brain type device 512, and the like. The auxiliary storage device 513 and the main storage device 515 may be collectively referred to as a storage device. Although not shown, a monitor, a remote terminal, and other terminals are connected to the brain type computer 510, and instructions can be given to the image recognition device 500 by an administrator, and image recognition results can be displayed, and the like.
[0042] The main memory device 515 is, for example, a volatile semiconductor memory, and temporarily stores (holds) the software loaded from the auxiliary memory device 513, various data, video sent from the camera 520, frame images extracted from the video, etc. The functions of the image recognition device 500 are realized by the processing device 511 executing the programs stored in the main memory device 515.
[0043] <Processing operation of image recognition device> 6 shows an example of image recognition processing in the image recognition device 500. When the processing device 511 starts the image recognition processing (S601), it takes in an image output from the camera 520 (S602) and stores it in the main storage device 515.
[0044] Next, the processing device 511 sequentially reads out the video stored in the main memory device 515, extracts predetermined frames, and extracts an inference part for each frame (S603). In extracting the inference part, the processing device 511 extracts a part of a person or the like that is to be inferred. The processing device 511 then converts the image of the extracted inference part into a predetermined image size required for input to the inference process, and stores it in the main memory device 515. The processing device 511 may also store it in the auxiliary memory device 513 for later confirmation.
[0045] Next, in the inference process (S604), the processing device 511 reads out an image processed for inference (image for inference) from the main memory device 515, and inputs it to the brain-type device 512. The brain-type device 512 is equipped with the SNN 101 configured by the process shown in Fig. 2, and the brain-type device 512 inputs the image for inference to the implemented SNN 101 and performs inference processing.
[0046] Next, in the result display (S605), the brain-type device 512 stores the result of the inference processing in the main memory device 515, and the processing device 511 reads the result of the inference processing from the main memory device 515 and displays it on a monitor (not shown) connected to the brain-type computer 510.
[0047] Thereafter, the processing device 511 determines whether or not operation has ended (whether or not the administrator has issued an instruction to the image recognition device 500 to end the image recognition processing) (S606), and if it determines that operation has ended, it ends the image recognition processing (S607), and if it determines that operation has not ended, it continues the image recognition processing (S602 to S605).
[0048] <An example of a surveillance camera system configuration> FIG. 7 shows an example of a surveillance camera system (surveillance camera system 700) according to this embodiment.
[0049] The surveillance camera system 700 includes a plurality of cameras 520, a LAN switch 710, and a management device 720 such as a PC (Personal Computer) for managing the cameras 520. The plurality of cameras 520 and the management device 720 are connected to the LAN switch 710 via a LAN 730. In addition, the camera 520 is connected to a processing device (not shown) that controls the camera in the camera 520, which is a brain type device 512. Alternatively, if there is a margin of power to be supplied to the camera 520, a computer (not shown) in which the brain type computer 510 shown in FIG. 5 is implemented in the form of a small PC may be connected to each camera 520 and the image of the camera 520 may be input. With such a configuration, it is not necessary to mount a GPU on the management device 720 to perform image recognition processing, and the energy consumption of the system can be reduced.
[0050] <Another example of the configuration of the surveillance camera system> In the configuration of the surveillance camera system 700 shown in Fig. 7, the number of brain-type devices 512 that can be implemented in a camera 520 and a small PC-sized housing is limited to 1 or 2, so it may not be possible to implement an SNN 101 that achieves the recognition accuracy required by the system requirements. In such cases, another configuration of the surveillance camera system (surveillance camera system 800) shown in Fig. 8 can be used.
[0051] The surveillance camera system 800 includes multiple cameras 520, a LAN switch 710, and a brain-type computer 510. The multiple cameras 520 and the brain-type computer 510 are connected to the LAN switch 710 via a LAN 730. In this configuration, since more brain-type devices 512 can be implemented in the brain-type computer 510, it is possible to implement an SNN of a larger scale and a minimum scale required to achieve the recognition accuracy required by the system requirements. In addition, by adopting such a configuration, inference processing can be performed by the low-power brain-type device 512 instead of the high-power GPU, so that the energy consumption of the system can be reduced.
[0052] As described above, according to this embodiment, an SNN can be configured to satisfy a predetermined number of brain-type devices and a predetermined recognition accuracy for an image to be inferred.
[0053] Also, for example, in a surveillance camera system, an SNN can be configured to satisfy a predetermined number of brain-type devices and a predetermined recognition accuracy in an image of an inference target. As a result, the system can be configured to satisfy the power consumption allowable depending on the installation location of the surveillance camera system.
[0054] The present invention can be widely applied to systems that perform image recognition of people, etc., such as recognizing a specific person in public places such as stations, airports, and streets, or recognizing a specific behavior of a worker at a workplace.
[0055] (II) Supplementary Note The above-described embodiment includes, for example, the following contents.
[0056] In the above embodiment, the processing unit reduces the number of connections of the basic block by one, but the present invention is not limited to this configuration. For example, the processing unit may reduce the number of connections of the basic block by multiple connections.
[0057] In the above embodiment, the processing unit reduces the number of convolutional layers by one, but the present invention is not limited to this configuration. For example, the processing unit may reduce the number of convolutional layers by multiple layers.
[0058] In the above-described embodiment, the output of information is not limited to display on a display. The output of information may be audio output from a speaker, output to a file, printing on a paper medium or the like by a printing device, projection on a screen or the like by a projector, or other form.
[0059] In addition, in the above description, information such as programs, tables, and files that realize each function can be placed in a storage device such as a memory, a hard disk, or an SSD, or in a recording medium such as an IC card, an SD card, or a DVD.
[0060] The above-described embodiment has the following characteristic configurations, for example.
[0061] (1) A spiking neural network configuration method for configuring a spiking neural network by converting a deep neural network that has completed learning of a deep neural network model, the deep neural network model being configured by connecting one or more basic blocks (e.g., basic block 102) each having one or more convolution layers (e.g., convolution layer 201), an input to the basic block (e.g., input 211) is connected to a first convolution layer (e.g., first convolution layer 201-1) of the basic block, and an output from the basic block (e.g., The method includes configuring a convolution layer of the deep neural network model so that a predetermined number of brain-type devices implementing the spiking neural network and a predetermined recognition accuracy of inference by the spiking neural network are satisfied, and a processing unit (e.g., a processing unit, a CPU, a processing device 511, a brain-type computer 510) configures the convolution layer of the deep neural network model so that an input (e.g., a first output 222) branched from the input (e.g., the second input 212) and an output (e.g., the first output 221) from the last convolution layer (e.g., the last convolution layer 201-2) of the basic block are combined.
[0062] In the above configuration, for example, the convolutional layer is configured to satisfy a predetermined number of brain-type devices and a predetermined inference recognition accuracy, so that a spiking neural network can be configured that satisfies the requirements of the number of brain-type devices permitted in the usage environment and the inference recognition accuracy.
[0063] (2) The deep neural network model is constructed by connecting a plurality of the basic blocks, and the processing unit determines whether the spiking neural network can be implemented in the predetermined number of brain-like devices (e.g., see S309), and if it determines that the spiking neural network cannot be implemented in the predetermined number of brain-like devices, reduces the number of connections of the basic blocks (e.g., see S411 and S403).
[0064] According to the above configuration, for example, by reducing the number of connections between basic blocks, the scale of the spiking neural network is reduced, making it possible to configure a spiking neural network that fits within a predetermined number of brain-type devices.
[0065] (3) The basic block has a plurality of convolutional layers, and when the processing unit determines that the spiking neural network can be implemented in the predetermined number of brain-type devices, if the recognition accuracy of the inference by the spiking neural network is less than the predetermined recognition accuracy, the processing unit reduces the number of convolutional layers constituting the basic block (e.g., see S411 and S405).
[0066] According to the above configuration, for example, the number of convolutional layers in the basic block can be reduced to improve the accuracy of inference recognition, making it possible to configure a spiking neural network that satisfies a predetermined level of accuracy in inference recognition.
[0067] (4) An image recognition device (e.g., image recognition device 500, brain-type computer 510) comprising a storage device for storing data and a processing device for processing the data stored in the storage device, further comprising a brain-type device (e.g., brain-type device 512) implementing a spiking neural network configured by the spiking neural network configuration method described in (1) above for performing recognition processing of a predetermined image, and an output unit (e.g., chipset 514, monitor, communication device) for outputting the results of the recognition processing by the brain-type device.
[0068] According to the above configuration, for example, image recognition processing can be performed by a brain-type device with low power consumption, thereby reducing the energy consumption of the image recognition device.
[0069] (5) The system includes one or more cameras (e.g., camera 520) connected to a brain-type device (e.g., brain-type device 512) that implements a spiking neural network configured by the spiking neural network configuration method described in (1) above and performs recognition processing of a specified image, a management device (e.g., management device 720) that manages the results of the recognition processing by the camera, and a switch (e.g., LAN switch 710) that connects the camera and the management device via a network.
[0070] According to the above configuration, for example, recognition processing by a brain-type device with low power consumption is possible for each camera, and power consumption can be reduced when the number of cameras is scaled up.
[0071] Furthermore, the above-described configurations may be modified, rearranged, combined, or omitted as appropriate without departing from the spirit and scope of the present invention.
[0072] It should be understood that items listed in the format "at least one of A, B, and C" can mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). Similarly, items listed in the format "at least one of A, B, or C" can mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). [Explanation of symbols]
[0073] 500...image recognition device, 510...brain-type computer, 512...brain-type device.
Claims
1. A spiking neural network configuration method for configuring a spiking neural network by converting a deep neural network that has completed learning of a deep neural network model, comprising: The deep neural network model is configured by connecting a plurality of basic blocks each having a plurality of convolution layers; An input to the basic block is connected to a first convolutional layer of the basic block, and an output from the basic block is an output obtained by combining an input branched from the input and an output from a last convolutional layer of the basic block; The processing unit configures a convolution layer of the deep neural network model so as to satisfy a predetermined number of brain-type devices implementing the spiking neural network and a predetermined recognition accuracy of inference by the spiking neural network; The processing unit includes: determining whether the spiking neural network can be implemented in the predetermined number of brain-type devices; If it is determined that the spiking neural network cannot be implemented in the predetermined number of brain-type devices, the number of connections of the basic blocks is reduced, When it is determined that the spiking neural network can be implemented in the predetermined number of brain-type devices, if the recognition accuracy of the inference by the spiking neural network is less than the predetermined recognition accuracy, reducing the number of convolution layers constituting the basic block. How to construct spiking neural networks.
2. An image recognition device comprising a storage device that stores data and a processing device that processes the data stored in the storage device, A brain-type device that implements a spiking neural network that performs recognition processing of a predetermined image, the spiking neural network being configured by the spiking neural network configuration method according to claim 1; an output unit that outputs a result of the recognition processing by the brain-type device; An image recognition device comprising:
3. One or more cameras connected to a brain-type device that implements a spiking neural network that performs recognition processing of a predetermined image, the brain-type device being configured by the spiking neural network configuration method according to claim 1; a management device for managing a result of the recognition process performed by the camera; a switch that connects the camera and the management device via a network; A surveillance camera system comprising:
Citation Information
Patent Citations
Control device for array comprising neuromorphic element, calculation method of discretization step size, and program
JP2019046072A