System and method for occupancy detection using frame-based sensors
By using event-based data acquisition and spiking neural network processing, the problem of high power consumption during continuous operation of frame-based sensors was solved, achieving low-power real-time occupancy detection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INNATERA NANOSYSTEMS BV
- Filing Date
- 2024-07-15
- Publication Date
- 2026-07-23
AI Technical Summary
Existing frame-based sensors suffer from high power consumption during continuous operation when performing occupancy detection, especially when the occupancy level is stable and still requires a large amount of power.
An event-based data acquisition and dataset assembly method is adopted, combined with a spiking neural network, to acquire and process event-based data through event-based sensors, thereby reducing unnecessary data acquisition and computation.
It enables high-performance presence sensing applications with limited power budgets, reducing power consumption while maintaining real-time monitoring capabilities.
Smart Images

Figure 2026524703000001_ABST
Abstract
Description
Technical Field
[0001]
[0001] The present invention relates to the field of occupancy detection, particularly to methods and devices for generating and / or using event-based data from an image sensor to be processed via a spiking neural network.
Background Art
[0002]
[0002] Occupancy detection is defined as the task of identifying whether a person is present in a particular room, space, or environment, or whether the occupancy state of the room, space, or environment changes. Further, additional information, such as the count of people or objects, may be required from a system capable of performing occupancy detection.
[0003] [[ID=第十六条]]
[0003] Applications for occupancy detection include building management, where automatic actions can be taken (e.g., turning off ambient lighting or air conditioning, activating security cameras or alarms) when no humans are present in a building's room or area, or staged actions can be taken when several humans are present (e.g., setting the air conditioner fan to medium when three humans are present, but to maximum when five or more humans are present).
[0004]
[0004] Typically, frame-based sensors are used to perform occupancy detection. These are sensors that capture data in discrete frames or snapshots, similar to how a camera captures an image. Frame-based sensors can utilize various technologies such as infrared (IR), ultrasonic, microwave, or even visual cameras, and capture data in frames or snapshots at regular intervals. For example, an IR sensor can take a snapshot of the thermal signature of a room, while a visual camera can capture an image of the monitored area.
[0005]
[0005] The captured frames are then analyzed to detect changes in occupancy within the monitored area. Depending on the sensor technology used, this analysis may involve detecting spatial and temporal features, movement, thermal signatures, or other relevant parameters associated with human presence. Data from frame-based sensors, such as frames of RGB values, are typically processed using algorithms designed to identify patterns associated with occupancy. Machine learning algorithms may also be used to improve the accuracy of occupancy detection over time.
[0006]
[0006] Frame-based sensors may be used to perform occupancy detection by using minimal preprocessing, data buffering, and coupling of frame-based sensors to a neural network for processing.
[0007]
[0007] Frame-based sensors may be configured to continuously capture data, such as images or sensor readings, at regular intervals to monitor occupancy in a space. This constant data collection ensures that any changes in occupancy are detected immediately. The data captured by the sensor is then fed into a neural network for analysis. The neural network processes the data to identify patterns associated with occupancy, such as movement, thermal signatures, or other relevant features. Since occupancy detection often requires real-time responsiveness, the neural network must process the data quickly to provide timely feedback on occupancy status. This real-time analysis is essential for applications such as smart building automation, security systems, and energy management.
[0008]
[0008] In order to maintain continuous monitoring and real-time analysis, both the frame-based sensors and the neural network must remain "always on," meaning they are always operational and actively processing data. This constant operation consumes processing power and energy, even during periods when the occupancy level is stable.
[0009]
[0009] Continuous operation of frame-based sensors and neural networks can lead to significant power consumption, particularly in battery-powered devices or systems with stringent energy efficiency requirements. [Overview of the Initiative]
[0010]
[0010] In order to solve the above-mentioned problems, the subject matter of the claims of the present invention is proposed.
[0011]
[0011] The proposed solution introduces a method for performing data acquisition and dataset assembly to enable training of high-performance networks. This method enables the realization of high-performance presence sensing applications that can operate with a limited power budget.
[0012]
[0012] According to a first aspect of the present disclosure, a method for detecting occupancy in a visual scene captured by an image sensor is disclosed. The method uses a spiking neural network. The method may include: acquiring event-based data, wherein the event-based data may be a representation of input from an image sensor in the form of one or more discrete events, each event may represent a change in a visual scene that exceeds a predetermined threshold; encoding the event-based data as input spikes, wherein the input spikes may be configured to serve as input for a spiking neural network; processing the input spikes using a spiking neural network, wherein the spiking neural network may produce one or more output spikes by processing the input spikes; decoding the output spikes into one or more classification results; and detecting occupancy in a visual scene using the classification results.
[0013]
[0013] According to one embodiment of the first aspect, the image sensor may be an event-based image sensor that captures event-based data when there is a change in the visual scene that exceeds a predetermined threshold. Preferably, the event-based image sensor captures changes at the individual pixel level and / or region pixel level. More preferably, when a change is detected in a particular region of the visual scene that exceeds a predetermined threshold, the event-based image sensor may transmit event-based data representing the location of the change, optionally along with information about the rate and timing of the change.
[0014]
[0014] According to one embodiment of the first aspect, the image sensor may be a frame-based image sensor that captures images in a frame-by-frame manner, where each frame represents a complete snapshot of the visual scene at a particular point in time. The step of acquiring event-based data may further include preprocessing the image captured from the frame-based image sensor by calculating the difference between the current frame and a background frame, wherein the background frame may be obtained from one or more frames previously captured by the image sensor, checking whether the difference is a change in the visual scene that exceeds a predetermined or immediately calculated threshold, and if the change exceeds the predetermined threshold, generating event-based data based on the difference.
[0015]
[0015] According to one embodiment of the first aspect, event-based data may include the position, time, size, shape, and / or intensity of a change in a visual scene.
[0016]
[0016] According to one embodiment of the first aspect, a predetermined threshold may be either a positive threshold or a negative threshold. A positive threshold may be used to check whether an increase in intensity in a particular area of a visual scene is an event. A negative threshold may be used to check whether a decrease in intensity in a particular area of a visual scene is an event.
[0017]
[0017] According to one embodiment of the first aspect, the step of encoding event-based data into spike sequences may include mapping the location of changes in a visual scene to corresponding neurons, mapping the size of changes in a visual scene to the firing rate of corresponding neurons, mapping the shape of changes in a visual scene to a specific subset of neurons corresponding to that shape, mapping the intensity of changes in a visual scene to spike timing, frequency, or amplitude, and / or mapping the time of changes in a visual scene to spike timing.
[0018]
[0018] According to one embodiment of the first aspect, a spiking neural network may consist of multiple layers of spiking neurons connected in a feedforward manner, allowing input spikes to flow from the input layer to the output layer. Recurrent and / or feedback connections may be present, connecting the layers to themselves or to upstream layers, respectively.
[0019]
[0019] According to one embodiment of the first aspect, the step of decoding one or more output spikes may include counting the spiking activity of the output neuron, and / or calculating a maximum or softmax function in order for the network to select a class to which the input is predicted to belong, and / or analyzing the membrane potential of the output neuron, preferably by calculating a first derivative of the membrane potential that represents the rate of change of the spike count reaching the output neuron over time or under the area of the membrane potential curve, in order to extract information about the timing and dynamics of the output neuron.
[0020]
[0020] According to one embodiment of the first aspect, the method may further include recording event timestamps, wherein the timestamps represent the features they represent or the timing of the events in relation to an input stimulus, correlating the event timestamps with features represented by output neurons, and determining, based on the correlation, when and how the network responds to a particular input spike.
[0021]
[0021] According to one embodiment of the first aspect, the method may further comprise determining the time-averaged count of a feature in an inter-feature interval and / or time interval based on a timestamp, and measuring the repeatability of a feature using the inter-feature interval and / or the intensity of a feature in a time interval using the time-averaged count.
[0022]
[0022] According to one embodiment of the first aspect, the method may further comprise taking action based on detected occupation, wherein the action may be either internal or external. Internal actions are actions that affect the method by changing the control parameters of an image sensor or a spiking neural network, while external actions may be actions that affect the environment of a visual scene.
[0023]
[0023] According to a second aspect of the present disclosure, an occupancy detection device for occupancy detection in a visual scene captured by an image sensor is disclosed. The occupancy detection device may comprise a processor, an encoder, a spiking neural network, and a decoder. The processor may be configured to acquire event-based data. The event-based data may be a representation of input from the image sensor in the form of one or more discrete events, each event representing a change in the visual scene that exceeds a predetermined threshold. The encoder may be configured to encode the event-based data as input spikes, which are configured to serve as inputs for a spiking neural network. The spiking neural network may be configured to produce one or more output spikes by processing the input spikes. The decoder may be configured to decode the output spikes into one or more classification results. The processor may be configured to use the classification results to detect occupancy in the visual scene.
[0024]
[0024] According to one embodiment of the second aspect, the occupancy detection device may further include an image sensor. The image sensor may be an event-based image sensor or a frame-based image sensor.
[0025]
[0025] An event-based image sensor can capture event-based data when there is a change in a visual scene that exceeds a predetermined threshold. Preferably, the event-based image sensor captures changes at the individual pixel level and / or region pixel level. More preferably, when a change exceeding a predetermined threshold is detected in a particular region of the visual scene, the event-based image sensor can transmit event-based data representing the change, along with information about the location and timing of the event.
[0026]
[0026] A frame-based image sensor can capture images in a frame-by-frame manner, where each frame represents a complete snapshot of the visual scene at a particular point in time. The processor is further configured to preprocess the images captured from the frame-based image sensor by calculating the differences between successive frames, check whether the differences are changes in the visual scene that exceed a predetermined threshold, and generate event-based data based on the differences when the changes exceed the predetermined threshold.
[0027]
[0027] According to an embodiment of the second aspect, the event-based data comprises the position, time, size, shape, and / or intensity of the change in the visual scene.
[0028]
[0028] According to an embodiment of the second aspect, the encoder is configured to map the position of the change in the visual scene to corresponding neurons, map the size of the change in the visual scene to the firing rate of the corresponding neurons, map the shape of the change in the visual scene to a specific subset of neurons corresponding to that shape, map the intensity of the change in the visual scene to the timing, frequency, or amplitude of the spikes, and / or map the time of the change in the visual scene to the spike timing, to map the features of the event-based data to the spiking activity of the neurons in the spiking neural network.
[0029]
[0029] According to an embodiment of the second aspect, the processor can be configured to take an action based on the detected occupancy. The action can be either internal or external. An internal action is an action that affects the method by changing the control parameters of the image sensor or the spiking neural network, while an external action can be an action that affects the environment of the visual scene.
[0030]
[0030] Next, embodiments are described by way of example only with reference to the accompanying drawings, where corresponding reference numerals indicate corresponding parts.
Brief Description of the Drawings
[0031]
[0031] [Figure 1] FIG. 1 shows a schematic diagram of an embodiment of an occupancy detection system according to the present invention.
[0032]
[0032] [Figure 2] FIG. 2 shows a schematic diagram of an embodiment of an occupancy detection method according to the present invention.
Embodiments for Carrying Out the Invention
[0033]
[0033] Hereinafter, certain specific embodiments are described in more detail. However, it should be recognized that these embodiments should not be construed as limiting the scope of protection for the present disclosure.
[0034]
[0034] FIG. 1 shows a schematic diagram of an embodiment of an occupancy detection system according to the present invention.
[0035]
[0035] The occupancy detection system 100 may include an imaging sensor 101, which can be, for example, a low-resolution sensor or an event-based sensor, and a processing device 102.
[0036]
[0036] The image sensor captures a visual scene. A visual scene refers to the entirety of what is captured within the field of view of the image sensor at a given moment. It encompasses all objects, backgrounds, colors, textures, and lighting conditions present in the image. In other words, a visual scene is a specific area monitored by the sensor where the presence or absence of an object or person should be determined, and preferably, the visual scene is captured by the imaging sensor and provides input data that a spiking neural network will process and analyze to determine whether and where an occupant is present. In this context, the field of view is the spatial extent captured by the image sensor and includes the edges and corners of the scene. Objects in the field of view can be any tangible items in the scene, such as furniture, vehicles, or individuals. The background is the part of the scene that serves as the background for an object, such as a wall, floor, or ground. Textures and colors comprise the surface characteristics and color variations of objects and backgrounds in the scene. Lighting conditions refer to the lighting in the scene, including natural and artificial light sources, and their effects on visibility and shadows.
[0037]
[0037] Low-resolution sensors capture images with a lower pixel count compared to high-resolution sensors. For example, low-resolution sensors may have a resolution of, for example, VGA (640 x 480 pixels) or even lower. These sensors are often used in applications where high image detail is not important, or where constraints such as cost, power consumption, or processing requirements are limiting factors.
[0038]
[0038] Event-based sensors, also known as event-driven sensors or neuromorphic sensors, operate differently from conventional frame-based sensors. Instead of capturing entire frames at fixed intervals, they only capture and transmit information when there is a significant change in the scene. Event-based sensors, for example, asynchronously detect and report changes at the individual pixel level in brightness, meaning they respond only to changes in the scene. When they detect a change in a pixel that exceeds a predetermined threshold, they send an event signal indicating the change, along with information about the location and timing of the event. Event-based sensors are particularly well-suited for applications requiring low latency, high speed, and high dynamic range sensing.
[0039]
[0039] The processing device 102 includes one or more peripheral interfaces 103 for transferring data from the imaging sensor 101 to the processing device 102. Examples include a serial peripheral interface (SPI), an inter-integrated circuit (I2C), a MIPI (Mobile Industry Processor Interface), a parallel interface, a USB port, a Bluetooth® interface, an HDMI® port, an Ethernet® port, and / or low-voltage differential signaling (LVDS). The transfer may be performed directly or indirectly, for example, via a server, over a wired or wireless connection.
[0040]
[0040] The processing device 102 further comprises a microprocessor 104, a memory 105, an interconnection structure 109, and a spiking neural network (SNN) accelerator.
[0041]
[0041] The microprocessor 104 organizes the operation of the system and performs data acquisition and general calculations and data manipulation. The microprocessor 104 may preprocess image data to enhance relevant features and / or reduce noise. The image sensor 101 may be controlled by the microprocessor 104 by configuring the image capture rate, exposure duration, low power or high power mode, for example, by writing to configuration registers or memory locations contained within the sensor. The microprocessor 104 places the data from the image sensor 101 into memory 105. The data from the image sensor 101 may be RGB, luminance / grayscale, event, or any other suitable format.
[0042]
[0042] RGB data represents an image using three primary colors: red, green, and blue. Each pixel in the image is typically represented by three values corresponding to the intensity of these three color channels. RGB data is typically represented as a matrix of pixel values. Luminosity or grayscale data represents an image in terms of overall brightness or intensity without color information, where the data is represented as a single channel. Event-based data comprises data for individual events triggered by significant changes in the scene, such as changes in brightness or motion. Each event typically includes information such as pixel position, timestamp, and polarity (indicating whether the change was an increase or decrease in intensity). Event-based sensors provide high temporal resolution and enable low-latency real-time response to dynamic scenes.
[0043]
[0043] The memory 105 comprises storage components used to temporarily hold data, instructions, and intermediate results during an image processing task. The memory 105 may include, for example, RAM (random access memory), cache memory, and / or non-volatile memory. The memory 105 may be used to store, for example, raw image data captured by the imaging sensor 101, processed images generated during an image processing algorithm, parameter values to be used by other components of the system 100, and output from the SNN accelerator.
[0044]
[0044] The SNN accelerator may include an encoder 106, a spiking neural network 107, and a decoder 108.
[0045]
[0045] The encoder 106 is the initial stage of the SNN accelerator. Its purpose is to convert the input data into a format suitable for processing by the spiking neural network. In the context of a spiking neural network, the input data can be sensory data such as images, audio signals, or other types of time-varying signals. The encoder converts this continuous or analog input into a discrete representation of spikes, which are the basic unit of communication in a spiking neural network. The encoder outputs the acquired spike sequence to the spiking neural network 107.
[0046]
[0046] The spiking neural network (SNN) 107 is the core computing unit of the accelerator. It receives a sequence of spikes obtained from the encoder 106 as input. It consists of interconnected neurons that communicate with each other through the spikes. Each neuron in the SNN integrates the input spikes over time and generates output spikes based on certain activation rules, typically modeled after biological neurons. The neurons are connected via synaptic elements, which have certain weights that identify how strong the connections between neurons are. The SNN can process the input spikes through layers of neurons and perform computations such as feature extraction, pattern recognition, or classification.
[0047]
[0047] The decoder 108 is the final stage of the SNN accelerator. Its purpose is to interpret the output spikes generated by the spiking neural network 107 and produce meaningful output based on the task being performed. Depending on the application, the decoder may perform tasks such as classification, regression, or decision-making based on the spike patterns generated by the SNN. Essentially, the decoder converts the neural activity encoded in the spikes back into a form that can be understood or used by the microprocessor 104.
[0048]
[0048] There are several reasons why using an SSN is beneficial when using event-based data as its input, which comes from, for example, an event-based sensor (or a frame-based sensor where the data is preprocessed in order to acquire event-based data).
[0049]
[0049] Event-based data is inherently asynchronous because, unlike conventional cameras which capture frames at fixed intervals, it only reports changes in a scene (event). SNNs are well-suited to processing such asynchronous data because they operate on the principle of spike timing, mimicking the behavior of neurons in the brain that fire in response to input spikes.
[0050]
[0050] SNN has event-driven processing, which means it can efficiently handle low-latency data such as event-based data.
[0051]
[0051] SNNs are inspired by the energy-efficient processing of the brain and require less computational resources compared to conventional artificial neural networks (ANNs). This makes them an ideal combination for event-based sensors, which consume power only when there is a change in the scene.
[0052]
[0052] Event-based data is generally sparse data because events only occur when there is significant activity in the scene. SNNs are accustomed to processing information in a sparse and distributed manner, so they handle sparse data well naturally.
[0053]
[0053] Event-based sensors are often more robust to noisy and high-dynamic-range scenes compared to conventional frame-based sensors. SNNs have the ability to encode and process information at the timing of spikes, and can further enhance this robustness by focusing on the most relevant information while filtering out noise.
[0054]
[0054] Figure 2 shows a schematic diagram of one embodiment of the occupation detection method 200 according to the present invention.
[0055]
[0055] System 100 may be configured to perform a series of operations or steps of method 200. These operations 201-206 may be performed as phases in an operation pipeline.
[0056]
[0056] The first step of the method comprises data acquisition 201. Proper data acquisition is important to maximize the performance of the solution. The captured data is transferred from the image sensor 101 to the processing device 102 via one or more peripheral interfaces 103. As described above, this transfer can occur through various types of interfaces. Upon receiving the image data, the microprocessor 104 may process the data to perform tasks such as image compression (if necessary), applying filters, adjusting the color balance, and performing any other desired image operations. The image data, such as events, frames, or sequences, can then be stored in the memory 105 of the processing device 102.
[0057]
[0057] If the image sensor is an event-based sensor, steps 202 and 203 below may be skipped; otherwise, the following steps are performed in order.
[0058]
[0058] The second step of the method comprises preprocessing the image data into event-based data 202. The difference between consecutive or subsequent frames is calculated such that N frames are converted into a sequence of N-1 difference frames. This can be done using frame difference, or any other suitable method, such as a method involving frame difference.
[0059]
[0059] Frame difference is performed by calculating the difference between two consecutive frames captured by the image sensor. To simplify processing, both images may be converted to grayscale. This can reduce the calculations required and simplify the comparison process. Next, the absolute difference between corresponding pixel values in the two grayscale images is calculated. This operation yields a new image in which each pixel represents the absolute difference in intensity between the same pixel positions in the two input images.
[0060]
[0060] The third step of the method comprises an encoding step 203. Time contrast / level sampling encoding is used to convert differences into events. Positive changes between subsequent frames are encoded as events / spikes if the difference is greater than a predetermined threshold, which can be trained or engineered. Conversely, events capturing negative changes will be generated if the difference exceeds a predetermined negative threshold.
[0061]
[0061] Accordingly, a threshold can be applied to an image showing an absolute difference in order to identify areas in the image where the change in intensity exceeds a certain positive or negative threshold. This step helps to distinguish significant changes from noise. Next, connected components can be identified in the thresholded image. Each connected component represents an area where a change occurred between two frames. Each connected component can be identified as an “event”. For each identified event, information such as the location, time, size, shape, and intensity of the change can be extracted. This information can be stored or transmitted as event-based data.
[0062]
[0062] Next, the encoder encodes event-based data such as the location, timing, size, shape, and intensity of changes into spike trains, which can be used as input to an SNN. Typically, this may involve mapping each of these features to a specific characteristic of the spike train.
[0063]
[0063] Spatial information, such as the location of changes in an image, can be encoded using the spatial positions of neurons in the input layer of a spiking neural network. Each neuron in the input layer corresponds to a specific location in the field of view of the image sensor. When an event occurs at a specific location in the image, the neuron corresponding to that location generates a spike.
[0064]
[0064] Encoding the time of change into a spike sequence involves representing time information through spike timing. To encode the timing data of events, for example, time coding, rate coding, group coding, and / or relative timing coding may be used.
[0065]
[0065] In time coding, the timing of spikes may be used to encode the time of change relative to a reference point. Each spike generated by a neuron corresponds to an event occurring at a specific time. The precise timing of the spike relative to the start of the event can encode time information such as the time of occurrence and the duration of the change. Earlier spikes may represent changes that occurred earlier in time, while later spikes indicate changes that occurred later. In rate coding, the firing rate of a neuron may be used to encode time information. Higher firing rates may represent changes that occur over shorter durations or more frequently in time. Lower firing rates may represent changes that occur over longer durations or less frequently. In population coding, a population of neurons may be used to collectively encode time information. Different neurons or groups of neurons may represent different time intervals or time features. The activation patterns across a population of neurons can encode complex time information about the change. Finally, relative timing coding may be used to encode the relative timing of changes between different spatial locations or features. Neurons representing different locations or features in an image may fire at different times relative to each other, encoding the temporal relationships between changes.
[0066]
[0066] The size of change in an image can be encoded using the firing rate of neurons. Neurons representing larger areas or objects in an image may have a higher baseline firing rate, while neurons representing smaller areas may have a lower firing rate. The firing rate of neurons can be adjusted based on the size of change detected in their receptive fields.
[0067]
[0067] Encoding the shape of changes in an image into spike sequences can be more complex. One approach is to use a population of neurons with spatially arranged receptive fields that collectively represent different shapes. When a change with a particular shape is detected, a specific subset of neurons in the input layer corresponding to that shape is activated.
[0068]
[0068] Intensity information, such as the magnitude of changes in pixel values, can be encoded using spike timing or frequency. Higher intensity changes may lead to higher spike rates or earlier spike timing, while lower intensity changes result in lower spike rates or delayed spikes. Alternatively, intensity information can be encoded using spike amplitude, with higher intensity leading to spikes with higher amplitude.
[0069]
[0069] Overall, the encoding process involves mapping various features of event-based data to the spiking activity of neurons in the input layer of a spiking neural network. This allows the network to process and interpret event-based data in a format suitable for spike-based computation. The specific encoding scheme selected may vary depending on the requirements of the task and the characteristics of the input data.
[0070]
[0070] Hereinafter, specific examples of the first, second, and third steps are given. In the first step, the processing device may acquire data from, for example, a frame-based image sensor. These may be frames having a specific width and height, for example, 400 x 400 pixels, and a specific color model, for example, an RGB color model.
[0071]
[0071] Next, these frames are preprocessed into event-based data. In this particular example, at any given point in time, many past frames are resized (if necessary) into frames containing only their respective luminances, from which the background image is obtained as the average of these luminance-only frames.
[0072]
[0072] For example, the last 30 frames in the original 400x400 RGB format are resized into 64x64 luminance-only frames by discarding the 8 border pixels and selecting every 6th pixel. This can be done by discarding the 8 border pixels and reducing the 400x400 image to 384x384 pixels. The cropped image is then downsampled by selecting every 6th pixel to obtain a 64x64 image. Next, each frame is converted from RGB to a grayscale (luminance) image. To convert an RGB image to grayscale (luminance), typically a weighted sum of the RGB channels is used to represent the intensity of each pixel. The conversion to luminance can also be done beforehand. Finally, the average of these 30 resized luminance frames is calculated to obtain the background image.
[0073]
[0073] As the next step, the latest incoming frame is resized in the same way as the previous frames. For example, the latest incoming frame in its original 400x400 RGB format is resized by discarding the 8 border pixels, selecting one every 6th to obtain a 64x64 image, and converting it to a luminance image.
[0074]
[0074] It is also possible to use the original frame, for example, 400x400RGB, without downscaling or brightness conversion. This keeps all original data channels in place and reduces the time required for preprocessing, but imposes higher demands on subsequent inference steps.
[0075]
[0075] An absolute integer difference is applied between the processed frame and the previously acquired background frame. One or more iterations of the dilation operation are applied to the resulting difference frame.
[0076]
[0076] First, the absolute integer difference between the processed frame and the previously acquired background frame is obtained. This operation calculates the absolute difference between corresponding pixels in the two frames. Mathematically, for each pixel (x,y) in the background frame and the processed frame, the absolute integer difference is calculated as follows: difference(x,y) = |processed frame(x,y) - background frame(x,y)|. After calculating the absolute integer difference, the difference frame is obtained. This frame highlights areas where significant changes have occurred between the current processed frame and the background frame. Areas with little change will have a low intensity value, while areas with significant change will have a high intensity value.
[0077]
[0077] As described above, the difference frame can also be determined as the difference between the current frame and a previous frame, such as the previous frame, and as a result, there is no need to perform a rolling average calculation. The frame obtained by calculating the previous frame or the rolling average can also be quantized before calculating the difference with the current frame. Here, quantization means rounding the pixel values to the nearest level-crossing point, which is defined as the nearest integer input to a threshold with a potential constant offset value.
[0078]
[0078] After obtaining the difference frame, the next step is to apply one or more iterations of the dilation operation to this frame. Dilation is a morphological operation that expands the area of foreground pixels (in this case, the area with a significant difference from the background) based on a structured element (kernel). A structured element (kernel) can be selected to determine the neighborhood used in the dilation operation. Common choices include rectangular or circular kernels. For example, a 2x2 kernel can be used, which is a 2x2 matrix where all elements are 1. When applying dilation to a difference frame, the goal is to enhance the area where this difference exceeds a certain threshold. The dilation operation will expand these areas of significant change, effectively "emphasizing" them compared to the surrounding area. A kernel can be placed at each pixel of the difference frame. If a kernel overlaps with any foreground pixel (if its intensity value exceeds a certain threshold), the corresponding pixel in the output image is set to the foreground.
[0079]
[0079] Finally, if necessary, the dilation operation may be applied multiple times. Each iteration further expands the foreground region. Each application of the dilation operation expands the foreground objects by the size and shape of the kernel, giving a more pronounced effect. By applying the dilation operation multiple times, these foreground objects can be further expanded, which can be useful in various image processing tasks. For example, two iterations of dilation may be applied by using a 2x2 kernel.
[0080]
[0080] Next, the encoding step is performed as the third step of the method in this example.
[0081]
[0081] The encoding step requires defining a threshold and the number of extended time steps. Thus, a threshold can be selected that determines how integer values will be encoded into spikes. This threshold is used to normalize the values for encoding. Next, the total number of extended time steps for encoding the spikes is defined.
[0082]
[0082] For previously acquired processed images, row and column sums are applied to obtain integer values representing the sum of pixel intensity across rows and columns, respectively. For example, for previously acquired 64x64 processed images (e.g., dilated difference frames), row and column sums are applied to obtain 128 integer values.
[0083]
[0083] Next, these integer values are encoded into spikes. To do this, normalization is performed by dividing each of the integer values by a selected threshold. For example, each of the 128 values is divided by the selected threshold, and as a result, an integer quotient of the division is obtained.
[0084]
[0084] Next, this integer quotient is used to select how many consecutive time steps (out of the given total extended time steps) will contain the spike, and the remainder of the given total extended time steps (if any) will represent the absence of the spike.
[0085]
[0085] Therefore, for example, for each of the 128 integer values, a spike pattern is generated, where the quotient determines how many consecutive time steps (out of the total extended time steps) will have spikes. The remaining time steps will represent the absence of spikes. If the extended time steps are set to a specific number of time steps, it is possible to determine how many of that specific number of time steps will contain spikes.
[0086]
[0086] Alternatively, the quotient of the row and column sums divided by a threshold may be used as the rate for other types of encoding (periodic rate, Poisson rate, or integral & fire encoding), and thus, by using the aforementioned techniques, spike sequences can be generated over extended time steps.
[0087]
[0087] Next, this process may be applied to a specific number of consecutive frames, for example 30, to provide input to the spiking neural network.
[0088]
[0088] The fourth step of the method comprises processing event-based data using a spiking neural network 204.
[0089]
[0089] A pre-trained spiking neural network (SNN) receives events produced by appropriate preceding steps as input. The SNN consists of multiple layers of spiking neurons connected in a feedforward manner, allowing input spikes to flow from the input layer to the output layer. Additional recurrent or feedback connections may be added depending on the final complexity of the application, and connecting layers to themselves (recurrent) or upstream layers (feedback) will have the effect of increasing the network's ability to identify longer-term connections between features.
[0090]
[0090] A spiking neural network is terminated at one or more output nodes / neurons, and its activity (the spiking events produced, or internal states if observable) is interpreted as the instantaneous presence of abstract features in which they are sensitive. For example, certain neurons may be tuned to respond to the contours of a person, while other neurons may be tuned to respond most responsively to features related to a pet or a curtain.
[0091]
[0091] In order to evaluate the network performance, test / validation data should be compiled in the same manner as during training.
[0092]
[0092] A fifth step of the method comprises decoding the state of the output neurons of the SNN. Decoding the state of the output neurons of a spiking neural network (SNN) involves interpreting the spiking activity of these neurons to extract meaningful information about the input data or the network's response to the input. Each output neuron may represent a specific feature, class, or category that the network is trained to recognize or respond to. The spiking activity of these neurons encodes information about the network's decision or response based on the input data.
[0093]
[0093] The state of the output neurons refers to the spiking activity pattern of the output neurons over time. The state of the output neurons may be processed by thresholding in the activity domain or first derivative, and the timestamps of the events are recorded in relation to the features they represent.
[0094]
[0094] The step of decoding one or more output spikes may, for example, include counting the spiking activity of each output neuron. Another option is to calculate the value obtained after applying an argmax or softmax function to the (raw) spike count for each output neuron, where each output neuron may represent an output class of the spiking neural network in order for the network to make a selection of the class to which the inputs are predicted to belong. Another option is to analyze the membrane potential of the output neuron, preferably by calculating the first derivative of the membrane potential, which represents the rate of change of the spike count reaching the output neuron over time or under the area of the membrane potential curve, in order to extract information about the timing and dynamics of the output neuron.
[0095]
[0095] Thresholding in the activity domain involves setting a threshold for the spiking activity of output neurons. Neurons above the threshold are considered activated or "firing" and indicate a positive response to an input stimulus or the presence of a specific feature. This thresholding process helps to distinguish significant responses in the network from background activity or noise.
[0096]
[0096] Alternatively, the first derivative of the spike train can be analyzed to extract information about the timing and dynamics of spiking activity. The first derivative represents the rate of change in the spike count over time and indicates periods of increased or decreased activity. Peaks or changes in the first derivative may correspond to significant events or transitions in the neuron's spiking activity.
[0097]
[0097] Once the spiking activity of the output neurons has been processed through thresholding or first derivative analysis, event timestamps are recorded. These timestamps represent the features they represent or the timing of the events in response to the input stimulus. By correlating the event timestamps with the features represented by the output neurons, it is possible to determine when and how the network responds to a particular input stimulus or detects a particular feature in the data.
[0098]
[0098] Evidence of which tokens / features are produced over time can be stored in a memory-bounded and efficient structure (e.g., a circular buffer) that overwrites itself during normal operation. A circular buffer is a fixed-size memory that overwrites older data as new data comes in, ensuring that memory usage remains constant.
[0099]
[0099] The interval between tokens can serve as a measure of the repeatability of a feature, while the time-averaged count of a feature represents the intensity of the feature over a short time interval. Shorter intervals between tokens indicate the frequent occurrence of the feature and suggest high repeatability, while longer intervals suggest lower repeatability. Higher time-averaged counts indicate higher intensity of the feature during that time interval, while lower counts suggest lower intensity.
[0100]
[0100] The sixth step of the method comprises taking a specific action 206 based on the decoded result. The time between features, as well as the number of each feature predicted by the SNN, are used to make a higher level prediction (e.g., about 5 people are in the room) and so that an appropriate action may be taken. The action may be internal or external. An internal action is an action that affects the system (e.g., changing the exposure time of a camera), while an external action is an action that affects the environment (e.g., turning off the lights in the room). The action is then recorded as a mechanism for the system to take into account changes in how the environment is perceived. The action may be initiated or controlled by the microprocessor 104 of the processing device. For example, the microprocessor 104 may determine specific parameter values that define a particular action 206. The decoded result of the SNN may also be transferred directly from the processing device 102 to a separate device. Commands or decoded results may be transferred to different devices using the peripheral interface 103.
[0101]
[0101] The following are some considerations for data collection and use.
[0102]
[0102] The training and validation data used to train and validate the spiking neural network should be recorded in a continuous sequence so that realistic time patterns can be extracted from the data. This process involves recording multiple frames using the same or similar image sensor mounting positions and under gradually changing environmental conditions. The number of frames to be recorded in a sequence is at least the length required for the neural network to produce an output and may be influenced by the complexity of the task in each output, the network architecture, or the internal confidence target threshold.
[0103]
[0103] Data labeling should be done frame by frame or in an event-based manner, and when a label changes within the recording, the time and the new label are registered for that particular frame.
[0104]
[0104] Semi-automatic or automatic labeling may be used in conjunction with larger models (e.g., larger neural networks, optionally combined with some pre- and post-processing) that provide labels that are considered ground truth (e.g., labels that are considered correct for the dataset). Labeling of the dataset can be iteratively improved by examining correct and incorrect predictions of smaller neural networks through random sampling.
[0105]
[0105] Recording should be done in a setting that will yield more accurate / less noisy results in order to provide a basis for introducing faults or noise, or simply dropping out data.
[0106]
[0106] The data acquisition / assembly method enables optimization of the entire solution pipeline, including heuristics, any necessary preprocessing, and neural networks to be trained.
[0107]
[0107] The methods and processing pipelines presented herein enable extremely low-power, vision-based occupancy detection, as well as derivative presence / movement detection where power is the primary constraint. This performance is made possible by evidence-based computation flows and the closure of sensor computation loops.
[0108]
[0108] The present invention can be applied to a number of technical fields. For example, in building management by controlling the wake-up and alarm of security cameras, building lighting, heating and other services.
[0109]
[0109] This study describes a method for data collection that enables the creation of an optimized pipeline so that the classification of occupancy status in any environment can be performed by forcing chunks of data to have smoothly changing environmental conditions to leverage background similarity during training.
[0110]
[0110] This study describes a method for producing occupation classification (whether a person is present or absent) using level sampling of inputs, which are presented directly in event format or converted into events / spikes and processed by a spiking neural network.
[0111]
[0111] The presented pipeline is lightweight enough and tunable to enable the application of occupation detection with a very low power budget, as preprocessing involves the calculation of pairwise differences and threshold comparisons, as well as the storage of up to two full data frames and encoded sparse events / spikes.
[0112]
[0112] Note that any of the features of the embodiments disclosed herein can be appropriately combined.
Claims
1. A method for detecting occupancy in a visual scene captured by an image sensor, wherein the method uses a spiking neural network, The acquisition of event-based data, wherein the event-based data is a representation of the input from the image sensor in the form of one or more discrete events, each event representing a change in the visual scene that exceeds a predetermined threshold. The event-based data is encoded as input spikes, wherein the input spikes are configured to function as inputs for the spiking neural network. The spiking neural network is used to process the input spikes, wherein the spiking neural network generates one or more output spikes by processing the input spikes. Decoding the aforementioned output spikes into one or more classification results, Using the classification results, detect the occupation in the visual scene, A method that includes [a certain feature].
2. The method according to claim 1, wherein the image sensor is an event-based image sensor which captures event-based data when there is a change in the visual scene that exceeds a predetermined threshold, preferably the event-based image sensor captures changes at the individual pixel level and / or region pixel level, and more preferably, when a change is detected in a particular region of the visual scene that exceeds the predetermined threshold, the event-based image sensor transmits event-based data representing the location of the change, optionally along with information regarding the rate of the change and the timing of the change.
3. The image sensor is a frame-based image sensor that captures images in a frame-by-frame manner, where each frame represents a complete snapshot of the visual scene at a specific point in time, and the step of acquiring the event-based data is: Preprocessing the captured image from the frame-based image sensor by calculating the difference between the current frame and the background frame, wherein the background frame may be obtained from one or more frames previously captured by the image sensor. The process involves checking whether the difference represents a change in the visual scene that exceeds a predetermined or immediately calculated threshold, If the aforementioned change exceeds the predetermined threshold, event-based data is generated based on the difference. The method according to claim 1, further comprising:
4. The method according to any one of claims 1 to 3, wherein the event-based data comprises the position, time, size, shape, and / or intensity of the change in the visual scene.
5. The method according to any one of claims 1 to 4, wherein the predetermined threshold is either a positive threshold or a negative threshold, the positive threshold is used to check whether an increase in intensity in a particular area of the visual scene is an event, and the negative threshold is used to check whether a decrease in intensity in a particular area of the visual scene is an event.
6. The step of encoding the event-based data into spike columns is: Mapping the location of the change in the aforementioned visual scene to the corresponding neuron, Mapping the magnitude of the change in the aforementioned visual scene to the firing rate of the corresponding neuron, Mapping the shape of the change in the aforementioned visual scene to a specific subset of neurons corresponding to that shape, Mapping the intensity of the changes in the aforementioned visual scene to the timing, frequency, or amplitude of the spikes, and / or Mapping the time of change in the aforementioned visual scene to spike timing. The method according to any one of claims 1 to 5, comprising mapping the features of the event-based data to the spiking activity of neurons in the spiking neural network.
7. The method according to any one of claims 1 to 6, wherein the spiking neural network comprises multiple layers of spiking neurons connected in a feedforward manner, enabling the input spikes to flow from the input layer to the output layer, and therein recurrent and / or feedback connections are present, connecting the layers to themselves or to upstream layers, respectively.
8. The step of decoding one or more output spikes is: Counting the spiking activity of the output neuron, and / or The network calculates a maximum value or softmax function in order to select the class to which the input is predicted to belong, and / or The method according to any one of claims 1 to 7, wherein, in order to extract information relating to the timing and dynamics of the output neuron, the membrane potential of the output neuron is analyzed, preferably by calculating the first derivative of the membrane potential, which represents the rate of change of the spike count reaching the output neuron over time or under the area of the curve of the membrane potential.
9. The aforementioned method, Record the timestamp of the aforementioned event, wherein the timestamp represents the characteristics they represent or the timing of the event in relation to the input stimulus. Correlating the timestamp of the event with the feature represented by the output neuron, The method according to any one of claims 1 to 8, further comprising determining when and how the network responds to a particular input spike based on the correlation.
10. The aforementioned method, Based on the aforementioned timestamp, the time-averaged count of features in the interval between features and / or in the time interval is determined, The repeatability of the features is measured using the interval between features, and / or the intensity of the features in the time interval is measured using the time-averaged count. The method according to claim 9, further comprising:
11. The aforementioned method, The method involves taking action based on the detected occupation, and the action may be either internal or external, with the internal action being an action that affects the method by changing the control parameters of the image sensor or the spiking neural network, while the external action is an action that affects the environment of the visual scene. The method according to any one of claims 1 to 10, further comprising:
12. An occupancy detection device for occupancy detection in a visual scene captured by an image sensor, The occupation detection device comprises a processor, an encoder, a spiking neural network, and a decoder. The processor is configured to acquire event-based data, wherein the event-based data is a representation of the input from the image sensor in the form of one or more discrete events, each event representing a change in the visual scene that exceeds a predetermined threshold. The encoder is configured to encode the event-based data as input spikes, wherein the input spikes are configured to function as inputs for the spiking neural network. The spiking neural network is configured to produce one or more output spikes by processing the input spikes. The decoder is configured to decode the output spike into one or more classification results. The processor is configured to use the classification result to detect the occupation in the visual scene. Occupancy detection device.
13. The image sensor further comprises an image sensor which is either an event-based image sensor or a frame-based image sensor. The event-based image sensor captures event-based data when a change in the visual scene exceeds a predetermined threshold, preferably the event-based image sensor captures changes at the individual pixel level and / or region pixel level, and more preferably, when a change exceeding a predetermined threshold is detected in a specific region of the visual scene, the event-based image sensor transmits event-based data representing the change, along with information regarding the location and timing of the event. The frame-based image sensor captures images in a frame-by-frame manner, where each frame represents a complete snapshot of the visual scene at a specific point in time, and the processor, Preprocessing the captured image from the frame-based image sensor by calculating the difference between subsequent frames, The process involves checking whether the difference is a change in the visual scene that exceeds the predetermined threshold, If the aforementioned change exceeds the predetermined threshold, event-based data is generated based on the difference. Further configured to perform, The occupation detection device according to claim 12.
14. The event-based data comprises the position, time, size, shape, and / or intensity of the change in the visual scene, and / or The encoder described above is Mapping the location of the change in the aforementioned visual scene to the corresponding neuron, Mapping the magnitude of the change in the aforementioned visual scene to the firing rate of the corresponding neuron, Mapping the shape of the change in the aforementioned visual scene to a specific subset of neurons corresponding to that shape, Mapping the intensity of the changes in the aforementioned visual scene to the timing, frequency, or amplitude of the spikes, and / or Mapping the time of change in the aforementioned visual scene to spike timing. The occupation detection device according to claim 12 or 13, configured to map the features of the event-based data to the spiking activity of neurons in the spiking neural network.
15. The occupancy detection device according to any one of claims 12 to 14, wherein the processor is configured to take action based on the detected occupancy, the action being either internal or external, the internal action being an action that affects the method by changing the control parameters of the image sensor or the spiking neural network, while the external action being an action that affects the environment of the visual scene.