Method and apparatus for generating dynamic vision sensor (DVS) data

By acquiring image frames and using a neural network containing temporal features to generate DVS data, the problem of high cost in generating high-quality DVS data in existing technologies is solved, achieving the effect of low-cost and rapid generation of high-quality DVS data.

CN115545063BActive Publication Date: 2026-08-04LYNXI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LYNXI TECH CO LTD
Filing Date
2021-06-29
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies lack methods for rapidly generating high-quality DVS data, and existing methods are either costly or generate poor-quality data, failing to meet the training requirements of network models.

Method used

By acquiring at least two image frames, a difference image is determined, and DVS data is generated using a neural network containing temporal features, including a spiking neural network model or a recurrent neural network model, to learn the correspondence between the difference image and the DVS data.

Benefits of technology

It enables the rapid generation of large amounts of high-quality DVS data at low cost, improving data accuracy and generation efficiency, and is suitable for application scenarios with limited budgets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115545063B_ABST
    Figure CN115545063B_ABST
Patent Text Reader

Abstract

The present disclosure provides a dynamic vision sensor (DVS) data generation method and device. The method comprises: acquiring at least two image frames; determining at least one frame difference image according to the at least two image frames; and generating DVS data corresponding to each frame difference image through a first neural network; wherein the first neural network is a neural network comprising a time sequence feature. This method uses a neural network comprising a time sequence feature to quickly generate DVS data. Since the network model has high accuracy, a large amount of high-quality DVS data can be quickly generated at low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method and apparatus for generating dynamic vision sensor (DVS) data. Background Technology

[0002] Event-based cameras are a relatively new type of camera that has emerged in recent years. Unlike traditional cameras, event cameras use a novel, asynchronous data transmission format, transmitting data via events. Event cameras can also be called Dynamic Vision Sensors (DVS). The data generated by a DVS is called DVS data.

[0003] In related technologies, there is a lack of methods for quickly generating high-quality DVS data. Summary of the Invention

[0004] In view of the above problems, this disclosure is made in order to provide a method and apparatus for generating dynamic vision sensor (DVS) data to overcome or at least partially solve the above problems.

[0005] According to one aspect of the present disclosure, a method for generating dynamic vision sensor (DVS) data is provided, comprising:

[0006] Acquire at least two image frames;

[0007] Based on the at least two image frames, at least one differential image frame is determined;

[0008] A first neural network is used to generate DVS data corresponding to each frame of the differential image; wherein the first neural network is a neural network that includes temporal features.

[0009] According to another aspect of the present disclosure, a neural network training method is provided, comprising:

[0010] The first neural network is trained based on the acquired differential sample images and the DVS data corresponding to the differential sample images to obtain the trained first neural network.

[0011] The trained first neural network is used to generate DVS data, and the first neural network is a neural network that includes temporal features.

[0012] According to another aspect of the present disclosure, an apparatus for generating dynamic visual sensor (DVS) data is provided, comprising:

[0013] The acquisition module is suitable for acquiring at least two image frames;

[0014] The determining module is adapted to determine at least one differential image frame based on the at least two image frames;

[0015] The generation module is adapted to generate DVS data corresponding to each frame of the difference image through a first neural network; wherein the first neural network is a neural network that includes temporal features.

[0016] According to another aspect of the present disclosure, a neural network training apparatus is provided, comprising:

[0017] The training module is adapted to train the first neural network based on the acquired differential sample images and the DVS data corresponding to the differential sample images, so as to obtain the trained first neural network.

[0018] The trained first neural network is used to generate DVS data, and the first neural network is a neural network that includes temporal features.

[0019] Additionally, embodiments of this disclosure provide an electronic device comprising:

[0020] One or more processors;

[0021] A storage device having one or more programs stored thereon, which, when executed by one or more processors, cause the one or more processors to perform at least one of the following methods:

[0022] The above-mentioned method for generating dynamic visual sensor (DVS) data, or the above-mentioned method for training neural networks.

[0023] Additionally, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements at least one of the following methods:

[0024] The above-mentioned method for generating dynamic visual sensor (DVS) data, or the above-mentioned method for training neural networks.

[0025] In the method and apparatus for generating Dynamic Visual Sensor (DVS) data provided in this disclosure, firstly, at least two image frames are acquired, and at least one differential image frame corresponding to the at least two image frames is determined. Then, DVS data corresponding to each differential image frame is generated using a first neural network. The neural network, including temporal features, can efficiently learn the correspondence between the differential images and the DVS data, thereby rapidly generating the DVS data corresponding to the image frames based on the learned first neural network. This method achieves rapid generation of DVS data by utilizing a neural network containing temporal features. Due to the high accuracy of the network model, a large amount of high-quality DVS data can be generated quickly and at low cost. Attached Figure Description

[0026] The accompanying drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:

[0027] Figure 1 A flowchart illustrating a method for generating dynamic vision sensor (DVS) data according to an embodiment of this disclosure;

[0028] Figure 2 A flowchart illustrating a method for generating dynamic vision sensor (DVS) data according to another embodiment of this disclosure;

[0029] Figure 3 A structural diagram of a device for generating dynamic visual sensor (DVS) data, provided in yet another embodiment of this disclosure;

[0030] Figure 4 A block diagram of an electronic device provided in an embodiment of this disclosure;

[0031] Figure 5 This is a block diagram illustrating the composition of a computer-readable medium provided in an embodiment of the present disclosure. Detailed Implementation

[0032] To enable those skilled in the art to better understand the technical solutions of this disclosure, the methods, systems, electronic devices and computer-readable media provided in this disclosure will be described in detail below with reference to the accompanying drawings.

[0033] Exemplary embodiments will be described more fully below with reference to the accompanying drawings; however, these exemplary embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will enable those skilled in the art to fully understand the scope of this disclosure.

[0034] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0035] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0036] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded.

[0037] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0038] In related technologies, there are two main methods for acquiring DVS data: the first is to directly capture DVS data using a DVS camera, and the second is to generate similar DVS data through software simulation. However, the inventors discovered that both methods have at least the following drawbacks in the process of developing this invention: When using the first method, the DVS camera equipment is too expensive, thus limiting the generation of large amounts of DVS data within a given cost. When using the second method, the software-generated DVS data differs significantly in similarity from the DVS data acquired through a DVS camera, making it unsuitable as high-quality DVS data for network model training. Therefore, it is evident that related technologies cannot generate large amounts of high-quality DVS data at low cost.

[0039] To address the aforementioned issues, this disclosure provides a method for generating dynamic visual sensor (DVS) data, which can generate a large amount of high-quality DVS data in a low-cost manner.

[0040] In some embodiments, this disclosure provides a method for generating dynamic vision sensor (DVS) data. Figure 1 A flowchart illustrating a method for generating dynamic vision sensor (DVS) data according to this disclosure is shown, as follows: Figure 1 As shown, the method includes:

[0041] Step S110: Acquire at least two image frames.

[0042] There are several ways to obtain at least two image frames. For example, at least two image frames can be extracted from a dynamic video; or at least two image frames can be obtained from multiple consecutively captured still images. At least two image frames can be obtained based on the original image data, or from target transformed image data generated after transforming the original image data, or from both the original image data and the target transformed image data. In short, this disclosure does not limit the specific number or method of obtaining at least two image frames.

[0043] Step S120: Determine at least one differential image based on at least two image frames.

[0044] Specifically, every two image frames are used to determine one difference image. When there are multiple image frames, a difference image can be determined based on every two adjacent image frames. "Adjacent" here includes both direct and indirect adjacency. When determining the difference image corresponding to two image frames, the difference between each pixel in the two image frames is calculated, and the difference image is obtained based on the calculation result.

[0045] Step S130: Generate DVS data corresponding to each frame of difference image through a first neural network; wherein, the first neural network is a neural network including temporal features.

[0046] Specifically, in the process of realizing this invention, the inventors discovered that since DVS data contains temporal information, training a neural network that includes temporal features can effectively learn the temporal relationship between the difference images and the DVS data. Accordingly, the first neural network can generate DVS data corresponding to each frame of the difference image.

[0047] In specific implementation, this can be achieved through at least one of the following two methods: In the first method, at least two image frames are first acquired, and preprocessed to obtain at least one corresponding difference image. Then, the difference image obtained through the above preprocessing is input into a first neural network, thereby determining the corresponding DVS data based on the output of the first neural network. In the second method, at least two image frames or a video containing at least two image frames can also be directly input into the first neural network, which then performs the operations of acquiring at least two image frames and determining at least one difference image in steps S110 and S120.

[0048] The difference between the two implementation methods is that in the first method, the input to the first neural network is a preprocessed difference image, while in the second method, the input to the first neural network is an unprocessed dynamic video (the preprocessing step is completed internally within the neural network). Both methods can be applied to this application; the specific method used depends on the sample format selected by the neural network during training. For example, in the first method, the sample format selected by the neural network during training is a difference sample image; in the second method, the sample format selected by the neural network during training is a dynamic video.

[0049] It can be seen that this method achieves rapid generation of DVS data by using a neural network containing temporal features. Due to the high accuracy of the network model, it can generate a large amount of high-quality DVS data quickly and at low cost.

[0050] Figure 2 A flowchart illustrating a method for generating dynamic vision sensor (DVS) data according to some embodiments of this disclosure is shown. Figure 2 As shown, the method includes:

[0051] Step S200: Generate a target transformed image corresponding to the original image based on the scene type label.

[0052] Specifically, scene type labels and the original image are input into a second neural network. The output of the second neural network yields a target-transformed image corresponding to the original image. When there are multiple scene type labels, there are multiple target-transformed images. To improve the image quality of the target-transformed images, the second neural network can be determined using an adversarial network, which includes an encoder, a generator, and a discriminator. The encoder extracts coded features from the original image to help the generator produce high-quality target-transformed images. The discriminator evaluates the target-transformed images generated by the generator during training, prompting the generator to continuously optimize its generation process based on the evaluation results. Through the interaction between the discriminator and the generator, the generator is able to produce high-quality target-transformed images.

[0053] In practice, firstly, the original image is input into the encoder to obtain the encoded features output by the encoder that correspond to the original image; then, the type label and the encoded features are input into the generator so that the generator can generate a target image corresponding to the original image.

[0054] Scene type tags include weather tags, season tags, and / or sentiment tags. For example, scene transformation processing can convert an original image in a sunny scene into a target transformed image in a rainy or snowy scene. Since the target transformed image is obtained based on the encoded features of the original image, the main content of the target transformed image is the same as that of the original image. Scene transformation is achieved by adding scene type tags.

[0055] This involves pre-collecting scene images corresponding to various scenarios, learning the scene features of various scene images through a network model, and then performing scene transformation processing on the original images based on these scene features.

[0056] In practice, the second neural network is trained and the target-converted image is generated as follows: First, the generator and discriminator within the adversarial network are trained. During training, random noise based on a normal distribution and the category label of the target-converted image are input into the generator. Then, the encoder is trained to obtain the encoded features contained in the input original image. Accordingly, the second neural network mainly consists of the encoder and generator within the trained adversarial network. During conversion, the original image is first input into the encoder to obtain the encoded features corresponding to the original image output by the encoder. Then, the scene type label and the encoded features output by the encoder are input into the generator to obtain the target-converted image.

[0057] Scene transformation processing can expand original images into target transformed images for different scenes. Typically, multiple scene type labels can be assigned to the same original image, thus converting it into multiple target transformed images corresponding to different scenes. Therefore, scene transformation processing can quickly generate a large amount of image content, facilitating the rapid acquisition of large amounts of DVS data and achieving the goal of efficient data acquisition.

[0058] Of course, in other embodiments of this disclosure, this step can be omitted, and at least two image frames can be directly obtained from the original image in subsequent processes. Alternatively, at least two image frames can be obtained from the original image and the target transformed image generated above. This disclosure does not limit this.

[0059] Step S210: Obtain at least two image frames based on the target converted image.

[0060] The target image to be converted can be either a static image or a dynamic video, depending on the type of the original image. When the original image is a static image, the resulting target image is also a static image; when the original image is a dynamic video, the resulting target image is also a dynamic video. The following sections will explain these two scenarios respectively:

[0061] In the first scenario, both the original image and the target-converted image are still images. In this case, at least two still images can be extracted from the target-converted images corresponding to multiple consecutively captured original images to obtain the aforementioned at least two image frames. Alternatively, at least two image frames can be obtained based on the original image and the target-converted image. For example, it can be assumed that the original image can be used as the previous frame of the target-converted image.

[0062] In the second scenario, both the original image and the target converted image are dynamic videos. In this case, since the target converted image is a video-type image containing multiple image frames, at least two image frames are extracted from the video-type image.

[0063] In the specific extraction process, each pair of adjacent image frames is extracted into a set of images, thereby obtaining at least two image frames from each set of images. Adjacent to each other includes both directly adjacent and indirectly adjacent.

[0064] Direct adjacency refers to the situation where, in a video image consisting of multiple sequentially arranged image frames, the first and second image frames in each extracted image set are consecutively numbered and there are no other image frames between them. For example, frames 1 and 2 are identified as one image set to determine difference image 1; frames 3 and 4 are identified as another image set to determine difference image 2, and so on.

[0065] Indirect adjacency refers to the following: In a video image consisting of multiple sequentially arranged image frames, the first and second image frames in each extracted image set are not consecutively numbered, and there is at least one image frame interval between them. For example, taking the example of one image frame interval between the first and second image frames, the first and third frames are determined as one image set, so that difference image 1 can be determined based on the first and third frames; the second and fourth frames are determined as one image set, so that difference image 2 can be determined based on the second and fourth frames, and so on.

[0066] In addition, those skilled in the art can flexibly use various methods to obtain at least two image frames. For example, two image frames can be randomly selected for differential operation to obtain the subsequent differential image. Of course, the time interval between the two randomly selected image frames should be less than a preset value.

[0067] Therefore, in this embodiment, the Mth and (M+N)th image frames in a video image can be considered as a set of images, and at least two image frames can be obtained from this set of images. Here, M and N are natural numbers. Specifically, N can be a fixed value, such as 1 or 2, which are natural numbers less than a preset value. The value of M can range from 1 to positive infinity. Specifically, the value of M can range from 1 to a preset upper limit value, which depends on the total number of image frames in the dynamic video. When N is 1, it corresponds to the processing method mentioned above of extracting two directly adjacent image frames into a set of images; when N is 2, it corresponds to the processing method mentioned above of extracting two indirectly adjacent image frames into a set of images.

[0068] Step S220: Determine at least one differential image based on at least two image frames.

[0069] Specifically, since each pair of adjacent image frames was extracted into a set of images in the previous step, in this step, at least one differential image is determined based on each set of images: for each set of images, the first image frame and the second image frame contained in the set are obtained, and the differential operation is performed on the first image frame and the second image frame to obtain the differential image corresponding to the first image frame and the second image frame.

[0070] For example, in the direct adjacency method mentioned above, frames 1 and 2 can be defined as a set of images, and difference image 1 can be determined based on frames 1 and 2; frames 2 and 3 can be defined as a set of images, and difference image 2 can be determined based on frames 2 and 3; or, frames 1 and 2 can be defined as a set of images, and difference image 1 can be determined based on frames 1 and 2; frames 3 and 4 can be defined as a set of images, and difference image 2 can be determined based on frames 3 and 4, and so on. Similarly, in the indirect adjacency method mentioned above, frames 1 and 3 are defined as a set of images, and difference image 1 is determined based on frames 1 and 3; frames 2 and 4 are defined as a set of images, and difference image 2 can be determined based on frames 2 and 4, and so on. Therefore, by performing a difference operation on each pair of adjacent image frames, the corresponding difference image can be obtained.

[0071] Step S230: Generate DVS data corresponding to each frame of difference image through a first neural network; wherein, the first neural network is a neural network including temporal features.

[0072] Specifically, neural networks that include temporal features mainly refer to neural networks that contain time information in their output, i.e., neural networks carrying time information. These can be spiking neural network (SNN) models or recurrent neural network (RNN) models, etc. The first neural network includes multiple neurons. Correspondingly, when generating DVS data corresponding to each frame of the difference image through the first neural network, the first neural network processes each frame of the difference image and determines the corresponding DVS data based on the information of the neurons that emit pulses.

[0073] Furthermore, there is a primary correspondence between the information of the neurons that fire pulses and the data items contained in the DVS data. Specifically, DVS data can be represented by DVS events, and correspondingly, the DVS data for a single DVS event includes the following data items: polarity information, time information, and position information. For example, a DVS data set can be represented by the following quadruple: (x, y, t, p), where x and y represent the position coordinates of a pixel, t represents the time point at which the pixel value changes, and p represents the polarity, which indicates the change in scene brightness: positive or negative. Therefore, DVS data primarily includes information in terms of time, position, and polarity.

[0074] Accordingly, the information of the neuron firing the pulse is used to describe the behavior of the neuron firing the pulse, including at least one of the following: polarity description information related to the pulse firing type, time information related to the pulse firing time, and orientation information describing the location of the neuron firing the pulse. Therefore, the first correspondence mentioned above is specifically as follows: the polarity description information in the information of the neuron firing the pulse corresponds to the polarity information in the DVS data, the time information in the information of the neuron firing the pulse corresponds to the time information in the DVS data, and the orientation information in the information of the neuron firing the pulse corresponds to the location information in the DVS data.

[0075] Furthermore, the neurons in the first neural network correspond to the pixels in the difference image; that is, there is a second correspondence between the neurons in the first neural network and the pixels in the difference image. This second correspondence is the correspondence between the orientation information in the information of the neurons that emit pulses and the position information in the DVS data.

[0076] For example, the number of neurons in the first neural network is the same as the number of pixels in the difference image, and the neurons in the first neural network correspond one-to-one with the pixels in the difference image, that is, each neuron corresponds to a pixel in the difference image.

[0077] For example, if the number of neurons in the first neural network is less than the number of pixels in the difference image, then one neuron corresponds to multiple pixels, and the changes of multiple pixels are described by one neuron, which is equivalent to reducing the image clarity. However, reducing the number of neurons can improve the processing speed, which is especially suitable for application scenarios with high real-time requirements.

[0078] For example, if the number of neurons in the first neural network is greater than the number of pixels in the difference image, then multiple neurons correspond to one pixel. Consequently, the change of one pixel can be described by multiple neurons. In this case, calculations can be performed on the output results of multiple neurons, and the change of the corresponding pixel can be determined based on the calculation results. The calculation methods can be averaging, medianing, etc. By combining the output results of multiple neurons, the change of the corresponding pixel can be determined. This method describes the change of one pixel through multiple neurons, which can further improve the accuracy, and is especially suitable for application scenarios with high data accuracy requirements.

[0079] Therefore, on the one hand, there is a first correspondence between the information of the neuron firing a pulse and the data items contained in the DVS data; on the other hand, there is a second correspondence between the neurons in the first neural network and the pixels in the difference image. Accordingly, based on the above first and second correspondences, DVS data corresponding to the difference image is generated. Specifically, when a neuron fires a pulse, firstly, the orientation information of the neuron is obtained, thereby determining the position information in the corresponding DVS data based on the above second correspondence; then, the polarity description information and time information of the neuron are obtained, thereby determining the polarity information and time information in the corresponding DVS data based on the above first correspondence.

[0080] The following section describes the specific generation methods for each data item in the DVS data:

[0081] (1) Polarity information in DVS data

[0082] Specifically, the polarity information in the DVS data corresponds to the polarity description information contained in the information of the neurons firing pulses. Therefore, the polarity information in the corresponding DVS data is determined based on the polarity description information contained in the information of the neurons firing pulses. The first neural network can be implemented using either a single-branch neural network or a two-branch neural network. Correspondingly, the polarity description information contained in the information of the neurons firing pulses can be implemented in two ways: in the first way, the polarity description information is type indicator information; in the second way, the polarity description information is branch information. The two implementation methods are described below:

[0083] In the first implementation, the first neural network is a single-branch neural network, and the aforementioned polarity description information is type indication information. The type indication information includes: a first type indication information for indicating pulse firing according to a first firing threshold, and a second type indication information for indicating pulse firing according to a second firing threshold. Accordingly, based on the type indication information of the neuron firing the pulse, the corresponding DVS data is determined; wherein, when the type indication information is the first type indication information, the DVS data is determined to be of the first polarity; when the type indication information is the second type indication information, the DVS data is determined to be of the second polarity. Since the first neural network is a single-branch neural network, when each neuron in this neural network fires a pulse, it may be a pulse corresponding to the first polarity (e.g., positive polarity) or a pulse corresponding to the second polarity (e.g., negative polarity).

[0084] Pulses corresponding to different polarities are fired using different firing thresholds. Therefore, the polarity information of the DVS data corresponding to this pulse firing behavior needs to be determined based on the type indication information. The first firing threshold and the second firing threshold can be positive and negative, respectively. When the neuron's input is greater than or equal to the positive threshold, a pulse corresponding to the positive polarity is fired; when the neuron's input is less than or equal to the negative threshold, a pulse corresponding to the negative polarity is fired.

[0085] The absolute values ​​of the first firing threshold and the second firing threshold can be the same or different, depending on the training process of the neural network. Therefore, in the first implementation, two firing thresholds are set within the single-branch neural network, used to fire pulses corresponding to positive and negative polarities, respectively.

[0086] In the second implementation, the first neural network is a two-branch neural network, and the polarity description information mentioned above is branch information. Specifically, the first neural network includes a first branch and a second branch; when determining the corresponding DVS data based on the information of the neuron firing the pulse, the branch information contained in the information of the neuron firing the pulse is obtained, and the corresponding DVS data is determined based on the branch information; wherein, when the branch information is the first branch information, the corresponding DVS data is the first polarity; when the branch information is the second branch information, the corresponding DVS data is the second polarity.

[0087] In this system, neurons in the first branch trigger impulse firing behavior based on a first threshold, while neurons in the second branch trigger impulse firing behavior based on a second threshold. For example, neurons in the first branch trigger impulse firing behavior corresponding to positive polarity based on a positive threshold, while neurons in the second branch trigger impulse firing behavior corresponding to negative polarity based on a negative threshold.

[0088] In specific implementation, the first threshold and the second threshold can both be positive numbers, or one can be positive and the other negative, wherein the absolute values ​​of the positive and negative numbers are the same. This disclosure does not limit the specific setting method of the thresholds. Therefore, in the second implementation, the difference image is input into the first branch and the second branch respectively, and the output results of the first branch and the second branch are summarized to determine the corresponding DVS data.

[0089] (2) Time information in DVS data

[0090] Specifically, the pulse firing time of the neuron firing the pulse is obtained, and the time information contained in the corresponding DVS data is determined based on the pulse firing time. When the differential image is determined based on the first image frame at the first time point and the second image frame at the second time point, the time information of each DVS data corresponding to this differential image is between the first and second time points.

[0091] In practical implementation, if the input content of the first neural network contains multiple difference images, it is necessary to predetermine the time range corresponding to each difference image, match the pulse firing time with the time range of each difference image, and determine the correspondence between DVS data and difference images based on the matching result. The time range corresponding to each difference image is determined based on the times corresponding to the two image frames that generated the difference image. For example, assuming the first difference image is obtained from the first image frame at the first time and the second image frame at the second time, the DVS data corresponding to the pulse firing behavior triggered between the first and second times corresponds to the first difference image.

[0092] For example, suppose the input dynamic video in the first neural network contains image frame 1 at time t1, image frame 2 at time t2, and image frame 3 at time t3. Differential image 1 is generated from image frames 1 and 2, and differential image 2 is generated from image frames 2 and 3. Correspondingly, the DVS data corresponding to the pulse fired between t1 and t2 corresponds to differential image 1, and the DVS data corresponding to the pulse fired between t2 and t3 corresponds to differential image 2. If the same neuron fires once between t1 and t2, and again between t2 and t3, then the DVS data for the two firings correspond to the two differential images, respectively.

[0093] (3) Location information in DVS data

[0094] Specifically, the directional information contained in the information of the neuron firing the pulse is obtained, and the positional information contained in the corresponding DVS data is determined based on the directional information. The directional information contained in the information of the neuron firing the pulse can be the neuron's location or sequence number; in short, any information that can uniquely locate a neuron can be used as the directional information. Since there is a correspondence between neurons and pixels, the positional information contained in the information of the neuron firing the pulse can be used to determine the positional information contained in the corresponding DVS data.

[0095] In summary, this method achieves rapid generation of DVS data by leveraging a neural network containing temporal features. Due to the high accuracy of the network model, it can quickly generate a large amount of high-quality DVS data at low cost. Furthermore, since there is a first correspondence between neurons in the first neural network and pixels in the difference image, and a second correspondence between the information of the neurons firing pulses and the data items contained in the DVS data, DVS data corresponding to the difference image can be generated based on these correspondences.

[0096] In this application, by mapping each information item in the neuron information of the neural network to each data item in the DVS data during the generation of DVS data, the precise generation process of DVS data can be achieved through neurons, thus improving the accuracy and quality of the DVS data. Furthermore, by using scene transformation, a large number of difference images can be quickly obtained, facilitating the rapid generation of large amounts of high-quality DVS data at low cost. Additionally, because the first neural network in this embodiment has temporal characteristics, it can generate DVS data quickly in a low-power, low-cost manner.

[0097] Some embodiments of this disclosure also provide a method for training a neural network. The method includes: training a first neural network based on acquired differential sample images and corresponding DVS data to obtain a trained first neural network; wherein the trained first neural network is used to generate DVS data, and the first neural network is a neural network including temporal features.

[0098] Optionally, before training the first neural network based on the acquired differential sample images and the DVS data corresponding to the differential sample images, the method further includes: acquiring a first sample image frame and a second sample image frame corresponding to the differential sample images; and calculating the DVS data corresponding to the differential sample images based on the first sample image frame and the second sample image frame.

[0099] Obtaining DVS data corresponding to the difference sample images through computation can solve the high cost problem of DVS camera acquisition of DVS data, providing a low-cost way to obtain DVS data as samples. Alternatively, to improve sample accuracy, DVS data and image frames can be acquired simultaneously using a DVS camera, and then time and space aligned to obtain the difference image and its corresponding DVS data. Both of these sample acquisition methods can be used individually or simultaneously.

[0100] It should be noted that, in this embodiment, the first neural network can be trained in the following two ways:

[0101] In the first training method, a difference operation is first performed on the first sample image frame and the second sample image frame to obtain the aforementioned difference sample image. Then, the first neural network is trained based on the difference sample image and the corresponding DVS. In this method, the first neural network is input to the preprocessed difference sample image and the annotation result of the difference sample image, i.e., DVS data, during the training process.

[0102] In the second training method, a sample video containing both the first and second sample image frames is directly input into the first neural network, allowing the network to learn the correlation between the sample video and its corresponding DVS data. In this method, the first neural network is trained using the unprocessed raw sample video and its annotation results (i.e., DVS data).

[0103] In addition, to expand the sample size, original samples and their corresponding DVS data can be generated first. Then, using the scene conversion method in step S200 of the previous embodiment, the original samples can be converted into target conversion samples with the help of scene type labels. The specific conversion method is the same as in the previous embodiment. The scene conversion method can expand the sample size and improve training efficiency. High-precision original samples and their corresponding DVS data can be generated using a DVS camera to improve the quality of the original samples.

[0104] In practice, the number of neurons in the neural network model is determined based on the number of pixels in the differential sample image to initialize the first neural network. Correspondingly, during the training process, the parameter values ​​of each parameter contained in the first neural network are determined based on the correlation between the behavior of the neurons firing pulses in the first neural network and the DVS data, thereby obtaining the trained first neural network.

[0105] Specifically, the training process of the network model is similar to the application process of the network model in Example 2: the neurons in the first neural network correspond to the pixels in the difference image, that is, there is a second correspondence between the neurons in the first neural network and the pixels in the difference image. Accordingly, when initializing the first neural network, processing is performed based on this second correspondence. For example, in one initialization method, the number of neurons in the first neural network is the same as the number of pixels in the difference image, and there is a one-to-one correspondence between the neurons in the first neural network and the pixels in the difference image. In another initialization method, the number of neurons in the first neural network is less than the number of pixels in the difference image. Yet another initialization method, the number of neurons in the first neural network is greater than the number of pixels in the difference image.

[0106] Furthermore, there is a primary correspondence between the information of the neurons that fire pulses and the data items contained in the DVS data. Specifically, DVS data can be represented by DVS events, and correspondingly, the DVS data corresponding to a DVS event includes the following data items: polarity information, time information, and position information. For example, a DVS data can be represented by the following quadruple form: (x, y, t, p), where x and y represent the position coordinates of the pixel, t represents the time point when the pixel value changes, and p represents the polarity, where polarity represents the change in scene brightness: positive or negative.

[0107] Therefore, DVS data mainly includes information on three aspects: time, location, and polarity. Correspondingly, the information of the neuron firing the pulse is used to describe the behavior of the neuron firing the pulse, including at least one of the following: polarity description information related to the pulse firing type, time information related to the pulse firing time, and orientation information describing the location of the neuron firing the pulse. Thus, the first correspondence mentioned above specifically means: the polarity description information in the information of the neuron firing the pulse corresponds to the polarity information in the DVS data; the time information in the information of the neuron firing the pulse corresponds to the time information in the DVS data; and the orientation information in the information of the neuron firing the pulse corresponds to the location information in the DVS data. Accordingly, the main purpose of the training process is to learn the above first and second correspondences, and determine the various parameters within the first neural network based on the learning results.

[0108] Optionally, the DVS data includes polarity information; then each neuron in the first neural network fires a pulse containing first type indication information based on a first firing threshold, and fires a pulse containing second type indication information based on a second firing threshold. The first firing threshold and the second firing threshold correspond to the single-branch neural network mentioned above. The specific setting method of the thresholds can be referred to the description of the previous embodiment.

[0109] Optionally, if the DVS data includes polarity information, then the first neural network includes a first branch and a second branch; wherein neurons in the first branch fire pulses corresponding to the first polarity based on a first threshold, and neurons in the second branch fire pulses corresponding to the second polarity based on a second threshold. The first and second thresholds correspond to the bi-branch neural network mentioned above. The specific setting method of the thresholds can be referred to the description in the previous embodiment.

[0110] The aforementioned thresholds can be adjusted and optimized during training using the loss function and backpropagation function.

[0111] Optionally, if the DVS data includes time information, then the pulse firing time contained in the information of the neurons firing pulses in the first neural network is used to determine the time information contained in the corresponding DVS data.

[0112] Optionally, the DVS data includes positional information, and the positional information of the neurons firing pulses in the first neural network is used to determine the positional information contained in the corresponding DVS data.

[0113] In summary, the training method in this embodiment can generate a neural network for producing DVS data. This method achieves rapid generation of DVS data by utilizing a neural network containing temporal features. Due to the high accuracy of the network model, it can quickly generate a large amount of high-quality DVS data at low cost. Furthermore, since there is a first correspondence between the neurons in the first neural network and the pixels in the difference image, and a second correspondence between the information of the neurons firing pulses and the data items contained in the DVS data, DVS data corresponding to the difference image can be generated based on the above correspondences.

[0114] In addition, another embodiment of this disclosure provides an apparatus for generating dynamic visual sensor (DVS) data. Figure 3 A schematic diagram of the device is shown, as follows. Figure 3 As shown, the device includes:

[0115] Acquisition module 31 is adapted to acquire at least two image frames;

[0116] The determining module 32 is adapted to determine at least one differential image frame based on the at least two image frames;

[0117] The generation module 33 is adapted to generate DVS data corresponding to each frame of difference image through a first neural network; wherein the first neural network is a neural network including temporal features.

[0118] Optionally, the first neural network includes multiple neurons, wherein the generation module is specifically adapted to:

[0119] The first neural network processes each frame of the differential image and determines the corresponding DVS data based on the information of the neurons that emit pulses.

[0120] Optionally, the DVS data includes polarity information, and the information of the neuron firing the pulse includes type indication information; the type indication information includes: a first type indication information for indicating that the pulse is fired according to a first firing threshold, and a second type indication information for indicating that the pulse is fired according to a second firing threshold; wherein, determining the corresponding DVS data based on the information of the neuron firing the pulse includes: determining the corresponding DVS data based on the type indication information of the neuron firing the pulse; wherein, when the type indication information is the first type indication information, the DVS data is determined to be of the first polarity; when the type indication information is the second type indication information, the DVS data is determined to be of the second polarity.

[0121] Optionally, the DVS data includes polarity information, and the first neural network includes a first branch and a second branch; wherein, the generation module is specifically adapted to:

[0122] The branch information contained in the information of the neuron that emits the pulse is obtained, and the corresponding DVS data is determined based on the branch information; wherein, when the branch information is the first branch information, the corresponding DVS data is the first polarity; when the branch information is the second branch information, the corresponding DVS data is the second polarity.

[0123] Optionally, the DVS data includes time information, wherein the generation module is specifically adapted to: obtain the pulse firing time contained in the information of the neuron firing the pulse, and determine the time information contained in the corresponding DVS data based on the pulse firing time.

[0124] Optionally, the differential image is determined based on a first image frame at a first time and a second image frame at a second time, wherein the time information of each DVS data corresponding to the differential image is between the first time and the second time.

[0125] Specifically, the generation module is adapted to: obtain the orientation information contained in the information of the neurons that emit pulses, and determine the position information contained in the corresponding DVS data based on the orientation information.

[0126] Optionally, the acquisition module is specifically adapted to: generate a target transformed image corresponding to the original image based on the scene type label; and acquire the at least two image frames based on the target transformed image.

[0127] Optionally, the acquisition module is specifically adapted to: input the scene type label and the original image into a second neural network, and obtain a target transformed image corresponding to the original image based on the output of the second neural network; wherein, when there are multiple scene type labels, the number of target transformed images is multiple.

[0128] Optionally, the second neural network is determined based on an adversarial network, and the adversarial network includes an encoder, a generator, and a discriminator. In this case, the acquisition module is specifically adapted to: input the original image into the encoder to acquire the encoded features output by the encoder corresponding to the original image; input the type label and the encoded features into the generator so that the generator can generate a target image corresponding to the original image.

[0129] Optionally, the acquisition module is specifically adapted to: when the target converted image is a video image containing multiple image frames, take the Mth image frame and the M+Nth image frame in the video image as a set of images, and obtain at least two image frames based on the set of images; where M and N are natural numbers.

[0130] The specific working principles of each of the above modules can be found in the description of the corresponding steps in the method embodiment, and will not be repeated here.

[0131] In addition, another embodiment of this disclosure provides a neural network training apparatus, including:

[0132] The training module is adapted to train a first neural network based on the acquired differential sample images and the DVS data corresponding to the differential sample images, to obtain a trained first neural network; wherein the trained first neural network is used to generate DVS data, and the first neural network is a neural network including temporal features.

[0133] Optionally, the device further includes:

[0134] The acquisition module is adapted to acquire a first sample image frame and a second sample image frame corresponding to the differential sample image; and to calculate DVS data corresponding to the differential sample image based on the first sample image frame and the second sample image frame.

[0135] Optionally, the DVS data includes polarity information; each neuron in the first neural network fires a pulse containing first type indication information based on a first firing threshold, and fires a pulse containing second type indication information based on a second firing threshold.

[0136] Optionally, the DVS data includes polarity information; the first neural network includes a first branch and a second branch; wherein, neurons in the first branch fire pulses corresponding to a first polarity based on a first threshold, and neurons in the second branch fire pulses corresponding to a second polarity based on a second threshold.

[0137] Optionally, the DVS data includes time information, wherein the pulse firing time contained in the information of the neurons firing pulses in the first neural network is used to determine the time information contained in the corresponding DVS data.

[0138] Optionally, the DVS data includes position information, and the position information of the neurons that fire pulses in the first neural network is used to determine the position information contained in the corresponding DVS data.

[0139] The specific working principles of each of the above modules can be found in the description of the corresponding steps in the method embodiment, and will not be repeated here.

[0140] Additionally, refer to Figure 4 This disclosure provides an electronic device, which includes:

[0141] One or more processors 101;

[0142] The memory 102 stores one or more programs that, when executed by one or more processors, enable the processors to implement the aforementioned method for generating dynamic visual sensor (DVS) data, or the aforementioned method for training neural networks.

[0143] One or more I / O interfaces 103 are connected between the processor and the memory and configured to enable information exchange between the processor and the memory.

[0144] The processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read / write interface) 103 is connected between the processor 101 and the memory 102, enabling information exchange between the processor 101 and the memory 102, including but not limited to a data bus (Bus).

[0145] In some embodiments, the processor 101, memory 102, and I / O interface 103 are interconnected via bus 104, and thus connected to other components of the computing device.

[0146] Additionally, refer to Figure 5This disclosure provides a computer-readable medium storing a computer program that, when executed by a processor, implements the method for generating dynamic visual sensor (DVS) data as described in any of the above embodiments, or the neural network training method described above.

[0147] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0148] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A method for generating dynamic vision sensor (DVS) data, comprising: Acquire at least two image frames; Based on the at least two image frames, at least one differential image frame is determined; A first neural network is used to generate DVS data corresponding to each frame of the difference image; wherein the first neural network is a neural network that includes temporal features; The acquisition of at least two image frames includes: Generate a target transformed image corresponding to the original image based on the scene type label; The at least two image frames are obtained based on the target converted image.

2. The method according to claim 1, wherein the first neural network comprises a plurality of neurons, wherein, The process of generating DVS data corresponding to each frame of the difference image through the first neural network includes: The first neural network processes each frame of the differential image and determines the corresponding DVS data based on the information of the neurons that emit pulses.

3. The method according to claim 2, wherein the DVS data includes polarity information, and the information of the neuron firing the pulse includes: Type indication information; The type indication information includes: first type indication information for indicating that the pulse is issued according to a first issuance threshold, and second type indication information for indicating that the pulse is issued according to a second issuance threshold; The step of determining the corresponding DVS data based on the information of the neuron that fired the pulse includes: Based on the type indication information of the neuron that emits the pulse, the corresponding DVS data is determined; wherein, when the type indication information is a first type indication information, the DVS data is determined to be of the first polarity; when the type indication information is a second type indication information, the DVS data is determined to be of the second polarity.

4. The method according to claim 2, wherein the DVS data includes polarity information, and the first neural network includes a first branch and a second branch; wherein, The step of determining the corresponding DVS data based on the information of the neuron that fired the pulse includes: The branch information contained in the information of the neuron that emits the pulse is obtained, and the corresponding DVS data is determined based on the branch information; wherein, when the branch information is the first branch information, the corresponding DVS data is the first polarity; when the branch information is the second branch information, the corresponding DVS data is the second polarity.

5. The method according to claim 2, wherein the DVS data includes time information. in, The step of determining the corresponding DVS data based on the information of the neuron that fired the pulse includes: obtaining the pulse firing time contained in the information of the neuron that fired the pulse, and determining the time information contained in the corresponding DVS data based on the pulse firing time.

6. The method according to claim 5, wherein, The differential image is determined based on the first image frame at the first time and the second image frame at the second time, and the time information of each DVS data corresponding to the differential image is between the first time and the second time.

7. The method according to claim 2, wherein the neurons of the first neural network correspond to the pixels in the difference image, and the DVS data includes location information. in, The step of determining the corresponding DVS data based on the information of the neuron that fired the pulse includes: The directional information contained in the information of the neuron that emits the pulse is obtained, and the position information contained in the corresponding DVS data is determined based on the directional information.

8. The method according to claim 1, wherein, The step of generating a target transformed image corresponding to the original image based on the scene type label includes: The scene type label and the original image are input into the second neural network, and the target transformed image corresponding to the original image is obtained based on the output of the second neural network. When there are multiple scene type labels, the number of target conversion images is also multiple.

9. The method according to claim 8, wherein, The second neural network is determined based on an adversarial network, and the adversarial network includes an encoder, a generator, and a discriminator. Therefore, inputting the scene type label and the original image into the second neural network includes: The original image is input into the encoder to obtain the encoded features output by the encoder that correspond to the original image; The scene type label and the encoded features are input into the generator so that the generator can generate a target image corresponding to the original image.

10. The method according to claim 1, wherein, The step of obtaining the at least two image frames based on the target converted image includes: When the target converted image is a video image containing multiple image frames, the Mth image frame and the (M+N)th image frame in the video image are taken as a set of images, and at least two image frames are obtained from this set of images; where M and N are natural numbers.

11. A neural network training method, comprising: The first neural network is trained based on the acquired differential sample images and the DVS data corresponding to the differential sample images to obtain the trained first neural network. The trained first neural network is used to generate DVS data, and the first neural network is a neural network that includes temporal features; Before acquiring the differential sample image, the method further includes: Obtain a first sample image frame and a second sample image frame corresponding to the differential sample image; obtain the differential sample image based on the first sample image frame and the second sample image frame; The step of obtaining the first sample image frame and the second sample image frame includes: generating a target transformed image corresponding to the original image based on the scene type label; and obtaining the first sample image frame and the second sample image frame based on the target transformed image.

12. The method according to claim 11, further comprising, before training the first neural network based on the acquired difference sample image and the DVS data corresponding to the difference sample image: Obtain the first sample image frame and the second sample image frame corresponding to the differential sample image; Calculate the DVS data corresponding to the differential sample image based on the first sample image frame and the second sample image frame.

13. The method of claim 11, wherein the DVS data includes polarity information; Each neuron in the first neural network fires a pulse containing a first type of indication information based on a first firing threshold, and fires a pulse containing a second type of indication information based on a second firing threshold.

14. The method of claim 11, wherein the DVS data includes polarity information; The first neural network includes: A first branch and a second branch; wherein, the neurons in the first branch fire pulses corresponding to a first polarity based on a first threshold, and the neurons in the second branch fire pulses corresponding to a second polarity based on a second threshold.

15. The method according to claim 11, wherein the DVS data includes time information. in, The pulse firing time included in the information of the neurons firing pulses in the first neural network is used to determine the time information contained in the corresponding DVS data.

16. The method according to claim 11, wherein, The DVS data includes location information, and the location information of the neurons that fire pulses in the first neural network is used to determine the location information contained in the corresponding DVS data.

17. A device for generating dynamic vision sensor (DVS) data, comprising: The acquisition module is suitable for acquiring at least two image frames; The determining module is adapted to determine at least one differential image frame based on the at least two image frames; The generation module is adapted to generate DVS data corresponding to each frame of the difference image through a first neural network; wherein the first neural network is a neural network that includes temporal features; The acquisition module is specifically used to: generate a target transformed image corresponding to the original image based on the scene type label; and acquire the at least two image frames based on the target transformed image.

18. A neural network training device, comprising: The training module is adapted to train the first neural network based on the acquired differential sample images and the DVS data corresponding to the differential sample images, so as to obtain the trained first neural network. The trained first neural network is used to generate DVS data, and the first neural network is a neural network that includes temporal features; The acquisition module is configured to acquire a first sample image frame and a second sample image frame corresponding to the differential sample image before acquiring the differential sample image; and acquire the differential sample image based on the first sample image frame and the second sample image frame. The acquisition module is specifically used to: generate a target transformed image corresponding to the original image based on the scene type label; and acquire the first sample image frame and the second sample image frame based on the target transformed image.

19. An electronic device comprising: One or more processors; A storage device having stored one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method for generating dynamic visual sensor (DVS) data according to any one of claims 1 to 10, or to implement the neural network training method according to any one of claims 11 to 16.

20. A computer-readable medium having a computer program stored thereon, the program, when executed by a processor, implementing at least one of the following methods: The method for generating dynamic visual sensor (DVS) data according to any one of claims 1 to 10, or the neural network training method according to any one of claims 11 to 16.