Event data processing method, apparatus and system

By training a neural network model to adapt to event data in different spatiotemporal domains, feature matching and alignment were achieved, solving the problem of low accuracy of existing models when processing data in different spatiotemporal domains, and improving processing efficiency and accuracy.

CN115546248BActive Publication Date: 2026-04-07HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing pre-trained neural network models have low accuracy when processing event data in different spatiotemporal domains and cannot effectively adapt to the needs of various scenarios.

Method used

The neural network model is trained using at least two types of sample event data, including data from different and the same spatiotemporal domains. By adjusting the network parameters, feature matching and alignment between different spatiotemporal domains can be achieved, thereby improving processing accuracy.

Benefits of technology

It improves the accuracy of neural network models in processing event data in different spatiotemporal domains, reduces resource waste, and meets the requirements of high efficiency and low power consumption in computer vision applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546248B_ABST
    Figure CN115546248B_ABST
Patent Text Reader

Abstract

This invention relates to an event data processing method, apparatus, and system, belonging to the field of machine vision technology. Addressing the problem of spatiotemporal mismatch between event data from different spatiotemporal domains, this application, after acquiring first event data collected by a dynamic visual sensing device, processes the first event data using a neural network model to obtain a first recognition result of the target object. This neural network model is trained using at least two types of sample event data. The neural network model trained using at least two types of sample event data exhibits high processing accuracy when processing event data within the same spatiotemporal domain as the at least two types of sample event data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and in particular to an event data processing method, apparatus and system. Background Technology

[0002] Dynamic Vision Sensors (DVS), also known as event cameras, capture dynamic changes in a scene based on an event-driven approach. Typically, pre-trained neural network models are used to process the event data output by the event camera to achieve goals such as target tracking or gesture recognition.

[0003] In practical applications, event cameras output event data in different spatiotemporal domains according to the needs of different scenarios. However, pre-trained neural network models are only suitable for processing event data in one spatiotemporal domain, and their processing accuracy for event data in other spatiotemporal domains is low. Summary of the Invention

[0004] This application provides an event data processing method, apparatus, and system to improve the accuracy of event data processing.

[0005] In a first aspect, embodiments of this application provide an event data processing method, comprising: acquiring first event data collected by a dynamic visual sensing device, wherein the first event data is used to indicate dynamic events of a target object; processing the first event data through a neural network model to obtain a first recognition result of the target object; the neural network model is trained based on at least two types of sample event data, wherein the at least two types of sample event data have different spatiotemporal domains, or in other words, the sample event data used to train the neural network model is divided into multiple types according to the spatiotemporal domains of the sample event data, and the spatiotemporal domains of any two of the obtained at least two types of sample event data are different. For example, the neural network model can be trained based on three types of sample event data, wherein the spatiotemporal domains of the three types of sample event data are all different. The at least two types of sample event data include sample event data with the same spatiotemporal domain as the first event data.

[0006] Specifically, at least two types of sample event data have different frame rates, or at least two types of sample event data have different event point densities or numbers. Alternatively, any two types of sample event data may have different frame rates, or at least two types of sample event data may have different event point densities or numbers.

[0007] The neural network model in this application embodiment is trained using at least two types of sample event data with different spatiotemporal domains, including sample event data with the same spatiotemporal domain as the first event data to be processed. That is, in addition to using event data with a different spatiotemporal domain than the first event data as sample event data, event data with the same spatiotemporal domain as the first event data is also collected as sample event data. The neural network model trained using the above-mentioned at least two types of sample event data exhibits high processing accuracy when processing event data in the same spatiotemporal domain as the at least two types of sample event data. Furthermore, it eliminates the need to train a separate neural network model for each spatiotemporal domain, reducing resource waste.

[0008] In one possible design, the distribution area of ​​feature points in the spatiotemporal domain features extracted by the neural network model for the first event data is consistent with the distribution area of ​​feature points in the spatiotemporal domain features extracted for the second event data. The spatiotemporal domain of the second event data is the same as the spatiotemporal domain of one of the sample event data of at least two sample event data, while the spatiotemporal domains of the first event data and the second event data are different.

[0009] For example, in one application scenario, a neural network model can be used to process first event data. Specifically, the neural network model extracts the spatiotemporal features of the first event data and obtains a first recognition result of the target object indicated by the first event data based on these features. In another application scenario, the neural network model can also be used to process second event data. Specifically, second event data collected by a dynamic visual sensing device is acquired, the spatiotemporal features of the second event data are extracted using a neural network model, and a second recognition result of the target object indicated by the second event data is obtained based on these features. Here, the first event data and the second event data are event data from different spatiotemporal domains. The sample event data used to train the neural network model includes event data with spatiotemporal domains different from both the first and second event data, as well as event data with the same spatiotemporal domain as the first event data and event data with the same spatiotemporal domain as the second event data. The distribution area of ​​feature points in the spatiotemporal domain features extracted by the neural network model trained in this way is basically consistent with the distribution area of ​​feature points in the spatiotemporal domain features extracted by the neural network model trained in the second event data. This can achieve matching and alignment of spatiotemporal domain features of event data from different spatiotemporal domains in the spatiotemporal dual dimensions, thereby effectively overcoming the problem of spatiotemporal domain mismatch between event data from different spatiotemporal domains and improving the processing accuracy of event data by the neural network model in practical applications.

[0010] In one possible design, the neural network model can include a spiking neural network (SNN) model, which is capable of extracting features from event data in both spatial and temporal dimensions. SNN models have low latency in processing input data; therefore, using a SNN-based neural network model to process event data acquired by an event camera can fully utilize the high temporal resolution of the event data for real-time output, meeting the requirements of high efficiency and low power consumption in computer vision applications.

[0011] In one possible design, the network parameters of the neural network model are adjusted based on the spatiotemporal discrimination results of at least two types of sample event data and the predicted recognition results of the objects indicated by at least two types of sample event data. For example, the at least two types of sample event data include sample event data from a first spatiotemporal domain. When training the neural network model, a total loss value can be determined based on the spatiotemporal discrimination results of each type of sample event data and the predicted recognition results of the objects indicated by the sample event data from the first spatiotemporal domain. The network parameters of the neural network model are then adjusted based on the determined total loss value, so that the distribution of spatiotemporal features extracted by the trained neural network model for sample event data from different spatiotemporal domains is basically consistent.

[0012] In another possible design, the network parameters of the neural network model are obtained by adjusting a first loss value and a second loss value. The first loss value is obtained by reversing the positive and negative values ​​of loss values ​​obtained from the spatiotemporal discrimination results of at least two types of sample event data. The second loss value is obtained from the predicted recognition results of the objects indicated by the at least two types of sample event data. For example, when training the neural network model, an auxiliary training spatiotemporal discrimination network can be used to perform spatiotemporal discrimination on the spatiotemporal features of the at least two types of sample event data extracted by the neural network model, obtaining the corresponding spatiotemporal discrimination results of the sample event data. The loss values ​​obtained from the spatiotemporal discrimination results of the at least two types of sample event data are then reversed to obtain the first loss value. The second loss value is obtained based on the predicted recognition results of the objects indicated by the sample event data in the first spatiotemporal domain of the at least two types of sample event data. The network parameters of the neural network model are then adjusted by combining the first and second loss values. Since the first loss value used for adjusting network parameters is obtained by reversing the positive and negative values ​​of the loss values ​​corresponding to the spatiotemporal domain discrimination results, the direction of adjusting network parameters is to make the spatiotemporal domain features extracted by the neural network model for sample event data in different spatiotemporal domains become closer and closer. This allows the trained neural network model to overcome the problem of spatiotemporal domain mismatch between event data in different spatiotemporal domains, and further improve the accuracy of the neural network model in processing event data in practical applications.

[0013] In another possible design, the neural network model includes a feature extraction network and an object recognition network. The feature extraction network extracts spatiotemporal features from at least two types of sample event data, while the object recognition network determines the predicted recognition result of the object indicated by the at least two types of sample event data. After training, the distribution of the spatiotemporal features extracted by the feature extraction network from at least two different spatiotemporal domain sample event data is essentially consistent, thus enabling the matching and alignment of spatiotemporal features of event data from different spatiotemporal domains in both spatiotemporal dimensions.

[0014] Secondly, embodiments of this application also provide an event data processing apparatus, which includes corresponding functional modules, each used to implement the steps in the above methods. For details, please refer to the detailed description in the method examples; further elaboration is not provided here. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example, the event data processing apparatus includes a data acquisition unit and a data processing unit. The data acquisition unit is used to acquire first event data collected by a dynamic visual sensing device, the first event data being used to indicate dynamic events of a target object; the data processing unit is used to process the first event data through a neural network model to obtain a first recognition result of the target object; wherein the neural network model is trained using at least two types of sample event data, the at least two types of sample event data having different spatiotemporal domains, and the at least two types of sample event data including sample event data with the same spatiotemporal domain as the first event data.

[0015] Thirdly, embodiments of this application provide an event data processing system, including a dynamic visual sensing device and a processor; the dynamic visual sensing device is used to collect first event data, the first event data being used to indicate dynamic events of a target object; the processor is connected to the dynamic visual sensing device and is used to execute the method described in the first aspect or any design of the first aspect. Specifically, the processor acquires the first event data collected by the dynamic visual sensing device and executes the method described in the first aspect or any design of the first aspect on the first event data.

[0016] Fourthly, this application provides a computer-readable storage medium storing a computer program or instructions that, when executed by a terminal device, cause the terminal device to perform the method described in the first aspect or any possible design of the first aspect.

[0017] Fifthly, this application provides a computer program product comprising a computer program or instructions that, when executed by a terminal device, implement the method described in the first aspect or any possible implementation thereof.

[0018] The technical effects that can be achieved by any of the second to fifth aspects mentioned above can be referred to the description of the beneficial effects in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the hardware structure of the electronic device in the embodiments of this application;

[0020] Figure 2 A schematic diagram of the spatiotemporal data stream acquired by the event camera;

[0021] Figure 3 A schematic diagram of event image frames at different frame rates output by the event camera;

[0022] Figure 4 A flowchart illustrating an example of an event data processing method provided in an embodiment of this application;

[0023] Figure 5 A schematic diagram illustrating an example of the network architecture used in the model training process provided in this application embodiment;

[0024] Figure 6 A schematic diagram illustrating an example of a feature extraction network provided in an embodiment of this application;

[0025] Figure 7 A schematic diagram illustrating an example of a spatiotemporal gradient inversion module provided in an embodiment of this application;

[0026] Figure 8 A schematic diagram illustrating an example of a spatiotemporal domain discrimination network provided in an embodiment of this application;

[0027] Figure 9 A schematic diagram illustrating another example of the spatiotemporal domain discrimination network provided in the embodiments of this application;

[0028] Figure 10 A schematic diagram illustrating an example of a predictive classification network provided in an embodiment of this application;

[0029] Figure 11 A flowchart illustrating an example of the training process of a neural network model provided in an embodiment of this application;

[0030] Figure 12 A comparison chart showing the source domain features and target domain features extracted before and after model training;

[0031] Figure 13 A schematic diagram of an example of an event data processing apparatus provided in an embodiment of this application;

[0032] Figure 14This is a schematic diagram of an example of an event data processing system provided in an embodiment of this application. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be described in detail below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only used to explain specific embodiments of this application and is not intended to limit this application. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0034] Before introducing the specific solutions provided in the embodiments of this application, some terms used in this application will be explained in a general way to facilitate understanding by those skilled in the art, and the terms used in this application will not be limited.

[0035] (1) Dynamic visual sensing device: also known as event camera, event-driven camera or event camera sensor, is a new type of camera that has emerged in recent years. Unlike traditional cameras that capture a complete image, event cameras capture "events". Event cameras take the brightness change of a single pixel in a real scene as an event. Event cameras capture dynamic changes in a scene based on an event-driven approach. When objects in a real scene change, the event camera generates a spatiotemporal data stream of a series of events.

[0036] Compared to traditional cameras, event cameras have the advantages of high temporal resolution, large dynamic range, and low time delay, and are of great use in fields such as high dynamic range image reconstruction, target tracking, and gesture recognition.

[0037] (2) Spatiotemporal domain: a general term for the time domain and the spatial domain. In the embodiments of this application, the frame rate of event data in different spatiotemporal domains is different, or the density or number of event points in event data in different spatiotemporal domains is different.

[0038] Specifically, the event camera compresses the acquired spatiotemporal data stream into frames according to a set period, outputting event data at a fixed frame rate, or event image frames at a fixed frame rate. This period can be called the temporal resolution. For example, a temporal resolution of 5ms means that the spatiotemporal data stream is compressed into one event image frame every 5ms. The relationship between frame rate and temporal resolution is that the product of frame rate and temporal resolution is 1 second. For example, event data with a frame rate of 200 frames / second has a temporal resolution of 5ms.

[0039] It is understandable that the two types of event data have different frame rates or different temporal resolutions, and the density or number of event points in these two types of event data are also different. It can be assumed that the spatiotemporal domains of these two types of event data are different.

[0040] (3) Neural network model: In this embodiment of the application, the neural network model can be used to perform subsequent processing on the event data output by the event camera.

[0041] The neural network model in this embodiment can be constructed based on spiking neural networks (SNNs). SNNs are a new generation of neural networks inspired by the brain's operating mechanisms, using pulse sequences as the data transmission format. Compared to traditional artificial neural networks (ANNs), spiking neural networks use a more biologically interpretable spiking neuron model as their basic unit, possessing advantages such as low latency and low energy consumption. They can simulate various neural signals and arbitrary continuous functions, and can process complex spatiotemporal information.

[0042] In this application embodiment, "multiple" refers to two or more. Therefore, in this application embodiment, "multiple" can also be understood as "at least two". "At least one" can be understood as one or more, such as one, two, or more. For example, "including at least one" means including one, two, or more, and it does not limit which ones are included. For example, including at least one of A, B, and C, then it could include A, B, C, A and B, A and C, B and C, or A and B and C. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0043] Unless otherwise stated, the ordinal numbers such as "first" and "second" mentioned in the embodiments of this application are used to distinguish multiple objects, and are not used to limit the order, sequence, priority or importance of multiple objects.

[0044] This application embodiment can be used in electronic devices with built-in or external dynamic visual sensing devices. The electronic device can be a device that provides users with video recording and / or data connectivity, a handheld device with wireless connectivity, or other processing devices connected to a wireless modem, such as: mobile phones (or "cellular" phones), smartphones, and can be portable, pocket-sized, handheld, wearable devices (such as smartwatches), tablets, personal computers (PCs), PDAs (Personal Digital Assistants), in-vehicle computers, drones, aerial photography devices, computers, etc.

[0045] For example, the electronic device may be a dynamic vision sensing device or a device equipped with a dynamic vision sensing device, such as a mobile phone, tablet computer, or vehicle terminal equipped with a dynamic vision sensing device. The electronic device may also be a device for processing event data, such as a computer or server. The computer can connect to the dynamic vision sensing device via wired or wireless means to receive and process the event data transmitted by the dynamic vision sensing device; the server can receive and process the event data remotely sent by the dynamic vision sensing device via a network.

[0046] In the following detailed description, the dynamic visual sensing device is illustrated using an event camera as an example. For instance, Figure 1 A schematic diagram of an optional hardware structure of an electronic device 100 to which embodiments of this application are applicable is shown.

[0047] Electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, an event camera 190, buttons 191, a camera 192, and a display screen 193. In some embodiments, the event camera 190 may be a sensor within the sensor module 180; in other embodiments, the event camera 190 may be independent of the sensor module 180.

[0048] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0049] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0050] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0051] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0052] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0053] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 192, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.

[0054] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0055] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0056] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.

[0057] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 193 and the camera 192. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 192 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 193 communicate via the DSI interface to enable the electronic device 100 to display images.

[0058] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 192, a display screen 193, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0059] The SIM interface is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM interface to make contact with and detach from the electronic device 100. The electronic device 100 can support one or N3 SIM interfaces, where N3 is a positive integer greater than 1. The SIM interface can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM interface simultaneously. The multiple cards can be of the same or different types. The SIM card interface is also compatible with different types of SIM cards. The SIM card interface is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0060] A USB interface is an interface that conforms to the USB standard specification, specifically including Mini USB, Micro USB, and USB Type-C interfaces. A USB interface can be used to connect a charger to charge electronic device 100, and also for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0061] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0062] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0063] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 193, camera 192, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0064] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0065] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0066] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0067] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 193. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0068] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0069] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), etc.

[0070] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0071] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0072] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.

[0073] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.

[0074] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0075] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0076] The sensor module 180 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0077] A pressure sensor is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor may be located on the display screen 193. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to the pressure sensor, the capacitance between the electrodes changes. The electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to the display screen 193, the electronic device 100 detects the intensity of the touch operation based on the pressure sensor. The electronic device 100 may also calculate the touch position based on the detection signal from the pressure sensor. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities may correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS message is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS message is executed.

[0078] A gyroscope sensor can be used to determine the motion attitude of an electronic device 100. In some embodiments, the gyroscope sensor can determine the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes). The gyroscope sensor can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor detects the angle of the electronic device 100's movement, calculates the distance the lens module needs to compensate based on the angle, and allows the lens to counteract the movement of the electronic device 100 through reverse motion, thus achieving image stabilization. The gyroscope sensor can also be used in navigation and motion-sensing gaming scenarios.

[0079] A barometric pressure sensor is used to measure air pressure. In some embodiments, the electronic device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor to assist in positioning and navigation.

[0080] An accelerometer can detect the magnitude of acceleration of an electronic device 100 in various directions (typically three axes). When the electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of the electronic device and is applied to applications such as screen orientation switching and pedometers.

[0081] A distance sensor is used to measure distance. Electronic device 100 can measure distance using infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor to measure distance for rapid focusing.

[0082] The proximity sensor may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 100 emits infrared light outward through the LED. The electronic device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that an object is near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that no object is near the electronic device 100. The electronic device 100 may use the proximity sensor to detect when a user holds the electronic device 100 close to their ear for a phone call, so as to automatically turn off the screen to save power. The proximity sensor can also be used in holster mode and pocket mode for automatic unlocking and screen locking.

[0083] An ambient light sensor is used to sense ambient light intensity. In some embodiments, the electronic device 100 can determine the exposure time of an image based on the ambient light intensity sensed by the ambient light sensor. In some embodiments, the electronic device 100 can adaptively adjust the brightness of the display screen 193 based on the sensed ambient light intensity. The ambient light sensor can also be used to automatically adjust the white balance when taking a picture. The ambient light sensor can also work in conjunction with a proximity sensor to detect whether the electronic device 100 is in a pocket to prevent accidental touches.

[0084] A fingerprint sensor is used to collect fingerprints. Electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.

[0085] A temperature sensor is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor to execute a temperature handling strategy. For example, when the temperature reported by the temperature sensor exceeds a threshold, the electronic device 100 performs thermal protection by reducing the performance of a processor located near the temperature sensor to reduce power consumption. In other embodiments, when the temperature is below another threshold, the electronic device 100 heats the battery to prevent abnormal shutdown of the electronic device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, the electronic device 100 boosts the battery's output voltage to prevent abnormal shutdown caused by low temperature.

[0086] A touch sensor, also known as a "touch device," can be located on the display screen 193. The touch sensor and the display screen 193 together form a touchscreen, also known as a "touchscreen." The touch sensor detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 193. In some embodiments, the touch sensor may also be located on the surface of the electronic device 100, in a different position than the display screen 193.

[0087] Bone conduction sensors can acquire vibration signals. In some embodiments, bone conduction sensors can acquire vibration signals from vibrating bone fragments in the human vocal cords. Bone conduction sensors can also contact the human pulse to receive blood pressure signals.

[0088] The event camera 190 can capture dynamic changes in the scene, generate a spatiotemporal data stream, and compress the acquired spatiotemporal data stream into frames according to a set time resolution to form event data for output. The electronic device 100 processes the event data output by the event camera 190 through the processor 110 to achieve target tracking or gesture recognition, etc. For example, when a user unlocks the electronic device by setting a gesture, the event camera 190 can be used to capture the user's gesture changes and output event data. The processor 110 is used to determine the user's gesture based on the event data output by the event camera 190. If the gesture used by the user is consistent with the set gesture for unlocking the electronic device, the desktop of the electronic device is displayed.

[0089] For example, the event camera 190 may include a plurality of light sensors and an event generator coupled to the light sensors for sensing dynamic changes in brightness in the scene. The plurality of light sensors are arranged in a matrix of rows and columns, and each light sensor is associated with a row value and a column value. Taking one of the light sensors as an example, the light sensor includes a photodiode connected in series with a resistor between a source voltage and a ground voltage. The voltage across the photodiode is proportional to the intensity (i.e., brightness) of the light incident on the light sensor.

[0090] The light sensor includes a first capacitor connected in parallel with a photodiode. Therefore, the voltage across the first capacitor is the same as the voltage across the photodiode and is proportional to the intensity of the light detected by the light sensor. The light sensor also includes a switch coupled between the first and second capacitors. The second capacitor is coupled between the switch and ground. Therefore, when the switch is closed, the voltage across the second capacitor is the same as the voltage across the first capacitor and is proportional to the intensity of the light detected by the light sensor. When the switch is open, the voltage across the second capacitor is fixed at the voltage level when the switch was last closed.

[0091] The voltages on the first and second capacitors are fed to a comparator. When the difference between the voltages on the first and second capacitors is less than a threshold value, the comparator outputs a unchanged voltage. When the voltage on the first capacitor is at least as high as the threshold value, the comparator outputs a rising voltage. When the voltage on the first capacitor is at least as low as the threshold value, the comparator outputs a falling voltage. When the comparator outputs a unchanged voltage, the event generator performs no operation, indicating that the brightness of the pixels in the real-world scene of the light sensor has not changed. When the comparator outputs a rising or falling voltage, the event generator receives the signal output by the comparator and generates a corresponding event by combining the current time with the row and column values ​​associated with the light sensor.

[0092] Electronic device 100 implements display functions through a graphics processing unit (GPU), a display screen 193, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 193 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0093] Display screen 193 is used to display images, videos, etc. Display screen 193 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N1 displays 193, where N1 is a positive integer greater than 1.

[0094] Electronic device 100 can perform shooting functions through an image signal processing unit (ISP), camera 192, video codec, GPU, display screen 193, and application processor.

[0095] The ISP (Image Signal Processor) is used to process data fed back from the camera 192. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set within the camera 192.

[0096] Camera 192 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, processor 110 can trigger the camera 192 to start according to a program or instruction in internal memory 121, thereby enabling camera 192 to acquire at least one image and perform corresponding processing on at least one image according to the program or instruction, such as removing rotational blur, removing translation blur, de-mosaic, denoising, or enhancement processing, as well as image post-processing. After processing, the processed image can be displayed on display screen 193. In some embodiments, electronic device 100 may include one or N2 cameras 192, where N2 is a positive integer greater than 1. For example, electronic device 100 may include at least one front-facing camera and at least one rear-facing camera. For example, electronic device 100 may also include side-facing cameras. In one possible implementation, electronic device 100 may include two rear-facing cameras, for example, a main camera and a telephoto camera; or, electronic device 100 may include three rear-facing cameras, for example, a main camera, a wide-angle camera, and a telephoto camera; or, electronic device 100 may include four rear-facing cameras, for example, a main camera, a wide-angle camera, a telephoto camera, and a mid-range camera.

[0097] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0098] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0099] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0100] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as a camera application), etc. The data storage area may store data created during the use of electronic device 100 (such as images captured by a camera), etc. Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory disposed in the processor. Internal memory 121 may also store corresponding data of the neural network model provided in the embodiments of this application. Internal memory 121 may also store code for processing event data output by the event camera through the neural network model. When the code stored in the internal memory 121, which processes the event data output by the event camera using a neural network model, is executed by the processor 110, the event data processing function is implemented through the neural network model. Of course, the corresponding data of the neural network model provided in this embodiment, and the code for processing the event data output by the event camera using the neural network model, can also be stored in external memory. In this case, the processor 110 can run the corresponding data of the neural network model stored in external memory and the code for processing the event data output by the event camera using the neural network model through the external memory interface 120 to implement the corresponding event data processing function.

[0101] Electronic devices may also include buttons, such as power buttons and volume buttons. These buttons can be mechanical or touch-sensitive. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of the electronic device 100.

[0102] Electronic devices may also include motors that can generate vibration alerts. Motors can be used for incoming call vibration alerts or for touch vibration feedback.

[0103] This application embodiment is mainly used for processing event data collected by an event camera. The event camera captures dynamic changes in a scene based on an event-driven approach, which can be understood as capturing changes in pixel brightness within the scene; that is, the event camera outputs the changes in pixel brightness. Specifically, when the brightness of a pixel in a real-world scene changes, the event camera generates an event. The data corresponding to an event can include four parts: (t, x, y, p). Here, x and y are the pixel coordinates of the event in two-dimensional space, i.e., the row and column values ​​of the light sensor corresponding to the pixel whose brightness changed; t is the timestamp of the event, i.e., the time when the pixel's brightness changed; and p is the polarity of the event, representing whether the brightness change is an increase or a decrease.

[0104] like Figure 2 As shown, when objects in a real-world scene change, the event camera generates a spatiotemporal data stream of a series of events. Each row represents an event, the first column is the timestamp of the event, the second column is the x-coordinate, the third column is the y-coordinate, and the fourth column is the polarity of the event, where "0" indicates a decrease in pixel brightness and "1" indicates an increase in pixel brightness.

[0105] The event data processing method provided in this application can be widely applied to fields such as image reconstruction, target tracking, and gesture recognition. In various application areas, post-processing of event data acquired by event cameras is required, such as target classification and target location recognition. To facilitate post-processing of event data, event cameras typically compress the acquired spatiotemporal data stream into frames according to a set time resolution, forming event image frames for output. That is, event cameras can output event data at a fixed frame rate. In different application areas, pre-trained neural network models can be used to perform different post-processing on the fixed frame rate event data output by the event camera to complete corresponding data processing tasks, such as target classification or target location recognition.

[0106] Generally speaking, for a neural network model used to complete a specific task, the source domain event data used to train the neural network model and the target domain event data input into the neural network model in actual application need to have the same frame rate, that is, the spatiotemporal features of the two need to match, so that the neural network model can have a better processing effect on the input data in actual application.

[0107] In practical applications, even for the same target task, event cameras will output event data in different spatiotemporal domains depending on the needs of different scenarios. For example, for target classification tasks, some scenarios require outputting event data at 200 frames per second, while others require outputting event data at 500 frames per second. In addition, differences in hardware settings or software parameters between different event cameras will also cause them to output event data in different spatiotemporal domains.

[0108] Figure 3 The diagram shows visualizations of event data collected during the left-hand waving process in different spatiotemporal domains. These visualizations can also be referred to as event image frames. Figure 3 Different time resolutions are used to represent different spatiotemporal domains. Figure 3 The four columns of event image frames (a), (b), (c), and (d) correspond to temporal resolutions of 3ms, 5ms, 10ms, and 15ms, respectively. It can be seen that the event image frames in columns (c) and (d) contain more and denser event points than those in columns (a) and (b). Therefore, it can be concluded that high-temporal-resolution event data contains more and denser event points, while low-temporal-resolution event data contains fewer and sparser event points, exhibiting spatial differences in event data at different frame rates. Figure 3 The first row (1) is the 10th event image frame extracted from the event data at each time resolution, the second row (2) is the 30th event image frame extracted from the event data at each time resolution, and the third row (3) is the 50th event image frame extracted from the event data at each time resolution. By comparing the event image frames in the same row, it can be seen that the event image frames of event data in different spatiotemporal domains exhibit different temporal correlations.

[0109] from Figure 3 As can be seen, event data from different spatiotemporal domains exhibit spatial differences and varying temporal correlations, meaning that there is a spatiotemporal domain mismatch between event data from different spatiotemporal domains.

[0110] In practical applications, event cameras output event data in different spatiotemporal domains depending on the needs of different scenarios. This results in the source domain event data used to train the neural network model and the target domain event data input into the model being event data in different spatiotemporal domains. When these event data are from different spatiotemporal domains, the spatiotemporal mismatch leads to significant differences in the distribution of spatiotemporal features extracted by the neural network model, affecting the accuracy of the neural network model in processing event data in practical applications.

[0111] Based on this, embodiments of this application provide an event data processing method that can be applied to various intelligent application scenarios of dynamic visual sensing devices. This method can be executed by electronic devices, such as... Figure 1 The electronic device 100 shown may be executed by a chip or chip system in the electronic device, or by a processor in the electronic device. For dynamic visual sensing devices, an event camera will be used as an example below. Figure 4 A flowchart illustrating an event data processing method provided in an embodiment of this application is shown. Figure 4 As shown, the method may include the following steps:

[0112] S401, acquire the first event data collected by the event camera.

[0113] For example, in gesture recognition applications, electronic devices collect first event data through an event camera. This first event data is used to indicate dynamic events of a target object. In this application scenario, the target object can be the user's hand.

[0114] S402, the first event data is processed through a neural network model to obtain the first recognition result of the target object.

[0115] The neural network model is trained using at least two types of sample event data. These at least two types of sample event data have different spatiotemporal domains; specifically, any two types of sample event data have different spatiotemporal domains. In other words, the frame rates of any two types of sample event data are different, or the density or number of event points of any two types of sample event data are different. Furthermore, the at least two types of sample event data include sample event data with the same spatiotemporal domain as the first event data. That is, in addition to using event data with a different spatiotemporal domain than the first event data as sample event data, event data with the same spatiotemporal domain as the first event data is also collected as sample event data. When the neural network model trained using these at least two types of sample event data processes event data with the same spatiotemporal domain as the sample event data, the extracted spatiotemporal features are essentially consistent with the distribution of the spatiotemporal features of the event data used to train the neural network model. This effectively overcomes the problem of spatiotemporal domain mismatch between the sample event data used to train the neural network model and the actual event data to be processed, thus improving the processing accuracy of the neural network model for event data in practical applications. Furthermore, it eliminates the need to train a separate neural network model for each spatiotemporal domain, thus reducing resource waste.

[0116] In one possible example, the neural network model may include a feature extraction network and an object recognition network. The feature extraction network is used to extract spatiotemporal features from the first event data, and the object recognition network is used to determine a first recognition result of the target object based on the spatiotemporal features of the first event data.

[0117] The first identification result of the target object can be its location or category. The function implemented by the object recognition network varies depending on the data processing task. For example, in a target classification task, the object recognition network can be a predictive classification network, implemented using a classification neural network, used to classify the target object based on the spatiotemporal features of the first event data, i.e., determining the category corresponding to the target object in the first event data. For example, assuming the target object in the first event data is a hand, the predictive classification network can determine whether the object in the first event data is a hand or another object. In a regression prediction task, the object recognition network can be a predictive regression network, used to determine specific predicted values ​​based on the spatiotemporal features of the first event data, such as determining the specific movement speed of the target object indicated by the first event data. In a target detection task, the object recognition network can be a predictive detection network, implemented using a detection neural network, used to determine the specific pixel coordinates of the target object indicated by the first event data based on the spatiotemporal features of the first event data.

[0118] In one possible example, the neural network model can include a spiking neural network model, or a model based on a spiking neural network, which has the characteristic of low latency in processing input data. Therefore, using a neural network model constructed with spiking neural networks to process event data acquired by an event camera can fully utilize the high temporal resolution of the event data for real-time output, meeting the requirements of high efficiency and low power consumption in computer vision applications.

[0119] For example, the neural network model can be a model pre-trained using source domain sample event data. The spatiotemporal domains of the source domain sample event data and the first event data are different. When processing the first event data output by the event camera for the first time using the neural network model, the event data output by the event camera can be collected first as target domain sample event data. The target domain sample event data and the first event data have the same spatiotemporal domains. By using both source domain and target domain sample event data, the network parameters of the pre-trained neural network model are adjusted, so that the adjusted neural network model can overcome the problem of spatiotemporal domain mismatch between the source domain and target domain sample event data. The process of adjusting the network parameters of the pre-trained neural network model using both types of sample event data can also be called the transfer training process. Using the neural network model with adjusted network parameters to process the first event data output by the event camera can improve processing accuracy and obtain more accurate recognition results.

[0120] In some embodiments, to ensure high processing accuracy for the neural network model when handling event data from multiple different spatiotemporal domains, the sample event data can include source domain event data and multiple target domain event data from different spatiotemporal domains during transfer training of the neural network model pre-trained using source domain sample event data. For example, in addition to collecting target domain event data with the same spatiotemporal domain as the first event data as sample event data, target domain event data with the same spatiotemporal domain as the second event data can also be collected as sample event data. Here, the first event data and the second event data are event data from different spatiotemporal domains.

[0121] The neural network model trained in this way, when processing the first event data, uses a feature extraction network to extract the spatiotemporal features of the first event data, and an object recognition network to determine the first recognition result of the target object indicated by the first event data based on the spatiotemporal features of the first event data. When processing the second event data, the feature extraction network extracts the spatiotemporal features of the second event data, and the object recognition network determines the second recognition result of the target object indicated by the second event data based on the spatiotemporal features of the second event data. The distribution area of ​​feature points in the spatiotemporal features extracted by the feature extraction network for the first event data is basically consistent with the distribution area of ​​feature points in the spatiotemporal features extracted for the second event data, and both are basically consistent with the distribution area of ​​the spatiotemporal features extracted for the source domain sample event data. Therefore, it can achieve matching and alignment of the spatiotemporal features of different target domain event data with the spatiotemporal features of the source domain event data in both spatiotemporal dimensions. This results in both the first recognition result obtained by the object recognition network based on the spatiotemporal features of the first event data and the second recognition result obtained by the object recognition network based on the spatiotemporal features of the second event data having high accuracy.

[0122] In the process of transferring training of a neural network model using at least two types of sample event data, the network parameters of the neural network model are adjusted based on the spatiotemporal discrimination results of at least two types of sample event data and the predicted recognition results of the objects indicated by at least two types of sample event data.

[0123] For example, the training process of a neural network model may include: acquiring sample event data including at least two spatiotemporal domains; wherein, the sample event data of the first spatiotemporal domain has a corresponding spatiotemporal domain label and an object recognition label, and the sample event data of other spatiotemporal domains besides the first spatiotemporal domain has its own corresponding spatiotemporal domain label.

[0124] Using a neural network model, spatiotemporal features of at least two types of sample event data are extracted, and the predicted recognition results of sample event data in the first spatiotemporal domain are determined. A spatiotemporal discriminant network with auxiliary training is used to perform spatiotemporal discriminant analysis on the spatiotemporal features of the at least two types of sample event data, obtaining spatiotemporal discriminant results for the at least two types of sample event data. Based on the spatiotemporal discriminant results of the at least two types of sample event data and their corresponding spatiotemporal labels, a domain discriminant loss value is determined, and the domain discriminant loss value is inverted to obtain a first loss value. Based on the predicted recognition results of the sample event data in the first spatiotemporal domain and their corresponding object recognition labels, a second loss value is determined. Based on the first and second loss values, the network parameters of the neural network model are adjusted. For example, the total loss value can be determined by the weighted sum of the first and second loss values, and the network parameters of the neural network model can be adjusted based on the total loss value.

[0125] The sample event data in the first spatiotemporal domain can be considered as source domain sample event data, and the sample event data in other spatiotemporal domains besides the first spatiotemporal domain can be considered as target domain sample event data. In some embodiments, the target domain sample event data may include only event data from one spatiotemporal domain. For example, assuming the source domain sample event data has a time resolution of 5ms, if the trained neural network model needs to handle event data with a time resolution of 50ms, then several event data with a time resolution of 50ms without object identification labels can be collected as target domain sample event data. The source domain sample event data with object identification labels and the target domain sample event data without object identification labels are used to form a training dataset to train the neural network model.

[0126] In other embodiments, the target domain sample event data may include event data with multiple time resolutions. For example, assuming the source domain sample event data has a time resolution of 5ms, if the trained neural network model needs to handle event data with time resolutions of 15ms, 35ms, and 50ms, then for each of these three time resolutions, several event data sets without object identification labels can be collected and used as target domain sample event data. That is, the target domain sample event data includes sample event data from three different target domains. A training dataset is composed of source domain sample event data with object identification labels and target domain sample event data without object identification labels but containing multiple time resolutions. This dataset is used to train the neural network model, which can handle event data with multiple time resolutions and achieve high processing accuracy.

[0127] The above example only uses sample event data from three different target domains. In actual applications, the training dataset may contain sample event data from more or fewer target domains, and this application does not limit this.

[0128] During the training process described above, the feature extraction network of the neural network model is used to extract the spatiotemporal features of at least two types of sample event data, and the object recognition network of the neural network model is used to determine the predicted recognition result of the object indicated by the sample event data in the first spatiotemporal domain among the at least two types of sample event data.

[0129] In one possible example, the spatiotemporal discriminant network used during training can be a spiking neural network. The output layer of the spatiotemporal discriminant network includes the same number of spiking neurons as the types of spatiotemporal domains of the sample event data, with different spiking neurons corresponding to different spatiotemporal domains. The spatiotemporal domain features of the sample event data are input into the spatiotemporal discriminant network, resulting in probability values ​​output by N spiking neurons. The probability value output by the first spiking neuron represents the probability that the spatiotemporal domain of the sample event data from which the sample spatiotemporal domain features originate is the spatiotemporal domain corresponding to the first spiking neuron. The first spiking neuron can be any one of the N spiking neurons, where N is the number of spiking neurons in the output layer. Based on the probability values ​​output by the N spiking neurons and their corresponding spatiotemporal domains, the spatiotemporal domain discrimination result of the sample event data can be obtained.

[0130] The neural network model in this application embodiment is trained using at least source domain sample event data and target domain sample event data. The target domain sample event data has the same frame rate as the target domain event data input into the neural network model in actual applications, while the source domain sample event data has a different frame rate. Training the neural network model using both source and target domain sample event data ensures that the spatiotemporal features extracted by the trained feature extraction network are substantially consistent between the source and target domain event data. This achieves spatiotemporal feature matching and alignment between the source and target domain event data in both dimensions, effectively overcoming the problem of spatiotemporal mismatch between the source and target domain event data and improving the processing accuracy of the neural network model for event data in practical applications.

[0131] To better understand the embodiments of this application, the training process of a neural network model provided in this application embodiment is described in detail below. This neural network model includes a feature extraction network and an object recognition network. The training process is illustrated using the object recognition network as a prediction and classification network as an example. The predicted object recognition result output by the prediction and classification network can be referred to as the predicted classification result.

[0132] like Figure 5 As shown, in the model training process, the network architecture used, in addition to the feature extraction network 501 and the prediction and classification network 502 of the neural network model 500 to be trained, may also include a spatiotemporal gradient inversion module 503 and a spatiotemporal discriminant network 504 for auxiliary training of the neural network model 500. The neural network model to be trained can be a neural network model that has already been pre-trained using source domain sample event data.

[0133] Feature extraction network 501 extracts spatiotemporal features from source domain sample event data and target domain sample event data, and inputs the extracted spatiotemporal features into spatiotemporal gradient inversion module 503 and prediction classification network 502, respectively. Spatiotemporal gradient inversion module 503 transmits the spatiotemporal features input from feature extraction network 501 to spatiotemporal discriminator network 504, and inverts the sign of the training gradient returned by spatiotemporal discriminator network 504 before transmitting it to neural network model 500. Spatiotemporal discriminator network 504 determines whether the input spatiotemporal features originate from source domain sample event data or target domain sample event data from a spatiotemporal dual-dimensional perspective, obtaining a spatiotemporal discriminator result, and backpropagates the first training gradient obtained based on the spatiotemporal discriminator result to spatiotemporal gradient inversion module 503. Prediction classification network 502 predicts the category of target objects in source domain sample event data based on the spatiotemporal features of source domain sample event data, obtaining a prediction classification result. The total training gradient is determined based on the second training gradient obtained from the predicted classification result and the inverted first training gradient. The network parameters of neural network model 500 are then updated based on this total training gradient to complete the training process. Specifically, the first training gradient is determined by the loss value calculated based on the spatiotemporal domain discrimination result; the first training gradient is inverted to obtain the first loss value. The second training gradient is calculated based on the predicted classification result, and the second loss value can be determined based on this second training gradient. Therefore, it can also be said that the network parameters of neural network model 500 are updated based on the first and second loss values.

[0134] The following section provides a detailed explanation of the neural network model and the various modules used in the model training process.

[0135] like Figure 6 As shown, the feature extraction network 501 is constructed based on a deep spiking neural network, comprising multiple spiking neural network layers, used to extract high-dimensional spatiotemporal features from the input event data. The neurons in the spiking neural network layers are built based on the LIF (Leaky Integrity and Fire) neuron model, and the function of the spiking neural network layers can be described by the following recursive iterative formula:

[0136]

[0137] Where n is the neuron number; W is the synaptic weight of the neuron; p represents the neuron's membrane potential, which is a continuous value, while the output z can only be a binary value, i.e., whether a pulse is triggered; z n,t This represents the output of the nth neuron at time t; e -dt / σ The leakage effect of the membrane potential is represented; the trigger function f(x) is a step function that f(x) = 1 when x > 0, and f(x) = 0 otherwise.

[0138] The LIF neuron model in spiking neural networks combines all the behaviors of spiking neurons, such as integration, firing, and resetting. It is suitable for processing time-series event data and can extract the characteristics of event data from both spatial and temporal dimensions, that is, extract the spatiotemporal features of event data.

[0139] A feature extraction network 501, comprising multiple LIF neurons, is used to extract spatiotemporal features from the input event data. During model training, source domain sample event data is input into the feature extraction network 501, which then outputs the spatiotemporal features of the source domain sample event data; similarly, target domain sample event data is input into the feature extraction network 501, which then outputs the spatiotemporal features of the target domain sample event data.

[0140] The spatiotemporal gradient inversion module 503 is used to perform a forward identity mapping of spatiotemporal features during forward propagation and to invert the sign of the training gradients during backward propagation. The process of spatiotemporal features being transmitted from the feature extraction network 501 to the spatiotemporal discriminator network 504 via the spatiotemporal gradient inversion module 503, along with the process of spatiotemporal features being transmitted from the feature extraction network 501 to the prediction classification network 502, is collectively referred to as forward propagation. The process of training gradients being transmitted from the spatiotemporal discriminator network 504 to the feature extraction network 501 via the spatiotemporal gradient inversion module 503, along with the process of training gradients being transmitted from the prediction classification network 504 to the feature extraction network 501, is collectively referred to as backward propagation. Figure 7 As shown, during forward propagation, the spatiotemporal gradient inversion module 503 transmits the input spatiotemporal feature I unchanged to the spatiotemporal discriminant network 504, a process known as the forward identity mapping of spatiotemporal features. During backward propagation, the spatiotemporal gradient inversion module 503 inverts the sign of the first training gradient H returned by the spatiotemporal discriminant network 504, and continues to propagate the resulting -H backward. The spatiotemporal gradient inversion module 503 is implemented using a function that can realize the above logic, and does not require parameter updates during model training. It is used for cascading between the feature extraction network 501 and the spatiotemporal discriminant network 504.

[0141] The spatiotemporal discriminant network 504 can employ a classification network, specifically a multilayer spiking neural network used for classification. The spatiotemporal discriminant network 504 is used to discriminate between input spatiotemporal features derived from source domain sample event data or target domain sample event data, based on a dual spatiotemporal dimension.

[0142] In some embodiments, such as Figure 8 As shown, the spatiotemporal discriminant network 504 can be constructed based on a multi-layer spiking neural network. Its output layer includes two spiking neurons, corresponding to the source domain and the target domain, respectively. One spiking neuron outputs the probability value that the spatiotemporal feature originates from the source domain sample event data, and the other spiking neuron outputs the probability value that the spatiotemporal feature originates from the target domain sample event data. The target domain sample event data includes event data at one or more time resolutions, but all are considered as one type of target domain data. The two spiking neurons can be identified with different labels; for example, the spiking neuron with label 0 corresponds to the source domain, and the spiking neuron with label 1 corresponds to the target domain. The label of the spiking neuron with the larger output probability value is used as the spatiotemporal discriminant result for the corresponding spatiotemporal feature. If the spatiotemporal discriminant result is 0, it indicates that the corresponding spatiotemporal feature originates from the source domain sample event data; if the spatiotemporal discriminant result is 1, it indicates that the corresponding spatiotemporal feature originates from the target domain sample event data.

[0143] In other embodiments, if the target domain sample event data includes event data with multiple temporal resolutions, or if the training dataset includes sample event data from multiple target domains, then the output layer of the spatiotemporal discriminant network 504 may include multiple spiking neurons, the number of which is the same as the sum of the number of source and target domains. For example, suppose the training dataset contains sample event data from three different target domains: the first target domain sample event data has a temporal resolution of 15ms, the second target domain sample event data has a temporal resolution of 35ms, and the third target domain sample event data has a temporal resolution of 50ms. Figure 9As shown, the output layer of the spatiotemporal discriminant network 504 includes four spiking neurons, corresponding to the source domain and three target domains, respectively. The first spiking neuron outputs the probability value that the spatiotemporal features originate from sample event data in the source domain; the second spiking neuron outputs the probability value that the spatiotemporal features originate from sample event data in the first target domain; the third spiking neuron outputs the probability value that the spatiotemporal features originate from sample event data in the second target domain; and the fourth spiking neuron outputs the probability value that the spatiotemporal features originate from sample event data in the third target domain. Similarly, the four spiking neurons can be identified with different labels. For example, the spiking neuron labeled 0 corresponds to the source domain, the spiking neuron labeled 1 corresponds to the first target domain, the spiking neuron labeled 2 corresponds to the second target domain, and the spiking neuron labeled 3 corresponds to the third target domain. The labels of the spiking neurons with higher output probability values ​​are used as the spatiotemporal discrimination results of the corresponding spatiotemporal features. If the spatiotemporal discrimination result is 0, it means that the corresponding spatiotemporal feature comes from the source domain sample event data; if the spatiotemporal discrimination result is 1, it means that the corresponding spatiotemporal feature comes from the first target domain sample event data; if the spatiotemporal discrimination result is 2, it means that the corresponding spatiotemporal feature comes from the second target domain sample event data; if the spatiotemporal discrimination result is 3, it means that the corresponding spatiotemporal feature comes from the third target domain sample event data.

[0144] The loss value of the loss function is calculated based on the cross-entropy of the spatiotemporal domain discrimination result of the obtained spatiotemporal domain features and the label of the domain where the input sample event data is located. This loss value is then backpropagated as the first training gradient H to the spatiotemporal domain gradient reversal module 503 as part of the loss function for joint training with the prediction classification network 502.

[0145] The predictive classification network 502 can also be a classification network, specifically a multi-layer spiking neural network used for classification. The predictive classification network 502 is used to predict the category of a target object in the source domain sample event data based on the spatiotemporal features of the source domain sample event data, thus obtaining the predicted classification result. For example... Figure 10 As shown, the prediction classification network 502 can be constructed based on a multi-layer spiking neural network. Its output layer includes multiple spiking neurons, the number of which is determined by the number of target object categories; that is, the number of spiking neurons is consistent with the number of categories of target object labels. Each spiking neuron corresponds to a category of the target object and is used to output the probability value of the target object in the sample event data belonging to that category. The spiking neurons of the prediction classification network 502 can also be labeled with different numbers, and the label of the spiking neuron with the larger output probability value is used as the prediction classification result of the target object in the corresponding source domain sample event data.

[0146] For example, in some embodiments, the predictive classification network 502 can be used to predict whether a target object in the source domain sample event data is a hand or not a hand. In other embodiments, the predictive classification network 502 can be used to predict whether a hand in the source domain sample event data is waving to the left or to the right. In still other embodiments, the predictive classification network 502 can also be used to predict whether the waving speed of a hand in the source domain sample event data is fast, relatively fast, moderate, relatively slow, or very slow.

[0147] The loss value of the loss function is calculated based on the mean squared error between the predicted classification result of the target object and the category label of the target object in the input source domain sample event data. This loss value is used as the second training gradient, which is another part of the loss function for joint training. Based on the second training gradient and the first training gradient inverted by the spatiotemporal gradient inversion module 503, the total training gradient is determined. The network parameters of the feature extraction network 501 and the prediction classification network 502 in the neural network model 500 are updated based on the total training gradient to complete the training process of the neural network model 500.

[0148] The training process of the neural network model provided in the embodiments of this application is described in detail below. For example... Figure 11 As shown, the training process may include the following steps:

[0149] S1101, Obtain the training dataset including source domain sample event data and target domain sample event data.

[0150] The source domain sample event data can be event data stored in a dataset obtained from a public server via a network. The target domain sample event data can be event data collected at the appropriate time resolution from the target domain as needed for processing. Since the source domain sample event data uses existing event data, its data volume is relatively large, while the target domain sample event data is event data collected as needed, and therefore its data volume is relatively small.

[0151] The sample event data in the training dataset all have domain labels to indicate whether the corresponding sample event data is from the source domain or the target domain. The source domain sample event data also carries a category label to indicate the category to which the target object belongs within the corresponding source domain sample event data. Since the target domain sample event data is event data collected as needed, it does not have category labels.

[0152] In some embodiments, the target domain sample event data may include only event data with one time resolution. For example, the source domain sample event data may be event data with a time resolution of 5ms, and the target domain sample event data may be event data with a time resolution of 50ms. In other embodiments, if multiple different target domain event data need to be processed, the target domain sample event data may include event data with multiple time resolutions. For example, the source domain sample event data may be event data with a time resolution of 5ms, and the target domain sample event data may include event data from three target domains, with time resolutions of 15ms, 35ms, and 50ms respectively.

[0153] S1102, randomly sample event data from the training dataset.

[0154] S1103, the extracted sample event data is input into the feature extraction network of the neural network model to be trained, and the spatiotemporal features of the sample event data output by the feature extraction network are obtained.

[0155] S1104, the spatiotemporal features of the sample event data are forward propagated to the spatiotemporal discriminant network through the spatiotemporal gradient inversion module, and the spatiotemporal discriminant results of the sample event data output by the spatiotemporal discriminant network are obtained.

[0156] During forward propagation, the spatiotemporal gradient inversion module does not alter the spatiotemporal characteristics of the transmitted data.

[0157] In some embodiments, if the target domain sample event data includes only event data at one time resolution, the spatiotemporal domain discrimination result output by the spatiotemporal domain discrimination network is used to indicate whether the corresponding sample event data belongs to the source domain sample event data or the target domain sample event data. In other embodiments, if the target domain sample event data includes event data from multiple target domains, the spatiotemporal domain discrimination result output by the spatiotemporal domain discrimination network is used to indicate whether the corresponding sample event data belongs to the source domain sample event data or the sample event data of a specific target domain.

[0158] S1105, determine the first training gradient based on the obtained spatiotemporal domain discrimination results and the domain labels of the sample event data.

[0159] The loss value of the loss function is calculated based on the cross-entropy between the spatiotemporal domain discrimination result output by the spatiotemporal domain discrimination network and the domain label of the input sample event data, and this loss value is used as the first training gradient.

[0160] S1106 uses the spatiotemporal gradient inversion module to invert the sign of the first training gradient and propagate it back to the neural network model.

[0161] S1107, input the spatiotemporal features of the source domain sample event data into the prediction classification network of the neural network model to be trained, and obtain the prediction classification result output by the prediction classification network.

[0162] The prediction classification results output by the prediction classification network are used to predict the category to which the target object belongs in the source domain sample event data.

[0163] S1108, determine the second training gradient based on the obtained predicted classification results and the object category labels of the sample event data.

[0164] The loss value of the loss function is calculated based on the mean squared error between the predicted classification result output by the prediction classification network and the corresponding category label of the input source domain sample event data, and this loss value is used as the second training gradient.

[0165] S1109, determine the total training gradient based on the second training gradient and the first training gradient after sign inversion.

[0166] S1110: Based on the total training gradient, determine whether the neural network model has converged; if yes, execute S1112; if no, execute S1111.

[0167] The neural network model is considered to have converged if the total training gradient, i.e. the total loss value of the neural network model, converges to the preset expected value, or if the magnitude of the change in the total training gradient converges to the preset expected value.

[0168] Step S1111: Adjust the network parameters of the neural network model according to the total training gradient.

[0169] The network parameters of the feature extraction network and the prediction classification network of the neural network model are adjusted separately based on the total training gradient. Optionally, when adjusting the network parameters of the neural network model, the network parameters of the spatiotemporal discriminant network can also be adjusted based on the first training gradient.

[0170] After adjusting the network parameters of the neural network model, return to execute S1102 to continue the next round of training.

[0171] S1112, use the current network parameters as the network parameters of the neural network model to obtain the trained neural network model.

[0172] During the training process described above, S1107 and S1108 can be executed before S1104, or in parallel with S1104.

[0173] Before model training, the spatial distribution and temporal correlation of the spatiotemporal features of the source domain sample event data and the target domain sample event data extracted by the feature extraction network in the neural network model are different. Their temporal correlation is also different. During model training, the first training gradient is inverted using the spatiotemporal gradient inversion module and then backpropagated to adjust the network parameters of the feature extraction network. This makes the spatial distribution and temporal correlation of the spatiotemporal features of the source domain sample event data and the target domain sample event data extracted by the feature extraction network increasingly similar. After training, the spatial distribution and temporal correlation of the spatiotemporal features of the source domain sample event data and the target domain sample event data extracted by the feature extraction network in the neural network model are very similar.

[0174] Figure 12 The spatial distribution of spatiotemporal features of source domain sample event data and target domain sample event data is shown. Specifically, a dot-based feature extraction network is used to extract features from the source domain sample event data, yielding the spatiotemporal features; a cross-based feature extraction network is used to extract features from the target domain sample event data, yielding the spatiotemporal features. Figure 12 (a) shows the spatial distribution of spatiotemporal features of source domain sample event data and target domain sample event data extracted by the feature extraction network before model training. Figure 12 (b) shows the spatial distribution of spatiotemporal features of the source domain sample event data and the target domain sample event data extracted by the feature extraction network after model training. Figure 12 It can be seen that before model training, there is a significant difference between the spatial distribution of the spatiotemporal features of the source domain sample event data extracted by the feature extraction network and the spatial distribution of the spatiotemporal features of the target domain sample event data; however, after model training, the spatial distribution of the spatiotemporal features of the source domain sample event data extracted by the feature extraction network is very close to the spatial distribution of the spatiotemporal features of the target domain sample event data.

[0175] Because the spatial distribution and temporal correlation of the spatiotemporal features of the source domain sample event data and the target domain sample event data extracted by the feature extraction network in the trained neural network model are very similar, it is possible to achieve spatiotemporal matching and alignment of the features of the source domain event data and the target domain event data, thus enabling spatiotemporal adaptation of the event data between the source and target domains. If the target domain sample event data includes sample event data from multiple different target domains, then the spatial distribution and temporal correlation of the spatiotemporal features of each target domain sample event data extracted by the trained feature extraction network are relatively close to those of the source domain sample event data. Even if the neural network model only learns the category labels of the source domain sample event data, and the target domain sample event data has no category labels, the neural network model can still be used to perform target classification tasks on the target domain event data with high accuracy because of the spatiotemporal adaptation of the event data between the source and target domains.

[0176] In other embodiments, the predictive classification network in the neural network model can also be replaced by a regression network or a detection network. During the model training phase, the regression network is used to predict the specific predicted value of the target object in the source domain sample event data; the detection network is used to predict the specific coordinate value of the target object in the source domain sample event data. In this embodiment, the training process of the neural network model can be performed in accordance with the above training process, and will not be repeated here.

[0177] The neural network model trained through the above process can achieve high processing accuracy not only for source domain event data but also for target domain event data. If the target domain sample event data includes sample event data from multiple different target domains, the trained neural network model can achieve high processing accuracy for event data from each target domain.

[0178] Based on the same inventive concept as the methods described above, such as Figure 13 As shown in the illustration, this application also provides an event data processing apparatus 1300. The event data processing apparatus is applied in electronic devices capable of processing event data, such as those used in… Figure 1 The electronic device 100 shown may include an event camera, and the event data processing device can be used to implement the functions of the above-described method embodiments, thus achieving the beneficial effects of the above-described method embodiments. The event data processing device may include a data acquisition unit 1301 and a data processing unit 1302. The event data processing device 1300 is used to implement the above-described... Figure 4 The function shown in the method embodiment. When the event data processing device 1300 is used to implement Figure 4 The function of the method embodiment shown is as follows: the data acquisition unit 1301 can be used to execute S401, and the data processing unit 1302 can be used to execute S402.

[0179] For example, the data acquisition unit 1301 is used to acquire first event data collected by the dynamic visual sensing device. The first event data is used to indicate the dynamic events of the target object.

[0180] The data processing unit 1302 is used to process the first event data through a neural network model to obtain the first recognition result of the target object.

[0181] The neural network model is trained using at least two types of sample event data, the at least two types of sample event data have different spatiotemporal domains, and the at least two types of sample event data include sample event data with the same spatiotemporal domain as the first event data.

[0182] In one possible implementation, the frame rates of the at least two types of sample event data are different, or the density or number of event points in the at least two types of sample event data are different.

[0183] In one possible implementation, the distribution area of ​​feature points in the spatiotemporal domain features extracted by the neural network model for the first event data is consistent with the distribution area of ​​feature points in the spatiotemporal domain features extracted for the second event data. The spatiotemporal domain of the second event data is the same as the spatiotemporal domain of one of the sample event data of at least two sample event data. The spatiotemporal domains of the first event data and the second event data are different.

[0184] In one possible implementation, the neural network model includes a spiking neural network model.

[0185] In one possible implementation, the network parameters of the neural network model are obtained by adjusting the spatiotemporal discrimination results of at least two types of sample event data and the predicted recognition results of the objects indicated by at least two types of sample event data.

[0186] In one possible implementation, the network parameters of the neural network model are obtained by adjusting a first loss value and a second loss value; wherein the first loss value is obtained by reversing the positive and negative values ​​of the loss values ​​obtained from the spatiotemporal discrimination results of at least two sample event data, and the second loss value is obtained from the prediction and recognition results of the objects indicated by at least two sample event data.

[0187] In one possible implementation, the neural network model includes a feature extraction network and an object recognition network; the feature extraction network is used to extract spatiotemporal features of at least two types of sample event data, and the object recognition network is used to determine the predicted recognition result of the object indicated by at least two types of sample event data.

[0188] The neural network model in this embodiment is trained using sample event data from at least two different spatiotemporal domains. The spatiotemporal features of the sample event data from at least two different spatiotemporal domains extracted by the feature extraction network are basically consistent, achieving matching and alignment of spatiotemporal features of sample event data from different spatiotemporal domains in both spatiotemporal dimensions. This effectively overcomes the problem of spatiotemporal mismatch between sample event data from different spatiotemporal domains and improves the processing accuracy of the neural network model for event data in practical applications.

[0189] Based on the same inventive concept as the methods described above, this application also provides an event data processing system, see [link to relevant documentation]. Figure 14 As shown, the event data processing system 1400 includes a processor 1401 and a dynamic vision sensing device 1402. The processor 1401 and the dynamic vision sensing device 1402 can be housed in the same electronic device or in different electronic devices. The dynamic vision sensing device 1402 is used to acquire first event data, wherein the first event data is used to indicate dynamic events of a target object. A more detailed description of the dynamic vision sensing device 1402 can be found above. Figure 1 The description of the event camera 190 shown will not be repeated here. The processor 1401 is connected to the dynamic vision sensing device 1402 and is used to execute... Figure 4 The method shown.

[0190] In some embodiments, the event data processing system 1400 may further include a memory for storing instructions or programs executed by the processor 1401, or storing input data required by the processor 1401 to run the instructions or programs, or storing data generated after the processor 1401 runs the instructions or programs. Figure 4 The functionality shown in the method embodiment. For example, when the event data processing system 1400 is used to implement... Figure 4 In the illustrated method, processor 1401 performs the functions of the data acquisition unit 1301 and data processing unit 1302. Exemplarily, data acquisition unit 1301 can use processor 1401 to call a program or instruction stored in memory to acquire first event data collected by dynamic visual sensing device 1402, wherein the first event data is used to indicate dynamic events of a target object. Data processing unit 1302 can use processor 1401 to call a program or instruction stored in memory to process the first event data through a neural network model to obtain a first recognition result of the target object. The neural network model is trained using at least two types of sample event data, the at least two types of sample event data having different spatiotemporal domains, and including sample event data with the same spatiotemporal domain as the first event data.

[0191] It should be noted that in some embodiments, the event data processing device may not include an event camera. For example, an event camera interface may be provided, through which the event camera is connected when needed. In other embodiments, the event data processing device may also acquire the event data to be processed via a network or other means. This event data may be collected by the event camera and stored on a network server or other storage medium.

[0192] It is understood that the processor 1401 in the embodiments of this application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.

[0193] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (eEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Additionally, the ASIC can reside in a terminal device. Alternatively, the processor and storage medium can exist as discrete components in the terminal device.

[0194] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).

[0195] In the various embodiments of this application, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments are consistent and can be referenced mutually. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or device is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.

[0196] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative examples of the solutions defined by the appended claims and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of this application.

[0197] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if these modifications and variations of the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

Claims

1. An event data processing method, characterized in that, The method includes: Acquire first event data collected by a dynamic visual sensing device, wherein the first event data is used to indicate dynamic events of the target object; The first event data is processed by a neural network model to obtain the first identification result of the target object; The neural network model is trained based on at least two types of sample event data, wherein the at least two types of sample event data have different spatiotemporal domains, and the at least two types of sample event data include sample event data with the same spatiotemporal domain as the first event data; the at least two types of sample event data have different spatiotemporal domains, which means that the at least two types of sample event data have different frame rates, or that the density or number of event points in the at least two types of sample event data are different.

2. The method as described in claim 1, characterized in that, The distribution area of ​​feature points in the spatiotemporal domain features extracted by the neural network model for the first event data is consistent with the distribution area of ​​feature points in the spatiotemporal domain features extracted for the second event data. The spatiotemporal domain of the second event data is the same as the spatiotemporal domain of one of the at least two sample event data. The spatiotemporal domains of the first event data and the second event data are different.

3. The method as described in claim 1 or 2, characterized in that, The neural network model includes a spiking neural network model.

4. The method as described in claim 1 or 2, characterized in that, The network parameters of the neural network model are obtained by adjusting the spatiotemporal discrimination results of the at least two types of sample event data and the predicted recognition results of the objects indicated by the at least two types of sample event data.

5. The method as described in claim 1 or 2, characterized in that, The network parameters of the neural network model are obtained by adjusting based on the first loss value and the second loss value; The first loss value is obtained by reversing the positive and negative values ​​of the loss values ​​obtained from the spatiotemporal discrimination results of the at least two sample event data, and the second loss value is obtained from the prediction and recognition results of the objects indicated by the at least two sample event data.

6. The method as described in claim 1 or 2, characterized in that, The neural network model includes a feature extraction network and an object recognition network; the feature extraction network is used to extract spatiotemporal features of the at least two types of sample event data, and the object recognition network is used to determine the predicted recognition result of the object indicated by the at least two types of sample event data.

7. An event data processing device, characterized in that, The device includes: The data acquisition unit is used to acquire first event data collected by the dynamic visual sensing device, wherein the first event data is used to indicate dynamic events of the target object; A data processing unit is configured to process the first event data using a neural network model to obtain a first recognition result of the target object; wherein the neural network model is trained based on at least two types of sample event data, the at least two types of sample event data having different spatiotemporal domains, and the at least two types of sample event data including sample event data with the same spatiotemporal domain as the first event data; the at least two types of sample event data having different spatiotemporal domains means that the at least two types of sample event data have different frame rates, or that the density or number of event points in the at least two types of sample event data are different.

8. The apparatus as claimed in claim 7, characterized in that, The network parameters of the neural network model are obtained by adjusting based on the first loss value and the second loss value; The first loss value is obtained by reversing the positive and negative values ​​of the loss values ​​obtained from the spatiotemporal discrimination results of the at least two sample event data, and the second loss value is obtained from the prediction and recognition results of the objects indicated by the at least two sample event data.

9. An event data processing system, characterized in that, include: A dynamic visual sensing device is used to collect first event data, which is used to indicate dynamic events of a target object; A processor, connected to the dynamic vision sensing device, is used to perform the method as described in any one of claims 1 to 6.

10. A computer-readable storage medium storing a computer program therein, characterized in that: When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Neural network training method and device, image processing method and device, equipment and storage medium

    CN111860823A