Intelligent recognition method and related device based on event camera
By using event cameras for information collection and processing in autonomous driving systems and combining them with pulse neural network models, the problems of high sensor cost and high power consumption are solved, and efficient driving decision-making and energy consumption optimization are achieved.
Patent Information
- Application Number
- CN202410391103.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-03-29
AI Technical Summary
In existing autonomous driving technologies, sensors are expensive, have low sampling rates, and consume high computing power, resulting in limited device endurance and high costs.
An event camera is used as the perception unit to collect information and perform pulse supervised learning and feature extraction, combined with scene segmentation and position prediction models to make driving decisions.
It provides a high sampling rate and high dynamic range perception method, reduces the deployment cost of high sampling rate perception units, and builds models through pulse neural networks to save energy and reduce computing resource requirements.
Smart Images

Figure CN118529061B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicles, and in particular to an intelligent recognition method based on an event camera and a related device. Background Art
[0002] Current autonomous driving solutions primarily rely on various sensors to collect signals, transmit them to a central processing unit (CPU), and process the corresponding algorithms to generate driving decisions. Current sensors primarily include cameras, ultrasonic radar, and lidar, and algorithms primarily based on second-generation deep learning algorithms. Achieving autonomous driving accuracy requires continuously expanding the model's parameters and depth. Furthermore, deploying large algorithmic models on devices significantly impacts autonomous driving solutions' range. Furthermore, improving image acquisition accuracy requires deploying a large number of sensors, which drives up sensor costs. This leads to the following challenges with current autonomous driving technology: high sensor costs, low sampling rates, and high computational power consumption. Summary of the Invention
[0003] The embodiments of the present application provide an intelligent recognition method and related devices based on an event camera. By adopting an event camera instead of a traditional camera, it is helpful to solve the problems of high sensor cost, low sampling rate, and high computing power consumption.
[0004] In a first aspect, an embodiment of the present application provides an event camera-based intelligent recognition method, which is applied to a vehicle, wherein the vehicle includes an event camera, and the method includes:
[0005] collecting first information through the event camera;
[0006] Processing the first information to obtain output information, where the output information corresponds to an event of the event camera;
[0007] Processing the output information to obtain a processing result;
[0008] Supervised learning is performed based on the processing results, driving information, and a decision model to obtain driving decision information.
[0009] In a second aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. The computer storage medium stores a computer program, and the computer program includes program instructions, which, when executed by a processor, cause the processor to execute some or all of the steps described in the first aspect.
[0010] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program includes instructions for executing some or all of the steps described in the first aspect of the embodiment of the present application.
[0011] In a fourth aspect, embodiments of the present application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0012] By implementing the embodiments of the present application, first information is collected through the event camera, and after processing the first information, output information is obtained. The output information corresponds to the event of the event camera. The output information is processed to obtain a processing result. Finally, supervised learning is performed based on the processing result, driving information, and a decision model to obtain driving decision information. In this way, by using the event camera as the perception unit, a high sampling rate and high dynamic range solution is provided for intelligent recognition, saving the deployment cost of high sampling rate perception units. At the same time, the back-end algorithm can construct multiple models, which can further save energy consumption and further reduce computing resource requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the background technology, the drawings required for use in the embodiments of the present invention or the background technology will be described below.
[0014] Figure 1 This is a system architecture diagram of an event camera-based intelligent recognition method provided in an embodiment of the present application;
[0015] Figure 2 This is a method flow chart of an intelligent recognition method based on an event camera provided in an embodiment of the present application;
[0016] Figure 3 This is a method flow chart of another intelligent recognition method based on an event camera provided in an embodiment of the present application;
[0017] Figure 4 This is a schematic diagram of an intelligent recognition scene based on an event camera provided in an embodiment of the present application;
[0018] Figure 5 This is a method flow chart of another intelligent recognition method based on an event camera provided in an embodiment of the present application;
[0019] Figure 61 is a schematic structural diagram of an event camera-based intelligent recognition device provided in an embodiment of the present application;
[0020] Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0022] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or electronic device comprising a series of steps or units is not limited to the listed steps or units, but may, in an optional example, also include steps or units not listed, or may, in an optional example, include other steps or units inherent to the process, method, product, or electronic device.
[0023] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0024] In response to the above problems, the present application provides an intelligent recognition method based on event cameras and related devices. By adopting event cameras as sensing units, a high sampling rate and high dynamic range solution is provided for intelligent recognition, saving the deployment cost of high sampling rate sensing units. At the same time, the back-end algorithm constructs multiple models based on pulse neural networks, which can save more energy consumption and further reduce the demand for computing resources.
[0025] See also Figure 1 , Figure 1 This is a system architecture diagram of an event camera-based intelligent recognition method provided in an embodiment of the present application. The event camera-based intelligent recognition system includes: a vehicle 100, an event camera 110, and a processing unit 120.
[0026] Among them, the processing unit 120 is installed in the vehicle, the event camera 110 can be installed on the outside of the vehicle body, and there can be multiple event cameras 110. In actual applications, the event cameras 110 can be arranged around the vehicle 100, so that the resolution angle range of multiple event cameras 110 can fully cover the four sides of the vehicle 100. The arrangement position of the event camera 110 can also be other positions, which is not limited here, and the number of event cameras 110 is also limited here. The resolution angle of the event camera 110 is not limited here and can be selected according to needs.
[0027] Among them, the vehicle 100 may be equipped with other sensors in addition to the event camera 110. The above-mentioned other sensors can also transmit the acquired event information to the processing unit 120 for processing by the processing unit 120. The results obtained after processing can be used as reference information for the results obtained after the data acquired by the event camera 110 is processed by the processing unit 120.
[0028] Among them, the processing unit 120 can obtain event information from the event camera 110. In a possible embodiment, the above-mentioned event information is the first information. The event information can also be transmitted from other sensors, which is not limited here. After being processed by multiple models in the processing unit, a driving decision is obtained, and the vehicle movement is controlled according to the driving decision to achieve the effect of intelligent driving.
[0029] The event camera 110 may acquire event information within a certain resolution angle range, and transmit the acquired event information, namely, first information, to the processing unit 120 , which processes the information to obtain decision information.
[0030] See also Figure 2 , Figure 2 This is a method flow chart of an intelligent recognition method based on an event camera provided in an embodiment of the present application, which is as follows:
[0031] S210: Collect first information through an event camera.
[0032] The method proposed in this embodiment may be based on the leaky integrate and fire (LIF) model. The first information may be event information. When the information acquired by the event camera can cause a sudden change in the neuronal membrane potential in the LIF model, the information is determined to be event information. The neuronal dynamics model of LIF is: Where u(t) is the neuron membrane potential at time t, τ is the time constant, and I(t) represents the presynaptic input from the previous neuron's spike activity or external stimulation. When the neuron membrane potential exceeds a specified threshold, the neuron releases a pulse and resets the membrane potential. The condition for generating an event is when the neuron membrane potential exceeds the specified threshold.
[0033] The event camera may be single or multiple, and the number of event cameras is not limited here. When there are multiple event cameras, the first information collected may also be multiple.
[0034] S220: Process the first information to obtain output information, where the output information corresponds to an event of the event camera.
[0035] Specifically, the first information is subjected to noise reduction training based on a pulse supervised learning network to obtain the second information.
[0036] Specifically, the first information includes positive event information and negative event information. The first information is subjected to noise reduction training by pulse neural network supervised learning to obtain the second information including:
[0037] 21. De-noise the positive event information and the negative event information of the first information through two isomorphic spiking neural network denoisers to obtain de-noised positive event information and de-noised negative event information respectively;
[0038] 22. Merge the denoised positive event information and the denoised negative event information to obtain the denoised first information;
[0039] 23. Minimize the loss between the first information after noise reduction and the reference information to obtain a first loss value, where the reference information is a reference value obtained by the second sensor after noise reduction, and the second sensor is the reference sensor;
[0040] 24. If the first loss value is within a first threshold range, output the denoised first information as the second information.
[0041] Among them, please refer to Figure 3 , Figure 3A flowchart of another method for intelligent recognition based on an event camera provided in this embodiment is shown. As shown in the figure, the first information output by the event camera has positive and negative pixel variations, that is, the first information has positive events and negative events. The positive events and negative events are input into two identically constructed neural network denoising models for denoising, obtaining denoised positive events and denoised negative events. The denoised positive events and denoised negative events are then merged to obtain the denoised first information. Adjacent images captured by cameras other than the event camera are subjected to autoencoder denoising to obtain information at time t and information at time t-1. The information at time t and the information at time t-1 are adjacent information. The information at time t and the information at time t-1 are merged to obtain reference information. This reference information is used to perform loss minimization calculations with the denoised first information obtained by the event camera. The reference information is used as the optimization criterion for the event camera, and the output result can be considered to have an error within the allowable range.
[0042] Among them, the minimization loss is calculated according to the following first formula to obtain the first loss value:
[0043]
[0044] Among them, Loss1 represents the first loss value, y s1 and o s1 The vector representing the reference information of the sth sample and the vector representing the first information after noise reduction are respectively ordered to minimize the mean square error within a given time window T. The loss minimization process ends when the first loss value decreases to within a first threshold range. The above-mentioned first threshold range is a preset acceptable error tolerance range and is not limited here.
[0045] S230: Process the output information to obtain a processing result.
[0046] Specifically, feature extraction is performed on the second information to obtain output information, where the output information corresponds to an event of the event camera.
[0047] Specifically, the feature extraction is pulse self-encoding, and feature extraction is performed on the second information to obtain output information, including: converting the second information into a first pulse signal; and performing feature extraction based on the first pulse signal to obtain output information.
[0048] The feature extraction method is pulse autoencoding feature extraction. The second information obtained after noise reduction undergoes unsupervised self-learning. First, the event information is pulse-encoded and converted into a first pulse signal. Feature extraction is then performed on the pulse-encoded first pulse signal to obtain the characteristic information of the second information, which is then output as the output information. The autoencoding feature extraction method can reduce the dimensionality of the input information, removing redundant information and extracting the key features of the data. This is very useful for processing large amounts of data and reducing computational costs.
[0049] Specifically, the method also includes: processing the second information according to the scene segmentation model to obtain a first output result, processing the second information according to the position prediction model to obtain a second output result, and merging the first output result and the second output result to output as a processing result.
[0050] Among them, the scene segmentation model and the position prediction model are the models that need to be trained. The scene segmentation and position prediction are constructed by two pulse neural networks respectively. Finally, an event merging module fuses the two information and outputs the processing results.
[0051] Before outputting the processing result, the method further includes: performing a minimization loss calculation on the fused first output result and the second output result according to the first formula to obtain a second loss value: Among them, Loss2 represents the second loss value, y s2 and o s2 The vector representing the reference image of the sth sample and the vector representing the fused first and second output results are respectively minimized within a given time window T. The loss minimization process ends when the second loss value falls within a second threshold range, and the fused first and second output results are output as the processing result. The second threshold range is a preset acceptable error range and is not limited here.
[0052] Specifically, the scene segmentation model training method includes: obtaining historical first information, and historical second information corresponding to the historical first information;
[0053] Performing feature extraction on the historical second information to obtain historical output information, where the historical output information corresponds to the event of the event camera;
[0054] Segmenting the historical output information by time to obtain a first feature information set corresponding to a plurality of different time periods, wherein the historical output information is obtained by extracting features from the historical second information, and the historical second information corresponds to the historical first information;
[0055] Mapping the historical output information included in the first feature information set corresponding to time t to obtain a first image, where the first image is the image corresponding to time t;
[0056] Minimizing a loss function calculation on the image and a reference image to obtain a second loss value, where the reference image is an image obtained by other cameras, and the other cameras are cameras other than the event camera;
[0057] When the second loss value is within a second threshold range, outputting the first image as the first output result;
[0058] The scene segmentation model is trained according to the historical output information, the first feature information set, the first image, and the reference image to obtain the scene segmentation model.
[0059] The specific method for obtaining the reference image includes: performing semantic segmentation processing on the first reference information obtained by other cameras, and performing omnidirectional mapping on the result of the semantic segmentation processing to obtain the reference image.
[0060] Specifically, the location prediction model training method includes:
[0061] Acquire historical first information and historical second information corresponding to the historical first information;
[0062] Performing feature extraction of location information on the historical second information to obtain historical output information, where the historical output information corresponds to an event of the event camera;
[0063] Segmenting the historical output information according to the location information to obtain a plurality of second feature information sets, each of the second feature information sets representing the historical output information of different location information;
[0064] Extracting the historical output information of each of the second feature information sets, obtaining a predicted position value according to a position prediction algorithm, and outputting the predicted position value as a second output result;
[0065] The position prediction model is trained according to the multiple historical output information, the second feature information set, and the position prediction algorithm to obtain the position prediction model.
[0066] The position prediction algorithm may be a particle filter method, a least squares method, etc., which is not limited here.
[0067] Among them, the scene segmentation model and the position prediction model can be trained separately or simultaneously. The scene segmentation module and the position prediction module can be assembled according to actual needs to achieve the effect of accelerating calculations.
[0068] S240 , performing supervised learning based on the processing results, driving information, and decision model to obtain driving decision information.
[0069] The driving information may include the vehicle speed and driving direction angle of the vehicle during driving, where the driving direction angle is the angle between the vehicle head direction and the road direction.
[0070] The decision model can be a preset model. After training, the decision model can also maintain a self-learning process to continuously learn information and improve the model. The decision model can be a model based on a spiking neural network or a model based on an artificial neural network, without limitation.
[0071] Among them, when the decision model is a model based on a pulse neural network, the decision model can be constructed by using pulse coding combined with one-hot coding. The specific model is: in a set of pulse sequences with multiple pulses representing execution contents, the first triggered pulse is used as the final decision value, where the condition for triggering each pulse is related to the input content, and the input content is mapped to the pulse of the execution content, thereby selecting the operation decision corresponding to the input content, where the above-mentioned input content is the processing result and driving information.
[0072] For example, if the output goal of the driving decision is to select a matching speed from 0-180 to output the driving decision, the first pulse result of a single event is taken as the final response result. If the decision speed of the first pulse mapped by the input content is 80, then 80 should be the first output result.
[0073] It should be noted that during autonomous driving based on intelligent recognition, the decision module's learning mechanism can be retained to continuously learn the user's driving patterns. Alternatively, during autonomous driving based on intelligent recognition, learning can be omitted and a preset model can be used. When the decision module's parameters change, the parameters can be reset via over-the-air download or other download methods to achieve a better driving method.
[0074] It should be noted that the specific type of pulse neural network selected by each of the above models can be selected according to the actual effect to be achieved, and is not limited here. The various functional modules can be mounted on different chips or on the same chip, and are not limited here.
[0075] As can be seen, in this embodiment, first information is collected via the event camera. Noise reduction training is performed on this first information based on a pulse supervised learning network to obtain second information. Feature extraction is performed on this second information to obtain output information, which corresponds to the event captured by the event camera. This output information is then processed according to a scene segmentation model to obtain a first output result. This output information is then processed according to a position prediction model to obtain a second output result. The first and second output results are then combined to form a processed result. Finally, supervised learning is performed based on the processed result, driving information, and a decision model to obtain driving decision information. Thus, by using an event camera as a sensing unit, a high sampling rate and high dynamic range solution is provided for intelligent recognition, reducing the deployment cost of high sampling rate sensing units. Furthermore, the back-end algorithm builds multiple models based on a pulse neural network, further saving energy consumption and reducing computing resource requirements.
[0076] In one possible embodiment, see Figure 4 , Figure 4 A schematic diagram of an intelligent scene recognition based on an event camera provided in an embodiment of the present application is shown as follows: Figure 4 As shown, the vehicle includes multiple event cameras: event camera A, event camera B, event camera C, event camera D, event camera E, event camera F, event camera G, and event camera H. When there are multiple event cameras, there are also multiple second information corresponding to the first information. It is necessary to construct omnidirectional events for the multiple second information to merge the multiple second information into an event set.
[0077] See also Figure 5 , Figure 5 A flow chart of another method for intelligent recognition based on event cameras provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the method further includes:
[0078] S510, collecting a plurality of first information by using a plurality of event cameras;
[0079] S520, performing noise reduction training on the plurality of first information based on a pulse supervised learning network to obtain a plurality of second information;
[0080] S530, constructing a first omnidirectional information set according to the plurality of second information;
[0081] S540, performing feature extraction on the first omnidirectional information set to obtain output information, where the output information corresponds to an event of the event camera;
[0082] S550: Process the output information according to the scene segmentation model to obtain a first output result, process the output information according to the position prediction model to obtain a second output result, and combine the first output result and the second output result to output as a processing result;
[0083] S560: Perform supervised learning based on the processing results, driving information, and decision model to obtain driving decision information.
[0084] Specifically, the event camera is a plurality of event cameras, wherein each event camera corresponds to a first information respectively. The method further includes: omnidirectionally constructing and processing the second information corresponding to the plurality of first information respectively to obtain a first omnidirectional information set; performing feature extraction on the second information to obtain output information, including: converting the plurality of second information included in the first omnidirectional information set into a second pulse signal; performing feature extraction based on the second pulse signal to obtain output information.
[0085] The first omnidirectional information set includes a plurality of second information, and the same pulse self-encoding and feature extraction operations as those for a single second information are performed on the first omnidirectional information set.
[0086] Specifically, the second information corresponding to the plurality of first information is omnidirectionally constructed to obtain a first omnidirectional information set, which satisfies the following formula: Wherein, event represents the first omnidirectional information set, N represents the number of event cameras, A represents the maximum detection angle of each event camera, and offset represents the angle between the probe of the event camera and the front direction of the vehicle.
[0087] Specifically, the event camera is a plurality of event cameras, wherein each event camera corresponds to a first information respectively. The method further includes: omnidirectionally constructing and processing the second information corresponding to the plurality of first information respectively to obtain a first omnidirectional information set; performing feature extraction on the second information to obtain output information, including: converting the plurality of second information included in the first omnidirectional information set into a second pulse signal; performing feature extraction based on the second pulse signal to obtain output information.
[0088] The feature extraction method uses pulse autoencoding feature extraction. Unsupervised self-learning is performed on the second information within the first omnidirectional information set. The event information is first pulse-encoded, converted into a first pulse signal, and then feature extraction is performed on the pulse-encoded first pulse signal to obtain the characteristic information of the second information, which is then output as the output information. The autoencoding feature extraction method can reduce the dimensionality of the input information, removing redundant information and extracting the key features of the data. This is very useful for processing large amounts of data and reducing computational costs.
[0089] Among them, the scene segmentation model and the position prediction model are the models that need to be trained. The scene segmentation and position prediction are constructed by two pulse neural networks respectively. Finally, an event merging module fuses the two information and outputs the processing results.
[0090] Before outputting the processing result, the method further includes: performing a loss minimization calculation on the fused first output result and the second output result according to the first formula to obtain a second loss value.
[0091] Specifically, in the case where the historical first information is a plurality of historical first information, the scene segmentation model training method includes: obtaining a plurality of historical first information, and a plurality of historical second information corresponding to the plurality of historical first information; omnidirectionally constructing and processing the plurality of historical second information corresponding to the plurality of historical first information to obtain a historical first omnidirectional information set; performing feature extraction on the historical first omnidirectional information set to obtain historical omnidirectional output information, the historical omnidirectional output information corresponding to the event of the event camera; segmenting the historical omnidirectional output information according to time to obtain a plurality of first feature information sets corresponding to different time periods, each first feature information set representing the first feature information set at a different time period. A feature information, historical omnidirectional output information is obtained by extracting features from historical second information, and the historical second information corresponds to the historical first information; all first feature information included in a first feature information set corresponding to time t is omnidirectionally mapped to obtain an omnidirectional image, and the omnidirectional image is the omnidirectional image corresponding to time t; the omnidirectional image and the reference omnidirectional image are minimized by loss function calculation to obtain a second loss value; when the second loss value is within a second threshold range, the omnidirectional image is output as the first output result; a scene segmentation model is trained according to the historical output information, the first feature information set, the omnidirectional image, and the reference omnidirectional image to obtain a scene segmentation model.
[0092] Among them, the specific method of obtaining the reference image includes: merging multiple first reference information obtained by multiple other cameras to obtain a second omnidirectional event set, performing semantic segmentation processing on the second omnidirectional event set, and performing omnidirectional mapping on the results of the semantic segmentation processing to obtain a reference image.
[0093] Specifically, the location prediction model training method includes:
[0094] Acquire historical first information and historical second information corresponding to the historical first information; perform feature extraction of location information on the historical second information to obtain historical output information, where the historical output information corresponds to an event of an event camera; segment the historical output information according to the location information to obtain multiple second feature information sets, where each second feature information set represents historical output information of different location information; extract the historical output information of each second feature information set, and obtain a predicted position value according to a position prediction algorithm, which is output as a second output result; train a position prediction model based on multiple historical output information, the second feature information sets, and the position prediction algorithm to obtain a position prediction model.
[0095] Specifically, the supervision model is a preset model, and supervised learning is performed based on the decision model, processing results and driving information coding to obtain driving decision information, including: mapping the processing results and the driving information coding to decision pulse coding according to the decision model, obtaining the first decision pulse coding among the matched decision pulse coding; and outputting the first decision pulse coding as the driving decision information.
[0096] The decision model can be a preset model. After training, the decision model can also maintain a self-learning process to continuously learn information and improve the model. The decision model can be a model based on a spiking neural network or a model based on an artificial neural network, without limitation.
[0097] Among them, when the decision model is a model based on a pulse neural network, the decision model can be constructed by using pulse coding combined with one-hot coding. The specific model is: in a set of pulse sequences with multiple pulses representing execution contents, the first triggered pulse is used as the final decision value, where the condition for triggering each pulse is related to the input content, and the input content is mapped to the pulse of the execution content, thereby selecting the operation decision corresponding to the input content, where the above-mentioned input content is the processing result and driving information.
[0098] It should be noted that during autonomous driving based on intelligent recognition, the decision module's learning mechanism can be retained to continuously learn the user's driving patterns. Alternatively, during autonomous driving based on intelligent recognition, learning can be omitted and a preset model can be used. When the decision module's parameters change, the parameters can be reset via over-the-air download or other download methods to achieve a better driving method.
[0099] It should be noted that the specific type of pulse neural network selected by each of the above models can be selected according to the actual effect to be achieved, and is not limited here. The various functional modules can be mounted on different chips or on the same chip, and are not limited here.
[0100] As can be seen, in this embodiment, multiple event cameras first collect multiple first information. These first information are then subjected to noise reduction training based on a pulse supervised learning network to obtain multiple second information. A first omnidirectional information set is constructed based on the multiple second information. Feature extraction is performed on the first omnidirectional information set to obtain output information, which corresponds to the events of the event cameras. The output information is then processed according to a scene segmentation model to obtain a first output result. The output information is then processed according to a position prediction model to obtain a second output result. The first and second output results are combined to output a processed result. Finally, supervised learning is performed based on the processed results, driving information, and a decision model to obtain driving decision information. In this way, by using multiple event cameras as perception units, a high sampling rate and high dynamic range solution is provided for intelligent recognition, reducing the deployment cost of high sampling rate perception units. Furthermore, the back-end algorithm builds multiple models based on a pulse neural network, further saving energy consumption and reducing computing resource requirements.
[0101] See Figure 6 , Figure 6 is a structural diagram of an event camera-based intelligent recognition device provided in an embodiment of the present application. The event camera-based intelligent recognition device 600 includes: an acquisition module 610, a noise reduction module 620, a feature extraction module 630, a processing module 640, and a decision module 650, wherein:
[0102] A collection module 610 is configured to collect first information via the event camera;
[0103] a noise reduction module 620 configured to perform noise reduction training on the first information based on a pulse supervised learning network to obtain second information;
[0104] a feature extraction module 630 configured to extract features from the second information to obtain output information corresponding to the event of the event camera;
[0105] a processing module 640 configured to process the output information according to the scene segmentation model to obtain a first output result, process the output information according to the position prediction model to obtain a second output result, and combine the first output result and the second output result to output as a processing result;
[0106] The decision module 650 is configured to perform supervised learning based on the processing results, driving information, and the decision model to obtain driving decision information.
[0107] In one possible implementation, the denoising module 620 performs denoising training on the first information through supervised learning of a spiking neural network to obtain the second information, specifically for:
[0108] Denoising the positive event information and the negative event information of the first information through two isomorphic spiking neural network denoisers to obtain denoised positive event information and denoised negative event information respectively;
[0109] Merging the denoised positive event information and the denoised negative event information to obtain the denoised first information;
[0110] Minimizing the loss calculation between the first information after noise reduction and reference information to obtain a first loss value, wherein the reference information is a reference value obtained by the second sensor after noise reduction, and the second sensor is the reference sensor;
[0111] If the first loss value is within a first threshold range, the noise-reduced first information is output as second information.
[0112] In a possible implementation, the feature extraction module 630 extracts features from the second information to obtain output information, specifically configured to:
[0113] converting the second information into a first pulse signal;
[0114] Feature extraction is performed based on the first pulse signal to obtain output information.
[0115] In a possible implementation, the feature extraction module, in extracting features from the second information to obtain output information, is further configured to:
[0116] Performing omnidirectional construction processing on the second information corresponding to the plurality of first information to obtain a first omnidirectional information set;
[0117] The extracting features of the second information to obtain output information includes: converting the plurality of second information included in the first omnidirectional information set into a second pulse signal;
[0118] Feature extraction is performed based on the second pulse signal to obtain output information.
[0119] In a possible implementation, the processing module 640 is specifically configured to:
[0120] Acquire historical first information and historical second information corresponding to the historical first information;
[0121] Performing feature extraction on the historical second information to obtain historical output information, where the historical output information corresponds to the event of the event camera;
[0122] Segmenting the historical output information by time to obtain a first feature information set corresponding to a plurality of different time periods, wherein the historical output information is obtained by extracting features from the historical second information, and the historical second information corresponds to the historical first information;
[0123] Mapping the historical output information included in the first feature information set corresponding to time t to obtain a first image, where the first image is the image corresponding to time t;
[0124] Minimizing a loss function calculation on the image and a reference image to obtain a second loss value, where the reference image is an image obtained by other cameras, and the other cameras are cameras other than the event camera;
[0125] When the second loss value is within a second threshold range, outputting the first image as the first output result;
[0126] The scene segmentation model is trained according to the historical output information, the first feature information set, the first image, and the reference image to obtain the scene segmentation model.
[0127] In a possible implementation, the processing module 640 is further configured to:
[0128] Acquire the plurality of historical first information and the plurality of historical second information corresponding to the plurality of historical first information;
[0129] Performing omnidirectional construction processing on the plurality of historical second information corresponding to the plurality of historical first information to obtain a historical first omnidirectional information set;
[0130] Performing feature extraction on the historical first omnidirectional information set to obtain historical omnidirectional output information, where the historical omnidirectional output information corresponds to an event of the event camera;
[0131] Segmenting the historical omnidirectional output information according to time to obtain a plurality of first feature information sets corresponding to different time periods, each first feature information set representing first feature information at a different time period, the historical omnidirectional output information being obtained by feature extraction of historical second information, the historical second information corresponding to the historical first information; performing omnidirectional mapping processing on all first feature information included in one of the first feature information sets corresponding to time t to obtain an omnidirectional image, the omnidirectional image being the omnidirectional image corresponding to time t;
[0132] Minimize the loss function calculation between the omnidirectional image and the reference omnidirectional image to obtain a second loss value;
[0133] When the second loss value is within a second threshold range, outputting the omnidirectional image as the first output result;
[0134] The scene segmentation model is trained according to the historical output information, the first feature information set, the omnidirectional image, and the reference omnidirectional image to obtain the scene segmentation model.
[0135] In one possible implementation, the processing module 640 is further configured to:
[0136] Acquire historical first information and historical second information corresponding to the historical first information;
[0137] Performing feature extraction of location information on the historical second information to obtain historical output information, where the historical output information corresponds to an event of the event camera;
[0138] Segmenting the historical output information according to the location information to obtain a plurality of second feature information sets, each of the second feature information sets representing the historical output information of different location information;
[0139] Extracting the historical output information of each of the second feature information sets, obtaining a predicted position value according to a position prediction algorithm, and outputting the predicted position value as a second output result;
[0140] The position prediction model is trained according to the multiple historical output information, the second feature information set, and the position prediction algorithm to obtain the position prediction model.
[0141] The position prediction algorithm may be a particle filter method, a least squares method, etc., which is not limited here.
[0142] Among them, the scene segmentation model and the position prediction model can be trained separately or simultaneously. The scene segmentation module and the position prediction module can be assembled according to actual needs to achieve the effect of accelerating calculations.
[0143] In one possible implementation, the decision module 650 performs supervised learning based on the decision model, the processing results, and the driving information encoding to obtain the driving decision information, specifically for:
[0144] Mapping the processing result and the driving information code to a decision pulse code according to the decision model, and obtaining a first decision pulse code among the matched decision pulse codes;
[0145] The first decision pulse code is output as the driving decision information.
[0146] It can be seen that the event camera-based intelligent recognition device described in the embodiment of the present application first collects first information through the event camera, performs noise reduction training on the first information based on a pulse supervised learning network to obtain second information, performs feature extraction on the second information to obtain output information, the output information corresponds to the event of the event camera, processes the output information according to the scene segmentation model to obtain a first output result, processes the output information according to the position prediction model to obtain a second output result, merges the first output result and the second output result into a processing result, and finally performs supervised learning based on the processing result, driving information, and a decision model to obtain driving decision information. In this way, by using an event camera as a perception unit, a high sampling rate and high dynamic range solution is provided for intelligent recognition, saving the deployment cost of a high sampling rate perception unit. At the same time, the back-end algorithm builds multiple models based on a pulse neural network, which can further save energy consumption and further reduce computing resource requirements.
[0147] See also Figure 7 , Figure 7 This is a structural diagram of an electronic device proposed in an embodiment of the present application. As shown in the figure, the electronic device 700 includes a processor 710, a memory 720, a communication interface 730 and one or more programs 721. The one or more programs 721 are stored in the memory 720 and are configured to be executed by the processor 710.
[0148] The processor 710, the memory 720, and the communication interface 730 are interconnected and perform communication between them.
[0149] The memory 720 can be a volatile memory such as a dynamic random access memory (DRAM) or a non-volatile memory such as a mechanical hard disk. The memory 720 is used to store a set of executable program codes, and the processor 710 is used to call one or more programs 721 stored in the memory 720 to execute some or all of the steps of any of the water-related imaging methods described in the above-mentioned water-related imaging method embodiments.
[0150] Among them, the electronic device 700 may include smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, PDAs, driving recorders, vehicle-mounted electronic devices, servers, laptops, mobile Internet electronic devices (MID, Mobile Internet Devices) or wearable electronic devices (such as smart watches, Bluetooth headsets), etc. The above are only examples and not exhaustive, including but not limited to the above electronic devices.
[0151] It can be seen that the electronic device described in the embodiment of the present application can actively turn on the underwater detection function when the vehicle is wading through water to obtain underwater images, and the detection accuracy is high, which can effectively increase the probability of avoiding underwater obstacles and improve the user experience.
[0152] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments, and the above computer includes an electronic device.
[0153] The present application also provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may comprise an electronic device.
[0154] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0155] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0157] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0158] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0159] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions for enabling a computer electronic device (which can be a personal computer, electronic device or network electronic device, etc.) to perform all or part of the steps of the above-mentioned methods of each embodiment of the present application. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0160] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0161] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An intelligent recognition method based on event camera, characterized in that: Applied to a vehicle, the vehicle including an event camera, the method comprising: collecting first information by the event camera, wherein the first information includes positive event information and negative event information; Performing noise reduction training on the first information based on a pulse supervised learning network to obtain second information, and performing feature extraction on the second information to obtain output information; wherein, performing noise reduction training on the first information based on the pulse supervised learning network to obtain the second information includes: performing noise reduction on the positive event information and the negative event information of the first information through two isomorphic pulse neural network denoisers to obtain noise-reduced positive event information and noise-reduced negative event information respectively, merging the noise-reduced positive event information and the noise-reduced negative event information to obtain noise-reduced first information, and performing loss minimization calculation on the noise-reduced first information and reference information to obtain a first loss value, wherein the reference information is a reference value obtained by noise reduction processing of information acquired by a second sensor, and the second sensor is a reference sensor; if the first loss value is within a first threshold range, outputting the noise-reduced first information as the second information; performing feature extraction on the second information includes: converting the second information into a first pulse signal, performing feature extraction based on the first pulse signal to obtain output information, wherein the feature extraction is pulse autoencoding, and the output information corresponds to an event of the event camera; Processing the output information according to the scene segmentation model to obtain a first output result, processing the output information according to the position prediction model to obtain a second output result, and combining the first output result and the second output result to output as a processing result; Supervised learning is performed based on the processing results, driving information, and a decision model to obtain driving decision information.
2. The method according to claim 1, characterized in that The event camera is a plurality of event cameras, wherein each event camera corresponds to one piece of the first information, and the method further includes: Performing omnidirectional construction processing on the second information corresponding to the plurality of first information to obtain a first omnidirectional information set; The extracting features of the second information to obtain output information includes: converting the plurality of second information included in the first omnidirectional information set into a second pulse signal; Feature extraction is performed based on the second pulse signal to obtain output information.
3. The method according to claim 2, characterized in that The second information corresponding to the plurality of first information is omnidirectionally constructed to obtain a first omnidirectional information set, which satisfies the following formula: Among them, event represents the first omnidirectional information set, N represents the number of the event cameras, A represents the maximum detection angle of each first sensor, and offset represents the angle between the probe of the first sensor and the front direction of the vehicle.
4. The method according to claim 3, characterized in that The method further comprises: Acquire historical first information and historical second information corresponding to the historical first information; Performing feature extraction on the historical second information to obtain historical output information, where the historical output information corresponds to the event of the event camera; Segmenting the historical output information by time to obtain a first feature information set corresponding to a plurality of different time periods, wherein the historical output information is obtained by extracting features from the historical second information, and the historical second information corresponds to the historical first information; Mapping the historical output information included in the first feature information set corresponding to time t to obtain a first image, where the first image is the image corresponding to time t; Minimizing a loss function calculation on the image and a reference image to obtain a second loss value, where the reference image is an image obtained by other cameras, and the other cameras are cameras other than the event camera; When the second loss value is within a second threshold range, outputting the first image as the first output result; The scene segmentation model is trained according to the historical output information, the first feature information set, the first image, and the reference image to obtain the scene segmentation model.
5. The method according to claim 4, characterized in that In a case where the historical first information is a plurality of historical first information, the method further includes: Acquire the plurality of historical first information and the plurality of historical second information corresponding to the plurality of historical first information; Performing omnidirectional construction processing on the plurality of historical second information corresponding to the plurality of historical first information to obtain a historical first omnidirectional information set; Performing feature extraction on the historical first omnidirectional information set to obtain historical omnidirectional output information, where the historical omnidirectional output information corresponds to an event of the event camera; Segmenting the historical omnidirectional output information according to time to obtain a plurality of first feature information sets corresponding to different time periods, each first feature information set representing first feature information at a different time period, the historical omnidirectional output information being obtained by feature extraction of historical second information, the historical second information corresponding to the historical first information; performing omnidirectional mapping processing on all first feature information included in one of the first feature information sets corresponding to time t to obtain an omnidirectional image, the omnidirectional image being the omnidirectional image corresponding to time t; Minimize the loss function calculation between the omnidirectional image and the reference omnidirectional image to obtain a second loss value; When the second loss value is within a second threshold range, outputting the omnidirectional image as the first output result; The scene segmentation model is trained according to the historical output information, the first feature information set, the omnidirectional image, and the reference omnidirectional image to obtain the scene segmentation model.
6. The method according to claim 1, characterized in that The method further comprises: Acquire historical first information and historical second information corresponding to the historical first information; Performing feature extraction of location information on the historical second information to obtain historical output information, where the historical output information corresponds to an event of the event camera; Segmenting the historical output information according to the location information to obtain a plurality of second feature information sets, each of the second feature information sets representing the historical output information of different location information; Extracting the historical output information of each of the second feature information sets, obtaining a predicted position value according to a position prediction algorithm, and outputting the predicted position value as a second output result; The position prediction model is trained according to the multiple historical output information, the second feature information set, and the position prediction algorithm to obtain the position prediction model.
7. The method according to claim 1, characterized in that The supervision model is a preset model, and supervised learning is performed based on the decision model, the processing results, and the driving information encoding to obtain driving decision information including: Mapping the processing result and the driving information code to a decision pulse code according to the decision model, and obtaining a first decision pulse code among the matched decision pulse codes; The first decision pulse code is output as the driving decision information.
8. A computer-readable storage medium, characterized in that A vehicle obstacle avoidance program based on an event camera is stored, including execution instructions. When a processor of an electronic device executes the execution instructions, the processor executes the method according to any one of claims 1 to 7.
9. An electronic device, characterized in that: The method comprises a processor and a memory storing execution instructions, wherein the memory stores one or more programs; when the processor executes the execution instructions stored in the memory, the processor executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for training driving behavior decision model
CN113835421A
Object classification based on frame and event camera processing and related applications
CN115668309A