Weather scene classification method, electronic device, vehicle, medium and computer program product
By using dynamic visual sensor DVS data and convolutional neural network pulse neuron model, the dynamic information of weather scenes is captured, the real-time problem of weather scene classification is solved, and classification decisions that quickly respond to severe weather are achieved.
Patent Information
- Application Number
- CN202510787032.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-03
AI Technical Summary
The real-time performance of weather scene classification in existing technologies is not high, and it is difficult to effectively process dynamically changing and fuzzy visual information.
The dynamic visual sensor (DVS) data is used to build a classification model in combination with convolutional neural networks and spiking neurons. By acquiring the event stream of DVS data, the dynamic information of time changes is captured and weather scene classification is performed.
It improves the real-time performance of weather scene classification, can quickly respond to changing weather scenes, and ensure the safe operation of systems such as intelligent driving and drones in severe weather.
Smart Images

Figure CN120747719A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer vision technology, and specifically relates to a weather scene classification method, electronic device, vehicle, medium and computer program product. Background Art
[0002] In the fields of intelligent driving, assisted driving, drones and navigation robots, weather scene classification has become an important research direction to ensure driving safety, flight safety and operational safety.
[0003] In related technologies, images of scenes are acquired through sensors such as cameras and lidars, traditional image processing methods such as edge detection and color segmentation are used to process visual data in severe weather scenes, and complex image enhancement techniques are used to improve perception effects.
[0004] However, the real-time performance of weather scene classification using related technical methods is not high. Summary of the Invention
[0005] Embodiments of the present application provide a weather scene classification method, electronic device, vehicle, medium, and computer program product to solve the problem of low real-time performance of weather scene classification in related technologies.
[0006] A first aspect of an embodiment of the present application provides a weather scene classification method, the method comprising: Acquire dynamic visual sensor (DVS) data, wherein the DVS data includes an event stream, wherein each event in the event stream includes time information, position information, and brightness information; The DVS data is input into a weather scene classification model to obtain a weather scene target classification result, wherein the weather scene classification model is trained based on a DVS sample data set.
[0007] Optionally, the weather scene classification model includes: at least two cascaded pulse modules, a flattening module, a classification module and a voting module; Inputting the DVS data into a weather scene classification model to obtain the weather scene target classification result includes: performing pulse processing on the DVS data by the at least two cascaded pulse modules to obtain target pulse characteristics; Flattening the spatial dimension of the target pulse feature by a flattening module to obtain a flattened target pulse feature; Performing feature mapping on the target pulse features through the classification module to obtain an initial classification result of the weather scene; The voting module performs a voting mechanism on the initial classification results of the weather scenes to obtain the target classification results of the weather scenes.
[0008] Optionally, the pulse processing of the DVS data by the at least two cascaded pulse modules to obtain a target pulse feature includes: For each pulse module, the previous level pulse features are input into the current level pulse module, spatial features and temporal features are extracted from the previous level pulse features, and the current level pulse features are output. The input of the first level pulse module is the DVS data, and the output of the last level pulse module is the target pulse feature.
[0009] Optionally, the pulse module includes: a convolution module and a pulse neuron module; the pulse neuron module includes: a pulse neuron unit and a maximum pooling unit; The performing spatial feature extraction and temporal feature extraction on the pulse feature of the previous stage and outputting the pulse feature of the current stage includes: The spatial feature extraction is performed on the pulse feature of the previous level through the convolution module to obtain the spatial feature; the temporal feature extraction is performed on the spatial feature through the pulse neuron module to output the pulse feature of the current level.
[0010] Optionally, the convolution module includes: a convolution unit and a batch normalization unit; The step of extracting spatial features from the previous-stage pulse features through a convolution module to obtain spatial features includes: Performing a convolution operation on the previous level pulse feature through the convolution unit to obtain a local spatial feature; The local spatial features are normalized by the batch normalization unit to obtain the spatial features.
[0011] The step of extracting the temporal features of the spatial features by the pulse neuron module and outputting the pulse features of the current level includes: The spiking neural unit emits pulses to the spatial feature to obtain a pulse sequence; The maximum pooling unit performs a maximum pooling operation on the pulse sequence to obtain the pulse features of the current level.
[0012] Optionally, the DVS sample data set includes: DVS sample data and annotation information, where the annotation information is used to identify weather scene categories. The training process of the weather scene classification model includes: The DVS sample data set is input into the weather scene classification model, forward propagation is performed on each batch of DVS sample data, and the weather scene classification model is updated based on the output result of the weather scene classification model, the labeling information of the DVS sample data and the mean square error loss function until the weather scene classification model converges.
[0013] A second aspect of an embodiment of the present application provides a weather scene classification device, the device comprising: An acquisition module is used to acquire dynamic visual sensor DVS data, wherein the DVS data includes: an event stream, wherein each event in the event stream includes: time information, position information and brightness information; The processing module is used to input the DVS data into a weather scene classification model to obtain a weather scene target classification result. The weather scene classification model is trained based on the DVS sample data set.
[0014] Optionally, the weather scene classification model includes: at least two cascaded pulse modules, a flattening module, a classification module and a voting module; The processing module is specifically used to perform pulse processing on the DVS data through the at least two cascaded pulse modules to obtain target pulse features; flatten the spatial dimensions of the target pulse features through the flattening module to obtain flattened target pulse features; perform feature mapping on the target pulse features through the classification module to obtain an initial classification result of the weather scene; and execute a voting mechanism on the initial classification result of the weather scene through the voting module to obtain a target classification result of the weather scene.
[0015] Optionally, the processing module is specifically used to input the previous level pulse features into the current level pulse module for each pulse module, perform spatial feature extraction and temporal feature extraction on the previous level pulse features, and output the current level pulse features, wherein the input of the first level pulse module is the DVS data, and the output of the last level pulse module is the target pulse feature.
[0016] Optionally, the pulse module includes: a convolution module and a pulse neuron module; The processing module is specifically used to extract the spatial features of the previous level pulse features through the convolution module to obtain spatial features; extract the temporal features of the spatial features through the pulse neuron module to output the pulse features of this level.
[0017] Optionally, the convolution module includes: a convolution unit and a batch normalization unit; the pulse neuron module includes: a pulse neuron unit and a maximum pooling unit; The processing module is specifically used to perform a convolution operation on the previous level pulse feature through the convolution unit to obtain a local spatial feature; perform normalization processing on the local spatial feature through the batch normalization unit to obtain the spatial feature; perform pulse emission on the spatial feature through the pulse neural unit to obtain a pulse sequence; perform a maximum pooling operation on the pulse sequence through the maximum pooling unit to obtain the pulse feature of the current level.
[0018] Optionally, the DVS sample data set includes: DVS sample data and annotation information, where the annotation information is used to identify weather scene categories. The training process of the weather scene classification model includes: The DVS sample data set is input into the weather scene classification model, forward propagation is performed on each batch of DVS sample data, and the weather scene classification model is updated based on the output result of the weather scene classification model, the labeling information of the DVS sample data and the mean square error loss function until the weather scene classification model converges.
[0019] A third aspect of an embodiment of the present application provides an electronic device, comprising: a processor, the processor being used to connect to a memory, the memory storing a computer program / instruction that can be run on the processor, and the computer program / instruction, when executed by the processor, implements the steps of the weather scene classification method as described in any one of the first aspects.
[0020] A fourth aspect of an embodiment of the present application provides a vehicle, comprising the electronic device as described in the third aspect.
[0021] A fifth aspect of an embodiment of the present application provides a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the weather scene classification method as described in any one of the first aspects are implemented.
[0022] A sixth aspect of an embodiment of the present application provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the weather scene classification method as described in any one of the first aspects.
[0023] The weather scene classification method, electronic device, vehicle, medium and computer program product provided in the embodiments of the present application obtain dynamic visual sensor DVS data, wherein the DVS data includes: an event stream, wherein each event in the event stream includes: time information, location information and brightness information; the DVS data is input into a weather scene classification model to obtain a weather scene target classification result, and the weather scene classification model is trained based on a DVS sample data set; that is, by obtaining the event stream of DVS data, dynamic information of weather scenes changing over time can be captured, and then the time series information of the DVS data is processed by the trained weather scene classification model, which can quickly respond to changing weather scenes and obtain weather scene classification results, thereby improving the real-time performance of weather scene classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A schematic diagram of the structure of a weather scene classification system provided in an embodiment of the present application; Figure 2 A schematic diagram of a flow chart of a weather scene classification method provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of another weather scene classification system provided in an embodiment of the present application; Figure 4 A flow chart of another weather scene classification method provided in an embodiment of the present application; Figure 5 A flow chart of another weather scene classification method provided in an embodiment of the present application; Figure 6 A schematic diagram of the structure of another weather scene classification system provided in an embodiment of the present application; Figure 7 A flow chart of another weather scene classification method provided in an embodiment of the present application; Figure 8 A schematic diagram of the structure of another weather scene classification system provided in an embodiment of the present application; Figure 9 A flow chart of another weather scene classification method provided in an embodiment of the present application; Figure 10 A flow chart of another weather scene classification method provided in an embodiment of the present application; Figure 11 A schematic structural diagram of a weather scene classification device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0026] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0027] The term "indication" in this application can be either a direct indication (or explicit indication) or an indirect indication (or implicit indication). A direct indication can be understood as the sender explicitly informing the receiver of specific information, the required operation, or the requested result, etc. in the instruction sent; an indirect indication can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the required operation or request result based on the judgment result.
[0028] In the related art, weather scene images are obtained through sensors such as cameras and lidars, and scene classification under complex weather conditions through methods such as edge detection has poor adaptability and cannot process dynamically changing and blurred visual information, resulting in low real-time performance of weather scene classification. To address the above problems, the present application provides a weather scene classification method, which can capture dynamic information of time changes by collecting dynamic vision sensor (DVS) event streams in high-speed mobile environments, and classify DVS event streams through a classification model constructed based on convolutional neural networks and pulse neurons, thereby improving the real-time performance of weather scene classification.
[0029] The dynamic vision sensor involved in the embodiments of the present application does not capture images at a fixed frame rate, but responds to the brightness changes of each pixel in the image frame. When a pixel detects a brightness change exceeding a set threshold, the pixel will generate an "event". These events contain information about the increase or decrease in brightness and the timestamp of the event. The dynamic vision sensor can capture changes in the scene with extremely high temporal resolution. For example, in unmanned driving, DVS can detect weather changes such as rain, snow and fog, which facilitates vehicles and aircraft to make timely decisions and ensure safe travel in bad weather.
[0030] The following describes the technical solution of the weather scene classification method of the application using several specific embodiments as examples: Figure 1 A structural diagram of a weather scene classification system provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the system includes: a weather scene classification model, which is used to obtain weather scene target classification results based on the DVS data.
[0031] Figure 2 A flow chart of a weather scene classification method provided in an embodiment of the present application is provided. The method of this embodiment is applied to Figure 1 The weather scene classification system shown in Figure 2 As shown, the process of the method of this embodiment is as follows: S21: Acquire dynamic visual sensor DVS data, the DVS data including: an event stream, wherein each event in the event stream includes: time information, position information, and brightness information.
[0032] The DVS data can be acquired by dynamic visual sensors such as DVS cameras, including but not limited to dynamic weather scenes such as rainy days, snowy days and foggy days. The rainy day event stream reflects the changes in the falling raindrops, the snowy day event stream reflects the changes in the fluttering snowflakes, and the foggy day event stream reflects the changes in light intensity. The identification of severe weather helps intelligent driving vehicles make decisions.
[0033] Furthermore, the event stream of the DVS data consists of a series of discrete "events" of weather scenes. The corresponding image frames when the "events" occur are used to provide background references to visualize the scenes, including but not limited to: highways, urban roads and country roads passed by smart driving vehicles; the "event" records the brightness change information of a certain pixel position in the image frame at a specific moment, such as the pixel brightness change caused by a raindrop falling on a highway. The "event" specifically includes: time information, position information and brightness information; the time information refers to the timestamp, which indicates the specific moment when the event occurred, in microseconds (μs); the position information refers to the pixel coordinates of the event in the image frame, corresponding to the position of the specific pixel point where the brightness change occurred, expressed in two-dimensional coordinates; the brightness information refers to the direction of change in pixel brightness, which is polarity information, for example, "+1" indicates an increase in brightness and "-1" indicates a decrease in brightness.
[0034] S22: Inputting the DVS data into a weather scene classification model to obtain a weather scene target classification result, wherein the weather scene classification model is trained based on the DVS sample data set.
[0035] The weather scene classification model is based on a convolutional neural network architecture, using spiking neurons instead of ordinary artificial neurons to simulate the spiking mechanism of biological neurons. This model can extract the spatial features of the DVS data and further encode information in the temporal dimension, thereby capturing key changes in the dynamic visual information of the weather scene. Classification decisions for the weather scene are then made based on the extracted features. Inputting the DVS data into the weather scene classification model yields the corresponding target classification results for the weather scene. For example, for an intelligent driving vehicle traveling on a bridge between 8 and 9 a.m., inputting the dynamic weather event stream captured by the DVS camera into the weather scene classification model yields a target classification result of "snow."
[0036] In this embodiment, dynamic visual sensor (DVS) data is acquired, the DVS data including an event stream, wherein each event in the event stream includes time information, location information, and brightness information; the DVS data is input into a weather scene classification model to obtain a weather scene target classification result, wherein the weather scene classification model is trained based on a DVS sample data set; that is, by acquiring the event stream of DVS data, dynamic information of weather scenes changing over time can be captured, and then the time series information of the DVS data is processed by the trained weather scene classification model, so that a rapid response to changing weather scenes can be obtained, and a weather scene classification result can be obtained, thereby improving the real-time performance of weather scene classification.
[0037] Figure 3 This is a structural diagram of another weather scene classification system provided in an embodiment of the present application. Figure 3 exist Figure 1 On the basis of the system shown, further, the weather scene classification model includes: at least two cascaded pulse modules, a flattening module, a classification module and a voting module, the at least two cascaded pulse modules are used to perform pulse processing on the DVS data to obtain target pulse features; the flattening module is used to flatten the spatial dimensions of the target pulse features through the flattening module to obtain flattened target pulse features; the classification module is used to perform feature mapping on the target pulse features to obtain the initial classification result of the weather scene; the voting module is used to execute a voting mechanism on the initial classification result of the weather scene to obtain the target classification result of the weather scene.
[0038] Figure 4 A flow chart of another weather scene classification method provided in an embodiment of the present application is provided. Figure 4 is Figure 2 Based on the description of a possible implementation of S22, the method of this embodiment is applied to Figure 3 The weather scene classification system shown in Figure 4 As shown, the process of the method of this embodiment is as follows: S221: Performing pulse processing on the DVS data through the at least two cascaded pulse modules to obtain target pulse characteristics.
[0039] Among them, a possible weather scene classification model structure includes 5 cascaded pulse modules, each of which outputs a different number of target pulse feature channels. By pulse-processing the DVS data through the 5 cascaded pulse modules, features with time series characteristics, namely the target pulse features, can be obtained from the image frames corresponding to the event stream of the DVS data. For example, assuming that the input DVS data is a DVS event stream of a rainy scene in a certain time period, the shape is , through 5 cascaded pulse modules, we can get the shape of The target pulse characteristics are represents the time step, Indicates height, Indicates width, Indicates the number of channels.
[0040] S222: Flatten the spatial dimension of the target pulse feature using a flattening module to obtain a flattened target pulse feature.
[0041] The flattening module is a flatten operation, which flattens the multidimensional feature map of the spatial dimension into a one-dimensional vector at each time step t, which can be easily input into the fully connected layer or other classifiers. For example, the target pulse feature shape is , after flattening, the shape is , which can be input into classification modules such as the fully connected layer for classification.
[0042] S223: Perform feature mapping on the target pulse features through the classification module to obtain an initial classification result of the weather scene.
[0043] Among them, one possible structure of the classification module is a fully connected (FC) layer, that is, a linear layer, in which each neuron is connected to all neurons in the previous layer. The fully connected layer can map the target pulse feature vector to the category space through the weight matrix and activation function, and output the category probability distribution of the weather scene at each time step, that is, the initial classification result of the weather scene. The number of neurons in the output layer is the number of categories, for example, it can be 3, 5, or 10, etc. The output of the fully connected layer can be expressed as:
[0044] in, is the weight matrix, is the bias term, is the activation function. For example, the target pulse feature obtained from the DVS data of the weather scene is input into the fully connected layer, and the time step is , the number of neurons in the output layer of the fully connected layer is , we can get The initial classification results of weather scenes in time steps are output at each time step. weather scene categories and the corresponding probability scores for each category.
[0045] S224: Executing a voting mechanism on the initial classification result of the weather scene through the voting module to obtain the target classification result of the weather scene.
[0046] The voting mechanism refers to averaging the initial classification results of the weather scene in the time dimension and selecting the category with the highest probability score as the final classification result. For example, based on the above example, The probability scores of each weather scene category at the time steps are averaged. Assuming that the probability score of rainy day is the highest, the weather scene target classification result is rainy day.
[0047] In this embodiment, the DVS data is pulsed by the at least two cascaded pulse modules to obtain target pulse features; the target pulse features are feature mapped by the classification module to obtain an initial classification result of the weather scene; and the voting module performs a voting mechanism on the initial classification result of the weather scene to obtain a target classification result of the weather scene, thereby realizing end-to-end weather scene classification based on DVS data.
[0048] Figure 5 A flow chart of another weather scene classification method provided in an embodiment of the present application is shown below. Figure 5 is Figure 4 Based on the description of a possible implementation of S221, the method of this embodiment is applied to Figure 3 The weather scene classification system shown in Figure 5 As shown, the process of the method of this embodiment is as follows: S2211: For each pulse module, input the pulse features of the previous level into the pulse module of the current level, perform spatial feature extraction and temporal feature extraction on the pulse features of the previous level, and output the pulse features of the current level.
[0049] The input of the first-stage pulse module is the DVS data, and the output of the last-stage pulse module is the target pulse feature.
[0050] Specifically, for example, for 5 cascaded pulse modules, the spatial feature extraction and time feature extraction are performed on the DVS data through the first-level pulse module, and the first-level pulse feature is output; the spatial feature extraction and time feature extraction are performed on the first-level pulse feature through the second-level pulse module, and the second-level pulse feature is output; and so on, the spatial feature extraction and time feature extraction are performed on the fourth-level pulse feature through the fifth-level pulse module, and the target pulse feature is output.
[0051] In this embodiment, for each pulse module, the pulse features of the previous level are input into the pulse module of the current level, spatial features and temporal features are extracted from the pulse features of the previous level, and the pulse features of the current level are output. The input of the first-level pulse module is the DVS data, and the output of the last-level pulse module is the target pulse feature. Thus, the deep spatial and temporal features of the DVS data are obtained, which helps to improve the accuracy of weather scene classification.
[0052] Figure 6 This is a structural diagram of another weather scene classification system provided in an embodiment of the present application. Figure 6 exist Figure 3 On the basis of this, further, the pulse module includes: a convolution module and a pulse neuron module; the convolution module is used to extract the spatial features of the previous level pulse features to obtain spatial features; the pulse neuron module is used to extract the temporal features of the spatial features and output the pulse features of this level.
[0053] Figure 7 A flow chart of another weather scene classification method provided in an embodiment of the present application is provided. Figure 7 is Figure 5 Based on the description of a possible implementation of S2211, the method of this embodiment is applied to Figure 6 The weather scene classification system shown in Figure 7 As shown, the process of the method of this embodiment is as follows: S7001: Perform spatial feature extraction on the previous level pulse feature through a convolution module to obtain spatial features.
[0054] Among them, one possible structure of the convolution module is a convolution layer and a batch normalization (BN) layer. The convolution module can extract and normalize the local spatial features of the previous level pulse features to obtain the spatial features of the current level. For example, for five cascaded pulse modules, the convolution module of the second level pulse module extracts and normalizes the local spatial features of the first level pulse pulse features to obtain the second level spatial features. The shape of the spatial features output by the convolution module of each level of pulse module is different, which is specifically determined based on the convolution kernel size, number of convolution kernels, stride size, and padding size of the convolution layer.
[0055] S7002: Extracting temporal features from the spatial features through the pulse neuron module, and outputting the pulse features of this level.
[0056] Among them, one possible structure of the spiking neuron module is a spiking neuron and maximum pooling layer. The spiking neuron module can be used to perform spiking and maximum pooling processing on the spatial features output by the convolution module of the current-level spiking module, extracting the temporal coding information and retaining the most important feature information to obtain the spiking features of the current level. For example, for five cascaded spiking modules, the spiking neuron module of the second-level spiking module performs spiking and maximum pooling processing on the second-level spatial features to obtain the second-level spiking features.
[0057] In this embodiment, the spatial features of the previous level pulse features are extracted through the convolution module to obtain spatial features; the temporal features of the spatial features are extracted through the pulse neuron module to output the pulse features of this level, thereby effectively extracting the spatiotemporal features of the input data, facilitating the effective analysis of time series data, and improving the accuracy of classification.
[0058] Figure 8 This is a structural diagram of another weather scene classification system provided in an embodiment of the present application. Figure 8 exist Figure 6 On the basis of this, further, the convolution module includes: a convolution unit and a batch normalization unit; the pulse neuron module includes: a pulse neuron unit and a maximum pooling unit; the convolution unit is used to perform a convolution operation on the previous level pulse feature to obtain a local spatial feature; the batch normalization unit is used to perform normalization on the local spatial feature to obtain the spatial feature; the pulse neuron is used to emit pulses on the spatial feature to obtain a pulse sequence; the maximum pooling unit is used to perform a maximum pooling operation on the pulse sequence to obtain the pulse feature of this level.
[0059] Figure 9 A flow chart of another weather scene classification method provided in an embodiment of the present application is provided. Figure 9 is Figure 7 Based on the description of a possible implementation of S7001, the method of this embodiment is applied to Figure 8 The weather scene classification system shown in Figure 9 As shown, the process of the method of this embodiment is as follows: S901: Perform a convolution operation on the previous level pulse feature through the convolution unit to obtain a local spatial feature.
[0060] The convolution unit is a convolution layer. Assuming there are M cascaded pulse modules, the M-1 level pulse feature size is , assuming that the convolution kernel size of the convolution layer is , the step size is , filled with , the output feature size is , that is, the local spatial feature, where The size is determined by the number of convolution kernels, and the spatial dimension of the output feature can be expressed as:
[0061]
[0062] S902: Normalizing the local spatial features by the batch normalization unit to obtain the spatial features.
[0063] The batch normalization unit is a BN layer, which normalizes the local spatial features of each channel at each time step so that its mean is close to 0 and its variance is close to 1, which can be expressed as:
[0064] in, represents the normalized eigenvalue, represents the local spatial eigenvalue of each channel at each time step, represents the mean, represents the variance, Represents the minimum value, which is a small constant added for numerical stability, such as 1e-5. Then multiply the normalized eigenvalue by a scaling factor And add an offset , to restore the scale and offset of the feature, and finally obtain the normalized and scaled feature value , that is, the spatial feature has the same shape as the input local spatial feature.
[0065] In this embodiment, the convolution unit performs a convolution operation on the previous-stage pulse feature to obtain a local spatial feature; the batch normalization unit performs normalization processing on the local spatial feature to obtain the spatial feature, thereby obtaining spatial information of the event occurrence in the DVS data, which helps to improve the efficiency of pulseization.
[0066] Figure 10 A flow chart of another weather scene classification method provided in an embodiment of the present application is provided. Figure 10 is Figure 7 Based on the description of a possible implementation of S7002, the method of this embodiment is applied to Figure 8 The weather scene classification system shown in Figure 10 As shown, the process of the method of this embodiment is as follows: S1001: The pulse neural unit emits pulses to the spatial feature to obtain a pulse sequence.
[0067] Among them, a possible structure of the spiking neural unit is the Leaky Integrate-and-Fire Node (LIF Node), which is a neuron model that mimics the behavior of biological neurons and fires a pulse when the input exceeds a certain threshold. Specifically, leakage refers to the fact that the membrane potential of the neuron decays back to the resting potential over time, integration refers to the neuron integrating the input current, and when the accumulation is sufficient to exceed the threshold potential, an action potential is generated, and firing refers to the neuron firing an action potential to reset the membrane potential in the dark. When the spatial features of the batch normalization unit output are passed to the LIF neuron model, the continuous activation values are converted into pulse signals according to the threshold trigger mechanism.
[0068] For example, suppose the spatial features of the batch normalization unit output are , the LIF neuron updates its membrane potential, which is expressed as:
[0069]
[0070]
[0071] in, It represents the membrane potential accumulated in the neuron by the input signal at the next moment. represents the membrane potential of the neuron, It represents the membrane time constant, which determines the rate at which the membrane potential changes over time and is a measure of how quickly a neuron responds to input signals. It represents the resting potential, which is the membrane potential of the neuron when there is no input signal; Represents the external input received by the neuron at the current moment, Indicates the pulse emission signal of the neuron at the next moment, represents the pulse emission threshold, represents a step function that determines whether a neuron sends a pulse if Exceeding the threshold , then it is 1, otherwise it is 0, that is, no pulse is sent. Represents the membrane potential after the neuron fires a pulse at the next moment; It represents the membrane potential adjustment coefficient after the release, which is used to calculate the membrane potential after the release of the pulse, and finally obtain the pulse sequence with the shape of .
[0072] S1002: Performing a maximum pooling operation on the pulse sequence through the maximum pooling unit to obtain the pulse features of this level.
[0073] Among them, the maximum pooling unit is a maximum pooling layer, which performs a maximum pooling operation on the spatial dimension at each time step t, specifically sets the pooling window size and stride, extracts the maximum value for each window, reduces the spatial dimension, retains the strongest pulse signal, and outputs the target pulse feature.
[0074] Specifically, assume that the input of the maximum pooling unit is of shape Pulse sequence, assuming the pooling window size is , the stride is , after maximum pooling, the output target pulse feature shape is ( ),in, ;
[0075] in, Indicates rounding down to an integer. represents the feature height before maximum pooling, represents the feature height after maximum pooling, represents the feature width before maximum pooling, Represents the feature width after maximum pooling.
[0076] In this embodiment, a pulse sequence is obtained by emitting pulses on the spatial features through the pulse neural unit; and the pulse features at this level are obtained by performing a maximum pooling operation on the pulse sequence through the maximum pooling unit, thereby capturing the temporal dynamic characteristics and reducing the dimension and noise of the features, thereby providing a stable input for subsequent classification.
[0077] In the above embodiment, optionally, the DVS sample data set includes: DVS sample data and annotation information, and the annotation information is used to identify the weather scene category. The training process of the weather scene classification model includes: inputting the DVS sample data set into the weather scene classification model, performing forward propagation on each batch of DVS sample data, and updating the weather scene classification model based on the output result of the weather scene classification model, the annotation information of the DVS sample data and the mean square error loss function until the weather scene classification model converges.
[0078] Specifically, the method of obtaining the DVS sample data includes but is not limited to dynamic visual sensors such as DVS cameras and simulation software such as autonomous driving simulators. The labeling information is manually labeled. For example, the weather scene categories are divided into 8, sunny days are labeled as 0, rainy days during the day are labeled as 1, rainy days at night are labeled as 2, snowy days during the day are labeled as 3, snowy days at night are labeled as 4, foggy days during the day are labeled as 5, foggy days at night are labeled as 6, and dense fog is labeled as 7.
[0079] Furthermore, the weather scene classification model is trained on a graphics processing unit (GPU). Specifically, torch.device('cuda' if torch.cuda.is_available() else 'cpu') automatically selects the appropriate CUDA device or CPU device for training. The labeled DVS sample dataset is input into the weather scene classification model. The forward propagation involves feeding a batch of DVS sample data into the model, starting from the input layer, passing through each hidden layer, and finally reaching the output layer. At each layer, the data is calculated based on the weights and biases of that layer, and the output of that layer is used as the input for the next layer until the output layer results are obtained. During training, each batch of DVS data is transposed to meet the model's requirements for processing time series data. The loss function of the model adopts the mean squared error (MSE) loss function. The output of the model and the annotation information of the corresponding DVS sample data are substituted into the MSE function to calculate the loss value. Then, the gradient of the MSE loss function is passed back to the model through the back propagation algorithm. Optimization algorithms such as gradient descent and adaptive moment estimation (Adam) optimizer can be used to update the model parameters according to the calculated gradient. Taking gradient descent as an example, the learning rate is set to 0.001. For the parameters , the update formula is:
[0080] in, represents the updated parameters, represents the learning rate, Represents the loss function gradient.
[0081] Taking the Adam optimizer as an example, initialize the first-order moment estimate and second-order moment estimates is a zero vector, and the learning rate Set to 0.001, the first-order moment decay rate Set to 0.9, the second-order moment decay rate Set to 0.999, for the parameter , the adaptive gradient update formula is:
[0082]
[0083]
[0084]
[0085]
[0086] in, represents the updated first-order moment estimate, represents the updated second-order moment estimate, represents the gradient, represents the corrected first-order moment estimate, represents the corrected second-order moment estimate, Represents a constant that prevents division by zero, such as 1e-8.
[0087] By continuously iterating the above optimization process, the parameters of the model are gradually adjusted to the direction of minimizing the MSE loss function, so as to achieve the convergence conditions of the weather scene classification model, make the output value of the model closer to the true value, and thus improve the performance of the model.
[0088] Optionally, in order to improve the stability of the weather scene classification model training, a dropout layer is also used to reduce overfitting. Specifically, during each forward propagation, the dropout layer sets the elements in the input tensor to 0 with a probability p. The neurons set to zero will not participate in the forward and backward propagation in the current batch.
[0089] Optionally, after the training is completed, the embodiments of the present application also provide multiple evaluation methods, including: overall testing, same-domain testing and heterogeneous testing, to test the generalization ability of the weather scene classification model; among them, the overall test is to evaluate the performance of the weather scene classification model on test data of all weather scenes, the same-domain test is to evaluate the performance of the weather scene classification model on test data with the same distribution as the training data set, and the heterogeneous test is to evaluate the performance of the weather scene classification model on test data with a different distribution from the training data set.
[0090] The weather scene classification model performs inference on each batch of test data. Specifically, each input data (frame) is transferred to the appropriate GPU or central processing unit (CPU). The test data is loaded and preprocessed using the get_DVSDataloader function. The input data dimensions are adjusted to meet the model's input requirements using the .transpose(0, 1) function. Forward propagation is performed using the net(frame) function to obtain the model's output. The inference time is calculated using the time.time() function, recording and accumulating the time spent on each inference. After each inference, the model's output (out) is compared with the true label (label) to calculate the number of correct predictions. A correct prediction is considered to be consistent with the label if out > 0.5. All correct predictions are accumulated to calculate the overall test accuracy.
[0091] Furthermore, during the testing phase, the model's performance is evaluated by calculating metrics such as accuracy, inference time, and inference speed. During inference, the model does not calculate gradients, reducing memory usage and computational overhead, thereby increasing inference speed and ensuring a fast model response.
[0092] Figure 11 A schematic diagram of the structure of a weather scene classification device provided in an embodiment of the present application is shown as follows: Figure 11 As shown, the device includes: an acquisition module 1101 and a processing module 1102, wherein: An acquisition module 1101 is used to acquire dynamic visual sensor (DVS) data, where the DVS data includes an event stream, wherein each event in the event stream includes time information, location information, and brightness information. A processing module 1102 is used to input the DVS data into a weather scene classification model to obtain a weather scene target classification result, where the weather scene classification model is trained based on a DVS sample data set.
[0093] Optionally, the weather scene classification model includes: at least two cascaded pulse modules, a flattening module, a classification module and a voting module; the processing module 1102 is specifically used to perform pulse processing on the DVS data through the at least two cascaded pulse modules to obtain target pulse features; flatten the spatial dimensions of the target pulse features through the flattening module to obtain flattened target pulse features; perform feature mapping on the target pulse features through the classification module to obtain an initial classification result of the weather scene; and execute a voting mechanism on the initial classification result of the weather scene through the voting module to obtain the target classification result of the weather scene.
[0094] Optionally, the processing module 1102 is specifically used to input the previous level pulse features into the current level pulse module for each pulse module, perform spatial feature extraction and temporal feature extraction on the previous level pulse features, and output the current level pulse features, wherein the input of the first level pulse module is the DVS data, and the output of the last level pulse module is the target pulse feature.
[0095] Optionally, the pulse module includes: a convolution module and a pulse neuron module; the processing module 1102 is specifically used to extract the spatial features of the previous level pulse features through the convolution module to obtain spatial features; and extract the temporal features of the spatial features through the pulse neuron module to output the pulse features of this level.
[0096] Optionally, the convolution module includes: a convolution unit and a batch normalization unit; the pulse neuron module includes: a pulse neuron unit and a maximum pooling unit; the processing module 1102 is specifically used to perform a convolution operation on the previous level pulse feature through the convolution unit to obtain a local spatial feature; perform normalization processing on the local spatial feature through the batch normalization unit to obtain the spatial feature; perform pulse emission on the spatial feature through the pulse neuron unit to obtain a pulse sequence; perform a maximum pooling operation on the pulse sequence through the maximum pooling unit to obtain the pulse feature of the current level.
[0097] Optionally, the DVS sample data set includes: DVS sample data and annotation information, where the annotation information is used to identify weather scene categories. The training process of the weather scene classification model includes: The DVS sample data set is input into the weather scene classification model, forward propagation is performed on each batch of DVS sample data, and the weather scene classification model is updated based on the output result of the weather scene classification model, the labeling information of the DVS sample data and the mean square error loss function until the weather scene classification model converges.
[0098] The device of this embodiment can be used to execute the solutions of the above-mentioned method embodiments. Its implementation principles and technical effects are similar and will not be described in detail here.
[0099] An embodiment of the present application also provides an electronic device, comprising: a processor, the processor being used to connect to a memory, the memory storing a computer program / instruction that can be run on the processor, and the computer program / instruction, when executed by the processor, implements the steps of the weather scene classification method as described above.
[0100] An embodiment of the present application also provides a vehicle, comprising the electronic device as described above.
[0101] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the weather scene classification method as described in any one of the above items are implemented.
[0102] An embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the weather scene classification method as described in any one of the above items.
[0103] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0104] Through the above description of the embodiments, those skilled in the art will clearly understand that the methods of the above embodiments can be implemented using a computer software product and the necessary general-purpose hardware platform, or alternatively, hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes a number of instructions for causing a terminal or network-side device to execute the methods described in the various embodiments of this application.
[0105] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms of implementation methods without departing from the purpose of this application and the scope of protection of the claims. These implementation methods are all within the protection of this application.
Claims
1. A weather scene classification method, characterized in that: The method comprises: Acquire dynamic visual sensor (DVS) data, wherein the DVS data includes an event stream, wherein each event in the event stream includes time information, position information, and brightness information; The DVS data is input into a weather scene classification model to obtain a weather scene target classification result, wherein the weather scene classification model is trained based on a DVS sample data set.
2. The method according to claim 1, characterized in that The weather scene classification model includes: at least two cascaded pulse modules, a flattening module, a classification module and a voting module; Inputting the DVS data into a weather scene classification model to obtain the weather scene target classification result includes: performing pulse processing on the DVS data by the at least two cascaded pulse modules to obtain target pulse characteristics; Flattening the spatial dimension of the target pulse feature by a flattening module to obtain a flattened target pulse feature; Performing feature mapping on the target pulse features through the classification module to obtain an initial classification result of the weather scene; The voting module performs a voting mechanism on the initial classification results of the weather scenes to obtain the target classification results of the weather scenes.
3. The method according to claim 2, characterized in that The pulse processing of the DVS data by the at least two cascaded pulse modules to obtain target pulse characteristics includes: For each pulse module, the previous level pulse features are input into the current level pulse module, spatial features and temporal features are extracted from the previous level pulse features, and the current level pulse features are output. The input of the first level pulse module is the DVS data, and the output of the last level pulse module is the target pulse feature.
4. The method according to claim 3, characterized in that The pulse module includes: a convolution module and a pulse neuron module; The performing spatial feature extraction and temporal feature extraction on the pulse feature of the previous stage and outputting the pulse feature of the current stage includes: The spatial feature extraction is performed on the pulse feature of the previous level through the convolution module to obtain the spatial feature; the temporal feature extraction is performed on the spatial feature through the pulse neuron module to output the pulse feature of the current level.
5. The method according to claim 4, characterized in that The convolution module includes: a convolution unit and a batch normalization unit; the pulse neuron module includes: a pulse neuron unit and a maximum pooling unit; The step of extracting spatial features from the previous-stage pulse features through a convolution module to obtain spatial features includes: Performing a convolution operation on the previous level pulse feature through the convolution unit to obtain a local spatial feature; Normalizing the local spatial features by the batch normalization unit to obtain the spatial features; The step of extracting the temporal features of the spatial features by the pulse neuron module and outputting the pulse features of the current level includes: The spiking neural unit emits pulses to the spatial feature to obtain a pulse sequence; The maximum pooling unit performs a maximum pooling operation on the pulse sequence to obtain the pulse features of the current level.
6. The method according to claim 1, characterized in that The DVS sample data set includes: DVS sample data and annotation information, wherein the annotation information is used to identify the weather scene category. The training process of the weather scene classification model includes: The DVS sample data set is input into the weather scene classification model, forward propagation is performed on each batch of DVS sample data, and the weather scene classification model is updated based on the output result of the weather scene classification model, the labeling information of the DVS sample data and the mean square error loss function until the weather scene classification model converges.
7. An electronic device, characterized in that: include: A processor, the processor being configured to be connected to a memory, the memory storing a computer program / instruction executable on the processor, the computer program / instruction implementing the steps of the weather scene classification method according to any one of claims 1 to 6 when executed by the processor.
8. A vehicle, characterized in that: Comprising the electronic device as claimed in claim 7.
9. A computer-readable storage medium, characterized in that The readable storage medium stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the weather scene classification method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the weather scene classification method according to any one of claims 1 to 6 are implemented.