A method and system for traffic violation identification based on fused data
By using cross-attention fusion of traffic scene images and radar point cloud data, the problem of unstable traffic violation recognition in existing technologies has been solved, achieving a more reliable recognition effect.
Patent Information
- Application Number
- CN202510929449.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-07-07
AI Technical Summary
In the identification of traffic violations, existing technologies struggle to capture comprehensive traffic scene information using a single sensing method, and the lack of multimodal data fusion methods leads to unstable identification results and low reliability.
By mining semantic vectors from traffic scene images and radar point cloud data, cross-attention mapping and cross-attention fusion are performed. Image occlusion parameters and point cloud occlusion parameters are used to control the cross-fusion of semantic vectors, forming a traffic scene fusion vector for illegal event identification.
It improves the reliability of traffic violation identification, ensures accurate representation of semantic information, avoids interference from invalid information, and enhances the reliability of identification results.
Smart Images

Figure CN120431533B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to a method and system for identifying traffic violations based on fused data. Background Technology
[0002] Traffic violation recognition technology plays a crucial role in intelligent transportation systems and is widely used in urban road monitoring and traffic management. However, existing technologies still face numerous challenges when dealing with complex traffic scenarios. For example, current technologies typically employ either image analysis or radar data for violation recognition, and relying on a single sensing method is insufficient to comprehensively capture key information within a traffic scene. Image analysis depends on high-resolution images and complex algorithms, while radar data is susceptible to factors such as weather and obstacles, leading to unstable recognition results. Furthermore, in complex urban traffic environments, occlusion from pedestrians and vehicles can significantly increase the risk of misjudgment when relying on a single data source. Existing technologies often depend on a single data source when dealing with occlusion, lacking effective occlusion handling methods, resulting in decreased recognition accuracy. Finally, existing technologies employ relatively simple methods for fusing multimodal data (such as image and radar data), lacking systematic research and failing to fully utilize the complementarity of both, thus affecting overall recognition performance and resulting in relatively low reliability. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method and system for identifying traffic violations based on fused data, so as to improve the problem of relatively low reliability of traffic violation identification in the prior art.
[0004] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:
[0005] A method for identifying traffic violations based on fused data includes:
[0006] The traffic scene semantic vector corresponding to the traffic scene image of the target traffic scene and the radar point cloud semantic vector corresponding to the radar point cloud data are extracted. The radar point cloud data is formed by collecting information about the target traffic scene using millimeter-wave radar.
[0007] The traffic scene semantic vector and the radar point cloud semantic vector are respectively cross-attention mapped to form a first attention vector of the traffic scene semantic vector, a second attention vector of the radar point cloud semantic vector and a third attention vector.
[0008] Based on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data, the first attention vector of the traffic scene semantic vector, the second attention vector and the third attention vector of the radar point cloud semantic vector are cross-attention fused to form a traffic scene fusion vector.
[0009] Based on the traffic scene fusion vector, illegal events are identified to form the illegal event identification result of the target traffic scene.
[0010] In some preferred embodiments, in the above-described traffic violation incident identification method based on fused data, the step of mining the traffic scene semantic vector corresponding to the traffic scene image of the target traffic scene and the radar point cloud semantic vector corresponding to the radar point cloud data includes:
[0011] A two-dimensional Fourier transform is performed on the traffic scene image of the target traffic scene to form a traffic scene spectrogram, and a depth convolution is performed on the traffic scene spectrogram to form a traffic scene semantic vector corresponding to the traffic scene image.
[0012] The discrete point cloud data in the radar point cloud data of the target traffic scene is expanded into a continuous three-dimensional grid to form radar point cloud three-dimensional data. In the radar point cloud three-dimensional data, the data at the position corresponding to the point cloud data is 1, and the data at other positions is 0.
[0013] A three-dimensional Fourier transform is performed on the radar point cloud three-dimensional data to form a radar point cloud spectrum.
[0014] The radar point cloud spectrum is subjected to depth convolution to form the radar point cloud semantic vector corresponding to the radar point cloud data.
[0015] In some preferred embodiments, in the above-described traffic violation incident identification method based on fused data, the step of performing deep convolution on the radar point cloud spectrogram to form a radar point cloud semantic vector corresponding to the radar point cloud data includes:
[0016] The radar point cloud spectrum is subjected to channel-wise convolution to form multiple channel convolution vectors;
[0017] The multiple channel convolution vectors are concatenated to form a channel concatenated vector;
[0018] Based on the dynamic adjustment parameters, the channel splicing vector is subjected to M time steps of association semantic mining operation to form a radar point cloud semantic vector. The dynamic adjustment parameters are used to adjust the association mining parameters at different time steps, and the association mining parameters are used to adjust the distribution of the mined association semantics in the association semantic mining operation.
[0019] In some preferred embodiments, in the above-described traffic violation incident identification method based on fused data, the step of performing M time-step association semantic mining operations on the channel splicing vector according to dynamically adjusted parameters to form a radar point cloud semantic vector includes:
[0020] For the a-th time step out of M time steps, the associated semantic vector of the (a-1)-th time step and the dynamic adjustment parameter of the a-th time step are concatenated to form the semantic vector to be mined for the a-th time step, where a is a positive integer less than or equal to M. When a=1, the associated semantic vector of the (a-1)-th time step is the channel concatenation vector.
[0021] Perform an association semantic mining operation on the semantic vector to be mined at the a-th time step to form the association semantic vector at the a-th time step. The association semantic mining operation includes, in sequence, an attention operation, a first vector compression operation, a nonlinear mapping operation, and a second vector compression operation.
[0022] If a is less than M, then a = a + 1 is determined, and the step of concatenating the associated semantic vector of the (a-1)th time step and the dynamic adjustment parameter of the ath time step to form the semantic vector to be mined in the ath time step is executed in reverse order.
[0023] If a equals M, then the associated semantic vector at the a-th time step is determined to be the radar point cloud semantic vector corresponding to the radar point cloud data.
[0024] In some preferred embodiments, in the above-described traffic violation incident identification method based on fused data, the step of performing semantic mining operations on the channel splicing vector for M time steps according to dynamic adjustment parameters to form a radar point cloud semantic vector further includes a step of determining the dynamic adjustment parameters for the a-th time step, which includes:
[0025] The configuration association mining parameters are determined, wherein the configuration association mining parameters are formed by learning the mapping relationship between traffic scene image samples, radar point cloud data samples and violation event tags;
[0026] For the a-th time step among the M time steps, when a is less than a first preset value, the configured association mining parameters are used as the dynamic adjustment parameters for the a-th time step;
[0027] When a is greater than or equal to the first preset value and less than the second preset value, the configuration association mining parameter and the first adjustment parameter are summed to form the dynamic adjustment parameter for the a-th time step;
[0028] When a is greater than or equal to the second preset value and less than or equal to M, the configuration associated mining parameter and the second adjustment parameter are summed to form the dynamic adjustment parameter for the a-th time step, wherein the second adjustment parameter is greater than the first adjustment parameter.
[0029] In some preferred embodiments, in the above-described traffic violation incident recognition method based on fused data, the step of performing cross-attention fusion of the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector of the radar point cloud semantic vector based on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data to form a traffic scene fusion vector includes:
[0030] The image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data are fused to form fused occlusion parameters. The image occlusion parameters and the point cloud occlusion parameters are formed by learning the mapping relationship between traffic scene image samples, radar point cloud data samples and violation event labels.
[0031] Based on the fusion masking parameters, the first attention vector of the traffic scene semantic vector, the second attention vector and the third attention vector of the radar point cloud semantic vector are cross-attention fused to form an attention fusion vector;
[0032] The attention fusion vector is mapped into a vector space to form a traffic scene fusion vector.
[0033] In some preferred embodiments, in the above-described traffic violation incident identification method based on fused data, the step of fusing the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data to form fused occlusion parameters includes:
[0034] A bitwise OR operation is performed on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data to form a first fusion occlusion parameter, wherein the image occlusion parameter belongs to a matrix composed of 0 and 1, and the point cloud occlusion parameter belongs to a matrix composed of 0 and 1.
[0035] A bitwise AND operation is performed on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data to form a second fused occlusion parameter;
[0036] The first fusion masking parameter and the second fusion masking parameter are used as fusion masking parameters.
[0037] In some preferred embodiments, in the above-described traffic violation incident identification method based on fused data, the step of performing cross-attention fusion of the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector based on the fusion occlusion parameters to form an attention fusion vector includes:
[0038] The first attention vector of the traffic scene semantic vector and the second attention vector of the radar point cloud semantic vector are matrix multiplied to form the attention parameter distribution;
[0039] The attention parameter distribution and the fusion masking parameters, including the first fusion masking parameter and the second fusion masking parameter, are respectively added bitwise to form the first adjustment parameter distribution and the second adjustment parameter distribution;
[0040] Based on the first adjustment parameter distribution and the second adjustment parameter distribution respectively, the third attention vector of the radar point cloud semantic vector is weighted and summed to form the first fusion vector and the second fusion vector;
[0041] An attention fusion vector is formed based on the first fusion vector and the second fusion vector.
[0042] In some preferred embodiments, in the above-described traffic violation incident recognition method based on fused data, the step of performing cross-attention fusion of the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector of the radar point cloud semantic vector based on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data to form a traffic scene fusion vector includes:
[0043] Based on the length of the image occlusion parameters of the traffic scene image, the first attention vector of the traffic scene semantic vector is segmented to form a local first attention vector. Based on the point cloud occlusion parameters of the radar point cloud data, the second attention vector and the third attention vector of the radar point cloud semantic vector are segmented to form a local second attention vector and a local third attention vector, respectively.
[0044] The image occlusion parameters and the local first attention vector are multiplied bitwise to form a local first adjustment vector. The point cloud occlusion parameters are multiplied bitwise with the local second attention vector and the local third attention vector to form a local second adjustment vector and a local third adjustment vector, respectively.
[0045] Based on the attention parameter distribution between the local first adjustment vector and the local second adjustment vector, the local third adjustment vector is weighted and summed to form a local attention fusion vector;
[0046] An attention fusion vector is formed based on the attention fusion vector of each locality.
[0047] This invention also provides a traffic violation incident identification system based on fused data, including a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the computer program to implement the above-described traffic violation incident identification method based on fused data.
[0048] The traffic violation incident recognition method and system based on fused data provided in this invention firstly mines the traffic scene semantic vector corresponding to the traffic scene image and the radar point cloud semantic vector corresponding to the radar point cloud data; secondly, the traffic scene semantic vector and the radar point cloud semantic vector are cross-attention mapped to form a first attention vector, a second attention vector, and a third attention vector; then, based on image occlusion parameters and point cloud occlusion parameters, the first attention vector, the second attention vector, and the third attention vector are cross-attention fused to form a traffic scene fusion vector; finally, the traffic scene fusion vector is used to identify traffic violations, resulting in a traffic violation incident recognition result. Based on the above method, since the cross-attention fusion process involves occlusion based on image occlusion parameters and point cloud occlusion parameters, the cross-fusion of the corresponding semantic vectors can be effectively controlled. This allows for attention to the effective semantic information in the corresponding semantic vectors, avoiding interference from invalid semantic information. Consequently, the semantic representation accuracy of the resulting traffic scene fusion vector is higher, thus ensuring the reliability of the violation event identification results obtained based on the traffic scene fusion vector. Compared to conventional techniques such as using data from a single data source for identification and analysis or simply fusing data from multiple data sources, this method has higher reliability. Therefore, it can improve the problem of relatively low reliability in traffic violation event identification in existing technologies.
[0049] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0050] Figure 1 This is a structural block diagram of a traffic violation incident recognition system based on fused data, provided in an embodiment of the present invention.
[0051] Figure 2 This is a schematic diagram of the modules included in the traffic violation incident identification device based on fused data provided in an embodiment of the present invention.
[0052] Figure 3 The flowchart illustrates the steps of the traffic violation incident identification method based on fused data provided in this embodiment of the invention.
[0053] Figure 4 This is a schematic diagram of semantic mining provided in an embodiment of the present invention.
[0054] Figure 5 This is a schematic diagram of depthwise convolution provided in an embodiment of the present invention.
[0055] Figure 6 This is a schematic diagram of occlusion parameter fusion provided in an embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0057] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0058] like Figure 1 As shown in the figure, this embodiment of the invention provides a traffic violation incident recognition system based on fused data. The traffic violation incident recognition system based on fused data may include a memory and a processor.
[0059] Specifically, the memory and processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, they can be electrically connected via one or more communication buses or signal lines. The memory may store at least one software functional module (such as...) that exists in the form of software or firmware. Figure 2 The traffic violation incident identification device based on fused data shown includes various modules. The processor can be used to execute executable computer programs stored in the memory, thereby implementing the traffic violation incident identification method based on fused data provided in the embodiments of the present invention (as described below).
[0060] Optionally, the traffic violation identification device based on fused data may include:
[0061] The semantic mining module is used to mine the traffic scene semantic vector corresponding to the traffic scene image of the target traffic scene and the radar point cloud semantic vector corresponding to the radar point cloud data. The radar point cloud data is formed by collecting information about the target traffic scene using millimeter-wave radar.
[0062] The attention mapping module is used to perform cross-attention mapping on the traffic scene semantic vector and the radar point cloud semantic vector respectively, to form a first attention vector of the traffic scene semantic vector, a second attention vector of the radar point cloud semantic vector and a third attention vector.
[0063] The attention fusion module is used to perform cross-attention fusion of the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector of the radar point cloud based on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data, to form a traffic scene fusion vector.
[0064] The illegal event identification module is used to identify illegal events based on the traffic scene fusion vector to form the illegal event identification result of the target traffic scene.
[0065] Optionally, the memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.
[0066] Optionally, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0067] and, Figure 1 The structure shown is for illustrative purposes only. The traffic violation identification system based on fused data may also include... Figure 1The more or fewer components shown, or having the same Figure 1 The different configurations shown may include, for example, a communication unit for exchanging information with other devices (such as image sensors, millimeter-wave radar, and other traffic scene information acquisition devices).
[0068] In one alternative example, the traffic violation identification system based on fused data can be a server with data processing capabilities.
[0069] Combination Figure 3 This invention also provides a traffic violation incident identification method based on fused data, which can be applied to the aforementioned traffic violation incident identification system based on fused data. The method steps defined in the relevant process of the traffic violation incident identification method based on fused data can be implemented by the traffic violation incident identification system based on fused data (hereinafter referred to as the identification system). The following will describe... Figure 3 The specific process shown will be explained in detail.
[0070] Step S110: Extract the traffic scene semantic vector corresponding to the traffic scene image of the target traffic scene and the radar point cloud semantic vector corresponding to the radar point cloud data.
[0071] In this embodiment of the invention, the recognition system can mine traffic scene semantic vectors corresponding to traffic scene images of the target traffic scene and radar point cloud semantic vectors corresponding to radar point cloud data. That is, semantic mining can be performed on the traffic scene images to form corresponding traffic scene semantic vectors, i.e., extracting potential semantic information from the traffic scene images and representing it in vector form. Similarly, semantic mining can be performed on the radar point cloud data to form corresponding radar point cloud semantic vectors, i.e., extracting potential semantic information from the radar point cloud data and representing it in vector form. The radar point cloud data is generated by collecting information from the target traffic scene (such as a traffic intersection) using millimeter-wave radar. The traffic scene images can be generated by collecting information from cameras installed in the target traffic scene.
[0072] Step S120: Perform cross-attention mapping on the traffic scene semantic vector and the radar point cloud semantic vector respectively to form a first attention vector of the traffic scene semantic vector, a second attention vector of the radar point cloud semantic vector, and a third attention vector.
[0073] In this embodiment of the invention, after mining the traffic scene semantic vector and the radar point cloud semantic vector, the recognition system can perform cross-attention mapping on the traffic scene semantic vector and the radar point cloud semantic vector respectively to form a first attention vector of the traffic scene semantic vector, a second attention vector of the radar point cloud semantic vector, and a third attention vector. For example, the first attention vector can also be called a query vector, the second attention vector can also be called a key vector, and the third attention vector can also be called a value vector.
[0074] Step S130: Based on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data, cross-attention fusion is performed on the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector to form a traffic scene fusion vector.
[0075] In this embodiment of the invention, after forming a first attention vector, a second attention vector, and a third attention vector, the recognition system can perform cross-attention fusion of the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector based on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data to form a traffic scene fusion vector. The image occlusion parameters and the point cloud occlusion parameters can be learned by studying the mapping relationship between traffic scene image samples, radar point cloud data samples, and violation event labels. That is, by learning the mapping relationship between traffic scene image samples, radar point cloud data samples, and violation event labels, it is possible to determine which semantic information is valid and which is invalid. Then, this semantic information relationship learned from the samples and labels is applied to the semantic fusion of traffic scene images and radar point cloud data, enabling the system to focus on valid semantic information and occlude or discard invalid semantic information during the fusion process, thereby ensuring that the formed traffic scene fusion vector has high semantic representation accuracy.
[0076] Step S140: Identify illegal events based on the traffic scene fusion vector to form the illegal event identification result of the target traffic scene.
[0077] In this embodiment of the invention, after forming the traffic scene fusion vector, the recognition system can perform violation event recognition based on the traffic scene fusion vector to form a violation event recognition result for the target traffic scene. For example, classification output can be performed based on the traffic scene fusion vector, such as binary classification or multi-class classification. The result of binary classification can be used to indicate whether there is a traffic violation in the target traffic scene, and the result of multi-class classification can be used to indicate the specific type of traffic violation in the target traffic scene (including no traffic violation and at least one other traffic violation, such as speeding, running a red light, crossing the line, etc.).
[0078] Based on the above method, since the cross-attention fusion process involves occlusion based on image occlusion parameters and point cloud occlusion parameters, the cross-fusion of the corresponding semantic vectors can be effectively controlled. This allows for attention to the effective semantic information in the corresponding semantic vectors, avoiding interference from invalid semantic information. Consequently, the semantic representation accuracy of the resulting traffic scene fusion vector is higher, thus ensuring the reliability of the violation event identification results obtained based on the traffic scene fusion vector. Compared to conventional techniques such as using data from a single data source for identification and analysis or simply fusing data from multiple data sources, this method has higher reliability. Therefore, it can improve the problem of relatively low reliability in traffic violation event identification in existing technologies.
[0079] In the first part, regarding step S110, it should be noted that the specific process of mining the traffic scene semantic vector corresponding to the traffic scene image of the target traffic scene and the radar point cloud semantic vector corresponding to the radar point cloud data is not restricted and can be configured according to actual needs.
[0080] For example, in one feasible implementation, the traffic scene image can be convolutionally processed to form a corresponding traffic scene semantic vector. Furthermore, the radar point cloud data can be processed based on a graph neural network (such as DGCNN, Dynamic Graph CNN, a deep learning model for point cloud processing) to form a corresponding radar point cloud semantic vector.
[0081] For example, in another feasible implementation, considering that in subsequent steps (such as step S130), the semantic vectors corresponding to the data of the two modalities will be cross-attention fused, in order to ensure the reliability of the cross-attention fusion, the semantic vectors formed can be in similar semantic spaces through corresponding semantic mining, thereby ensuring the adaptability between semantic vectors during fusion. Based on this, the above-mentioned step S110 can further include steps S111, S112, S113 and S114, the specific contents of each step are as follows (in conjunction with...). Figure 4 (As shown).
[0082] Step S111: Perform a two-dimensional Fourier transform on the traffic scene image of the target traffic scene to form a traffic scene spectrogram, and perform a depth convolution on the traffic scene spectrogram to form a traffic scene semantic vector corresponding to the traffic scene image.
[0083] In this embodiment of the invention, a two-dimensional Fourier transform can be performed on the traffic scene image of the target traffic scene to form a traffic scene spectrogram. Then, a depthwise convolution can be performed on the traffic scene spectrogram to form a traffic scene semantic vector corresponding to the traffic scene image. That is, the traffic scene image can first be transformed from the spatial domain to the frequency domain to obtain the traffic scene spectrogram. Then, a depthwise convolution can be performed on the traffic scene spectrogram to capture its high-level, abstract semantic information, thereby obtaining the traffic scene semantic vector. The formula for performing a two-dimensional Fourier transform on the traffic scene image is as follows:
[0084] ;
[0085] Wherein, the traffic scene image has a size of M*N, f(x,y) is the pixel value at position (x,y) in the traffic scene image, and F(u,v) is the value of the transformed spectrogram at position (u,v). M and N are the number of rows and columns of the traffic scene image, respectively. j is the imaginary unit. Furthermore, the network architecture for performing depthwise convolution on the traffic scene spectrogram may include:
[0086] 2D convolutional layer (convolutional kernel 1): The convolutional kernel size is 3×3, the number of output channels is 64, and the activation function is ReLU (Rectified Linear Unit).
[0087] Max pooling layer 1: pooling size is 2x2, step size is 2;
[0088] 2D convolutional layer (kernel 2): kernel size 3×3, output channels 128, activation function ReLU;
[0089] Max pooling layer 2: pooling size is 2x2, step size is 2;
[0090] 2D convolutional layer (kernel 3): The kernel size is 3×3, the number of output channels is 256, and the activation function is ReLU;
[0091] Max pooling layer 3: pooling size is 2x2, step size is 2;
[0092] Fully connected layer: output size is 1x512, activation function is ReLU.
[0093] Step S112: Expand the discrete point cloud data in the radar point cloud data of the target traffic scene into a continuous three-dimensional grid to form radar point cloud three-dimensional data.
[0094] In this embodiment of the invention, discrete point cloud data in the radar point cloud data of the target traffic scene can be extended into a continuous three-dimensional grid to form three-dimensional radar point cloud data. In this three-dimensional radar point cloud data, the data at the position corresponding to the point cloud data is 1, and the data at other positions is 0. For example, firstly, the minimum and maximum coordinate ranges of the three-dimensional grid need to be determined, i.e., the minimum and maximum coordinate values of the radar point cloud data. For example, if the x-range of the radar point cloud data is [-10, 10], the y-range is [-10, 10], and the z-range is [-5, 5], then the grid range can be set as x: -10 to 10, y: -10 to 10, z: -5 to 5. Then, the grid resolution is set, i.e., the resolution of the grid in each dimension is determined, i.e., how many grids will be divided into on each axis. For example, each dimension can be divided into 100 grids, so the entire grid will contain 100 × 100 × 100 = 1,000,000 grid points. Next, the origin can be set at the center of the radar point cloud data to reduce errors caused by offset. For example, if the x-range of the point cloud is [-10, 10], the origin is set at x=0, and the y and z are handled similarly. Thus, for each grid point, if the grid point has point cloud data, it is assigned a value of 1; if the grid point does not have point cloud data, it is assigned a value of 0, thereby forming the three-dimensional radar point cloud data.
[0095] Step S113: Perform a three-dimensional Fourier transform on the radar point cloud three-dimensional data to form a radar point cloud spectrum.
[0096] In this embodiment of the invention, after forming the three-dimensional radar point cloud data, a three-dimensional Fourier transform can be performed on the radar point cloud three-dimensional data to form a radar point cloud spectrum, that is, converting the radar point cloud three-dimensional data in the spatial domain to the frequency domain. The formula for performing the three-dimensional Fourier transform on the radar point cloud three-dimensional data is as follows:
[0097] ;
[0098] Wherein, the dimensions of the radar point cloud 3D data are M*N*P, f(x,y,z) is the value at position (x,y,z) in the radar point cloud 3D data, and F(k,l,m) is the value at position (k,l,m) in the transformed spectrum. M, N, and P are the number of rows, columns, and channels of the radar point cloud 3D data, respectively. j is the imaginary unit.
[0099] Step S114: Perform depth convolution on the radar point cloud spectrum to form a radar point cloud semantic vector corresponding to the radar point cloud data.
[0100] In this embodiment of the invention, after forming the radar point cloud spectrogram, a deep convolution can be performed on the radar point cloud spectrogram to form a radar point cloud semantic vector corresponding to the radar point cloud data. Thus, since both the traffic scene semantic vector and the radar point cloud semantic vector are formed by deep convolution of frequency domain data (i.e., the spectrogram), the reliability of cross-attention fusion can be improved in subsequent steps due to their similar semantic spaces.
[0101] Optionally, in step S114 above, the specific process of performing depth convolution on the radar point cloud spectrum is not limited and can be configured according to actual needs.
[0102] For example, in one feasible implementation, in order to more accurately capture the correlation between key semantic information and thus improve the accuracy of semantic vectors, the above step S114 may further include steps S114a, S114b, and S114c, the specific contents of each step of which are as follows (in conjunction with...). Figure 5 ).
[0103] Step S114a: Perform channel-wise convolution on the radar point cloud spectrum to form multiple channel convolution vectors.
[0104] In this embodiment of the invention, the radar point cloud spectrogram can be convolved channel by channel to form multiple channel convolution vectors. As mentioned earlier, the radar point cloud spectrogram is three-dimensional, i.e., x*y*z, where z represents the number of channels. Therefore, convolution can be performed on the two-dimensional spectrogram of each channel separately, thus forming multiple channel convolution vectors. Based on this, by performing channel-specific convolution on different feature layers of the spectrogram, unique features in different channels can be effectively extracted. This helps to capture multi-dimensional information and enrich the diversity of feature representations.
[0105] Step S114b: The multiple channel convolution vectors are concatenated to form a channel concatenated vector.
[0106] In this embodiment of the invention, after forming multiple channel convolution vectors, the multiple channel convolution vectors can be concatenated to form a channel concatenated vector. In this way, multiple channel convolution vectors can be initially fused simply and quickly to obtain a semantically rich channel concatenated vector, that is, the features extracted from different channels are integrated together to form a comprehensive semantic description.
[0107] Step S114c: Based on the dynamic adjustment parameters, perform M time-step association semantic mining operations on the channel splicing vector to form a radar point cloud semantic vector.
[0108] In this embodiment of the invention, after forming the channel splicing vector, the channel splicing vector can be subjected to M time-step association semantic mining operations based on dynamically adjusted parameters to form a radar point cloud semantic vector. The dynamically adjusted parameters are used to adjust the association mining parameters at different time steps, and these parameters are used to adjust the distribution of mined association semantics during the association semantic mining operations. Based on this, by adjusting the association mining parameters, the feature associations at different time steps can be flexibly focused on, enhancing the ability to capture complex semantic information and enabling more accurate capture of the associations between key features, thereby improving the accuracy of the semantic vector.
[0109] Optionally, in step S114c above, the specific process of performing M time-step association semantic mining on the channel concatenation vector is not limited. For example, in a feasible implementation, in order to capture effective semantic information during the association semantic mining operation, step S114c above may further include the following:
[0110] First, for the a-th time step out of M time steps, the associated semantic vector of the (a-1)-th time step and the dynamic adjustment parameter (which can be a matrix or vector) of the a-th time step can be concatenated to form the semantic vector to be mined for the a-th time step, where a is a positive integer less than or equal to M. When a=1, the associated semantic vector of the (a-1)-th time step is the channel concatenation vector. That is, in the first time step, the channel concatenation vector and the dynamic adjustment parameter of the first time step can be concatenated to form the semantic vector to be mined for the first time step. In the second time step, the associated semantic vector of the first time step and the dynamic adjustment parameter of the second time step can be concatenated to form the semantic vector to be mined for the second time step. And so on, the semantic vector to be mined for the third time step can be formed, and so on.
[0111] Secondly, the semantic vector to be mined at the a-th time step can be subjected to an association semantic mining operation to form the association semantic vector at the a-th time step. The association semantic mining operation includes, in sequence, an attention operation (e.g., which can be implemented by a self-attention network layer to achieve association mining of different channel convolution vectors and historical semantic information, wherein the dynamically adjusted parameters are related to the learned traffic scene image samples, radar point cloud data samples and illegal event labels, i.e., representing historical semantic information), a first vector compression operation (e.g., which can be implemented by a pooling network layer), a nonlinear mapping operation (e.g., which can be implemented by a fully connected network layer), and a second vector compression operation (e.g., which can be implemented by a pooling network layer).
[0112] If a is less than M, then a = a + 1 is determined (that is, after the first time step is completed, the associated semantic mining operation of the second time step is performed), so as to execute the step of concatenating the associated semantic vector of the (a-1)th time step and the dynamic adjustment parameter of the ath time step for the M time steps to form the semantic vector to be mined in the ath time step. In this way, the semantic vector to be mined and the associated semantic vector of each time step can be obtained.
[0113] Furthermore, if a equals M, then the associated semantic vector of the a-th time step is determined to be the radar point cloud semantic vector corresponding to the radar point cloud data, that is, the associated semantic vector of the last time step is determined to be the radar point cloud semantic vector corresponding to the radar point cloud data.
[0114] Alternatively, in one feasible implementation, step S114c may further include a specific process for determining the dynamic adjustment parameters for the a-th time step, such as:
[0115] First, the configuration association mining parameters can be determined. These parameters are formed by learning the mapping relationship between traffic scene image samples, radar point cloud data samples, and violation event tags, which represents the semantic information of the samples and tags.
[0116] Secondly, for the a-th time step among the M time steps, when a is less than the first preset value, the configured association mining parameters are used as the dynamic adjustment parameters for the a-th time step. In other words, in the early stage of the association semantic mining operation, the configured association mining parameters can be directly used as the dynamic adjustment parameters, so as to achieve full fusion of semantic information in samples and labels.
[0117] Then, when a is greater than or equal to the first preset value and less than the second preset value, the configuration association mining parameters and the first adjustment parameters are summed to form the dynamic adjustment parameters for the a-th time step. That is, in the middle of the association semantic mining operation, the configuration association mining parameters can be adjusted to a certain extent, so that the semantic information in the samples and labels will be deformed to a certain extent. Therefore, when performing association semantic mining, the degree of attention is reduced due to the decrease in correlation, thereby achieving attention to the channel splicing vector.
[0118] Finally, when a is greater than or equal to the second preset value and less than or equal to M, the configuration association mining parameters and the second adjustment parameters are summed to form the dynamic adjustment parameters for the a-th time step. The second adjustment parameter is greater than the first adjustment parameter. In other words, in the later stages of the association semantic mining operation, the configuration association mining parameters can be adjusted to a large extent, causing the semantic information in the samples and labels to be significantly deformed. Therefore, during the association semantic mining operation, the degree of attention is also greatly reduced due to the significant reduction in correlation, thereby achieving further attention to the channel splicing vector.
[0119] For example, M can be equal to 10. Based on this, the first preset value can be equal to 3, and the second preset value can be equal to 8. Thus, in the first two time steps, the configuration association mining parameters can be directly used as dynamic adjustment parameters. In the third to seventh time steps, the configuration association mining parameters can be increased to a certain extent. In the eighth to tenth time steps, the configuration association mining parameters can be increased to a greater extent. Furthermore, the first adjustment parameter can be greater than 0 and less than 0.5, and the second adjustment parameter can be greater than 0.5 and less than 1.
[0120] In the second part, regarding step S120, it should be noted that the specific process of performing cross-attention mapping on the traffic scene semantic vector and the radar point cloud semantic vector is not limited and can be selected according to actual needs.
[0121] For example, in one feasible implementation, cross-attention mapping can be achieved through a cross-attention network layer. This cross-attention network layer may have a first weight matrix, a second weight matrix, and a third weight matrix (formed by learning the mapping relationship between traffic scene image samples, radar point cloud data samples, and violation event labels). Then, the first weight matrix can be multiplied with the traffic scene semantic vector to form a first attention vector of the traffic scene semantic vector. Furthermore, the second weight matrix can be multiplied with the radar point cloud semantic vector to form a second attention vector of the radar point cloud semantic vector, and the third weight matrix can be multiplied with the radar point cloud semantic vector to form a third attention vector of the radar point cloud semantic vector.
[0122] The third part, regarding step S130, should be explained as follows: the specific process of cross-attention fusion of the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector is not restricted and can be selected according to actual needs.
[0123] For example, in one feasible implementation, in order to achieve the mining and capture of more granular correlation information, step S130 above may include the following:
[0124] The first step involves segmenting the first attention vector of the traffic scene semantic vector based on the length of the image occlusion parameters of the traffic scene image, forming local first attention vectors (that is, each time a local vector of the same length as the image occlusion parameters is extracted from the first attention vector to obtain a corresponding local first attention vector, wherein two adjacent local first attention vectors may or may not intersect). Then, based on the point cloud occlusion parameters of the radar point cloud data, the second and third attention vectors of the radar point cloud semantic vector are segmented respectively to form local second attention vectors and local third attention vectors (that is, each time a local first attention vector of the same length as the image occlusion parameters is extracted from the first attention vector, a corresponding local first attention vector is obtained; where adjacent local first attention vectors may or may not intersect). Each time, a local vector of the same length as the point cloud occlusion parameter is extracted from the adjacent second attention vector and third attention vector to obtain a corresponding local second attention vector and a local third attention vector. The two adjacent local second attention vectors may or may not intersect; the two adjacent local third attention vectors may or may not intersect. In addition, the image occlusion parameter and the point cloud occlusion parameter are both matrices composed of 0 and 1, and are formed by learning the mapping relationship between traffic scene image samples, radar point cloud data samples and violation event labels. Initially, they can be random matrices, all-zero matrices, or all-one matrices.
[0125] The second step involves performing a bitwise multiplication operation between the image occlusion parameters and the local first attention vector to form a local first adjustment vector (thus, non-essential semantic information in the local first attention vector can be occluded). The point cloud occlusion parameters are then performed bitwise multiplication operations between the local second attention vector and the local third attention vector to form a local second adjustment vector and a local third adjustment vector (thus, non-essential semantic information in the local second attention vector and the local third attention vector can be occluded).
[0126] The third step involves calculating a weighted sum of the local third adjustment vector based on the attention parameter distribution between the local first adjustment vector and the local second adjustment vector (for example, the attention parameter distribution can be obtained by multiplying the transpose of the local first adjustment vector and the local second adjustment vector). This results in a local attention fusion vector.
[0127] The second step is to form an attention fusion vector based on each local attention fusion vector. For example, each local attention fusion vector (a corresponding local attention fusion vector can be formed based on each segmentation) can be concatenated, and then convolution and pooling can be performed to achieve compression, thereby obtaining the attention fusion vector.
[0128] For example, in another feasible implementation, in order to achieve global fusion and make full use of the complementary information of image and point cloud data through the cross-attention mechanism, the above step S130 may further include steps S131, S132 and S133, as detailed below.
[0129] Step S131: The image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data are fused to form fused occlusion parameters.
[0130] In this embodiment of the invention, the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data can be fused to form fused occlusion parameters. Thus, the fused occlusion parameters can be used to jointly represent the occlusion parameters of the two dimensions. The image occlusion parameters and the point cloud occlusion parameters are formed by learning the mapping relationship between traffic scene image samples, radar point cloud data samples, and violation event labels; that is, they are formed as parameters in the corresponding neural network model during the model training process.
[0131] Step S132: Based on the fusion masking parameters, the first attention vector of the traffic scene semantic vector, the second attention vector and the third attention vector of the radar point cloud semantic vector are cross-attention fused to form an attention fusion vector.
[0132] In this embodiment of the invention, after forming the fusion masking parameters, the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector can be cross-attention fused based on the fusion masking parameters to form an attention fusion vector. That is, during the cross-attention fusion process, local attention information is masked through the fusion masking parameters to increase the focus on important semantic information, thereby achieving the mining of important semantic information.
[0133] Step S133: Map the attention fusion vector to a vector space to form a traffic scene fusion vector.
[0134] In this embodiment of the invention, after forming the attention fusion vector, the attention fusion vector can be mapped into a vector space to form a traffic scene fusion vector. For example, the attention fusion vector can be processed through a fully connected network layer to capture more complex semantic information, thereby obtaining the corresponding traffic scene fusion vector.
[0135] Optionally, in step S131 above, the specific process of fusing the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data is not limited. For example, in a feasible implementation, in order to fully utilize the occlusion parameters of both dimensions to capture complex semantic information, step S131 above may further include the following (in conjunction with...). Figure 6 As shown, where black represents 0 and white represents 1):
[0136] The first step is to perform a bitwise OR operation on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data (i.e., if any one of the same position is 1, the result is 1) to form the first fusion occlusion parameter, wherein the image occlusion parameter belongs to a matrix composed of 0 and 1, and the point cloud occlusion parameter belongs to a matrix composed of 0 and 1.
[0137] The second step is to perform a bitwise AND operation on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data (i.e., if any one of the same position is 0, the result is 0) to form the second fused occlusion parameters.
[0138] The third step is to use the first fusion masking parameter and the second fusion masking parameter as fusion masking parameters, that is, the fusion masking parameters include the first fusion masking parameter and the second fusion masking parameter.
[0139] Optionally, in step S132 above, the specific process of cross-attention fusion of the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector is not limited. For example, in a feasible implementation, as mentioned above, the fusion occlusion parameters include a first fusion occlusion parameter and a second fusion occlusion parameter. Based on this, step S132 above can further include the following:
[0140] The first step is to perform matrix multiplication on the first attention vector of the traffic scene semantic vector and the second attention vector of the radar point cloud semantic vector to form an attention parameter distribution;
[0141] Secondly, the attention parameter distribution and the first fusion occlusion parameter and the second fusion occlusion parameter included in the fusion occlusion parameter can be added bitwise to form a first adjustment parameter distribution and a second adjustment parameter distribution; that is, the attention parameter distribution and the first fusion occlusion parameter can be added bitwise to form a first adjustment parameter distribution, and the attention parameter distribution and the second fusion occlusion parameter can be added bitwise to form a second adjustment parameter distribution.
[0142] Then, the third attention vector of the radar point cloud semantic vector can be weighted and summed based on the first adjustment parameter distribution and the second adjustment parameter distribution respectively to form a first fusion vector and a second fusion vector; that is, the third attention vector can be weighted and summed based on the first adjustment parameter distribution to form a first fusion vector, and the third attention vector can be weighted and summed based on the second adjustment parameter distribution to form a second fusion vector.
[0143] Finally, an attention fusion vector can be formed based on the first fusion vector and the second fusion vector. For example, the first fusion vector and the second fusion vector can be averaged or summed to form the attention fusion vector. Alternatively, the first fusion vector and the second fusion vector can be concatenated and then processed by convolution, pooling, etc., to form the corresponding attention fusion vector.
[0144] In the fourth part, regarding step S140, it should be noted that the specific process for identifying illegal events based on the traffic scene fusion vector is not limited and can be selected according to the actual situation.
[0145] For example, in one feasible implementation, the traffic scene fusion vector can be fully connected to map it into a fully connected vector of size 1*b, where b equals the number of specific types of traffic violations (e.g., b equals 2 if it is necessary to identify whether there is a traffic violation). Then, the fully connected vector can be mapped (e.g., using a function such as softmax) to obtain a probability distribution. Each value in this probability distribution represents the probability of a traffic violation. The traffic violation with the highest probability can then be used as the violation event identification result of the target traffic scene.
[0146] In summary, the traffic violation incident recognition method and system based on fused data provided by this invention firstly mines the traffic scene semantic vector corresponding to the traffic scene image and the radar point cloud semantic vector corresponding to the radar point cloud data; secondly, it performs cross-attention mapping on the traffic scene semantic vector and the radar point cloud semantic vector respectively to form a first attention vector, a second attention vector, and a third attention vector; then, based on image occlusion parameters and point cloud occlusion parameters, it performs cross-attention fusion on the first attention vector, the second attention vector, and the third attention vector to form a traffic scene fusion vector; finally, it performs violation incident recognition based on the traffic scene fusion vector to form a violation incident recognition result. Based on the above method, since the cross-attention fusion process involves occlusion based on image occlusion parameters and point cloud occlusion parameters, the cross-fusion of the corresponding semantic vectors can be effectively controlled. This allows for attention to the effective semantic information in the corresponding semantic vectors, avoiding interference from invalid semantic information. Consequently, the semantic representation accuracy of the resulting traffic scene fusion vector is higher, thus ensuring the reliability of the violation event identification results obtained based on the traffic scene fusion vector. Compared to conventional techniques such as using data from a single data source for identification and analysis or simply fusing data from multiple data sources, this method has higher reliability. Therefore, it can improve the problem of relatively low reliability in traffic violation event identification in existing technologies.
[0147] In the several embodiments provided in this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0148] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0149] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0150] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying traffic violations based on fused data, characterized in that, include: The traffic scene semantic vector corresponding to the traffic scene image of the target traffic scene and the radar point cloud semantic vector corresponding to the radar point cloud data are extracted. The radar point cloud data is formed by collecting information about the target traffic scene using millimeter-wave radar. The traffic scene semantic vector and the radar point cloud semantic vector are respectively cross-attention mapped to form a first attention vector of the traffic scene semantic vector, a second attention vector of the radar point cloud semantic vector and a third attention vector. A first fusion occlusion parameter is formed by performing a bitwise OR operation on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data, wherein the image occlusion parameter and the point cloud occlusion parameter are both matrices composed of 0s and 1s; a second fusion occlusion parameter is formed by performing a bitwise AND operation on the image occlusion parameters of the traffic scene image and the point cloud occlusion parameters of the radar point cloud data; the first and second fusion occlusion parameters are used as the fusion occlusion parameters, wherein the image occlusion parameters and the point cloud occlusion parameters are formed by learning the mapping relationship between traffic scene image samples, radar point cloud data samples, and violation event tags; based on the fusion occlusion parameters, the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector are cross-attention fused to form an attention fusion vector; the attention fusion vector is then mapped in vector space to form a traffic scene fusion vector; Based on the traffic scene fusion vector, illegal events are identified to form the illegal event identification result of the target traffic scene.
2. The traffic violation incident identification method based on fused data as described in claim 1, characterized in that, The steps of mining the traffic scene semantic vector corresponding to the traffic scene image of the target traffic scene and the radar point cloud semantic vector corresponding to the radar point cloud data include: A two-dimensional Fourier transform is performed on the traffic scene image of the target traffic scene to form a traffic scene spectrogram, and a depth convolution is performed on the traffic scene spectrogram to form a traffic scene semantic vector corresponding to the traffic scene image. The discrete point cloud data in the radar point cloud data of the target traffic scene is expanded into a continuous three-dimensional grid to form radar point cloud three-dimensional data. In the radar point cloud three-dimensional data, the data at the position corresponding to the point cloud data is 1, and the data at other positions is 0. A three-dimensional Fourier transform is performed on the radar point cloud three-dimensional data to form a radar point cloud spectrum. The radar point cloud spectrum is subjected to depth convolution to form the radar point cloud semantic vector corresponding to the radar point cloud data.
3. The traffic violation incident identification method based on fused data as described in claim 2, characterized in that, The step of performing depth convolution on the radar point cloud spectrogram to form the radar point cloud semantic vector corresponding to the radar point cloud data includes: The radar point cloud spectrum is subjected to channel-wise convolution to form multiple channel convolution vectors; The multiple channel convolution vectors are concatenated to form a channel concatenated vector; Based on the dynamic adjustment parameters, the channel splicing vector is subjected to M time steps of association semantic mining operation to form a radar point cloud semantic vector. The dynamic adjustment parameters are used to adjust the association mining parameters at different time steps, and the association mining parameters are used to adjust the distribution of the mined association semantics in the association semantic mining operation.
4. The traffic violation incident identification method based on fused data as described in claim 3, characterized in that, The step of performing semantic mining operations on the channel splicing vector for M time steps based on dynamically adjusted parameters to form a radar point cloud semantic vector includes: For the a-th time step out of M time steps, the associated semantic vector of the (a-1)-th time step and the dynamic adjustment parameter of the a-th time step are concatenated to form the semantic vector to be mined for the a-th time step, where a is a positive integer less than or equal to M. When a=1, the associated semantic vector of the (a-1)-th time step is the channel concatenation vector. Perform an association semantic mining operation on the semantic vector to be mined at the a-th time step to form the association semantic vector at the a-th time step. The association semantic mining operation includes, in sequence, an attention operation, a first vector compression operation, a nonlinear mapping operation, and a second vector compression operation. If a is less than M, then a = a + 1 is determined, and the step of concatenating the associated semantic vector of the (a-1)th time step and the dynamic adjustment parameter of the ath time step to form the semantic vector to be mined in the ath time step is executed in reverse order. If a equals M, then the associated semantic vector at the a-th time step is determined to be the radar point cloud semantic vector corresponding to the radar point cloud data.
5. The traffic violation incident identification method based on fused data as described in claim 4, characterized in that, The step of performing semantic mining operations on the channel splicing vector for M time steps based on the dynamic adjustment parameters to form a radar point cloud semantic vector further includes the step of determining the dynamic adjustment parameters for the a-th time step, which includes: The configuration association mining parameters are determined, wherein the configuration association mining parameters are formed by learning the mapping relationship between traffic scene image samples, radar point cloud data samples and violation event tags; For the a-th time step among the M time steps, when a is less than a first preset value, the configured association mining parameters are used as the dynamic adjustment parameters for the a-th time step; When a is greater than or equal to the first preset value and less than the second preset value, the configuration association mining parameter and the first adjustment parameter are summed to form the dynamic adjustment parameter for the a-th time step; When a is greater than or equal to the second preset value and less than or equal to M, the configuration associated mining parameter and the second adjustment parameter are summed to form the dynamic adjustment parameter for the a-th time step, wherein the second adjustment parameter is greater than the first adjustment parameter.
6. The traffic violation incident identification method based on fused data as described in claim 1, characterized in that, The step of performing cross-attention fusion of the first attention vector of the traffic scene semantic vector, the second attention vector of the radar point cloud semantic vector, and the third attention vector based on the fusion masking parameters to form an attention fusion vector includes: The first attention vector of the traffic scene semantic vector and the second attention vector of the radar point cloud semantic vector are matrix multiplied to form the attention parameter distribution; The attention parameter distribution and the fusion masking parameters, including the first fusion masking parameter and the second fusion masking parameter, are respectively added bitwise to form the first adjustment parameter distribution and the second adjustment parameter distribution; Based on the first adjustment parameter distribution and the second adjustment parameter distribution respectively, the third attention vector of the radar point cloud semantic vector is weighted and summed to form the first fusion vector and the second fusion vector; An attention fusion vector is formed based on the first fusion vector and the second fusion vector.
7. A traffic violation incident identification system based on fused data, characterized in that, It includes a processor and a memory, the memory being used to store a computer program, and the processor being used to execute the computer program to implement the traffic violation event identification method based on fused data as described in any one of claims 1-6.
Citation Information
Patent Citations
Traffic target detection method and system based on cross-modal cross attention mechanism
CN117173399A
Detection method and device of protection equipment and electronic equipment
CN117994722A