Traffic participant perception method and apparatus, electronic device, and storage medium
By extracting target features and uncertainty quality from multimodal sensing data and performing adaptive fusion, the problem of decreased sensing accuracy caused by direct splicing of multi-source data is solved, achieving higher accuracy and robustness in sensing task processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA UNICOM SMART CONNECTION TECH LTD
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-31
AI Technical Summary
In existing traffic environment perception methods, multi-source data are directly spliced or simply fused, which fails to effectively distinguish perception data of different qualities, resulting in a decrease in the accuracy of perception tasks in complex environments.
By acquiring multimodal perception data of traffic participants, target features and uncertainty quality of each modality are extracted, and adaptive fusion is performed based on uncertainty quality to reduce the weight of high uncertainty data and improve the accuracy of fused features.
It achieves dynamic fault tolerance for multimodal sensing data, reduces accuracy degradation caused by single sensor failure or environmental interference, and improves the accuracy and robustness of sensing task processing.
Smart Images

Figure CN122493216A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of intelligent transportation technology, and in particular to methods, devices, electronic devices, and storage media for sensing traffic participants. Background Technology
[0002] In traffic environment perception scenarios, roadside multi-sensor systems are typically used for environmental perception. The acquired multi-source data is then mapped to a bird's-eye view (BEV) space for task processing, achieving a clear spatial structure and facilitating fusion and planning of the environment. In related technologies, BEV perception methods usually involve directly stitching or simply adding and fusing multi-source data, and then performing task detection based on the fused features. However, in actual traffic environment perception scenarios, environmental and weather conditions are often diverse, and the collected multi-source data is easily interfered with. Using it directly without differentiation may reduce the accuracy of data features, thereby affecting the accuracy of the task. Summary of the Invention
[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for sensing traffic participants.
[0004] In a first aspect, this disclosure provides a method for traffic participant perception, the method comprising:
[0005] Acquire multimodal perception data of traffic participants;
[0006] For each modal sensing data in the multimodal sensing data, feature extraction is performed to obtain the target features and uncertainty quality corresponding to each modal sensing data;
[0007] Based on the uncertainty quality corresponding to each modal sensing data, the target features corresponding to each modal sensing data are fused to obtain fused features;
[0008] Based on the fusion features, a perception task is processed to determine the perception result.
[0009] Secondly, this disclosure provides a traffic participant sensing device, the device comprising:
[0010] The acquisition module is used to acquire multimodal perception data of traffic participants;
[0011] The first determining module is used to extract features from each modal sensing data in the multimodal sensing data to obtain the target features and uncertainty quality corresponding to each modal sensing data.
[0012] The fusion module is used to fuse the target features corresponding to each modal sensing data according to the uncertainty quality corresponding to each modal sensing data to obtain fused features;
[0013] The second determining module is used to perform perception task processing based on the fusion features and determine the perception result.
[0014] Thirdly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the traffic participant perception method described above.
[0015] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the traffic participant perception method described above.
[0016] Fifthly, this disclosure provides a computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is executed in a processor of an electronic device, the processor in the electronic device performs the traffic participant perception method described above.
[0017] In this embodiment of the disclosure, multimodal perception data of traffic participants is acquired. The uncertainty quality of each modal perception data can be automatically identified and quantified, and the target features of each modal perception data can be determined. Then, the target features can be adaptively fused based on the uncertainty quality to obtain fused features. Perception task processing is performed based on the fused features to determine the perception result. This can distinguish perception data of different quality, reduce the weight of failed or highly uncertain modal perception data, improve the accuracy of fused features, thereby achieving dynamic fault tolerance of multimodal perception data fusion, reducing the problem of accuracy degradation caused by single sensor failure or environmental interference, and improving the accuracy of perception task processing.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0020] Figure 1A flowchart of a traffic participant perception method provided in this disclosure embodiment;
[0021] Figure 2 This is a system architecture diagram of the traffic participant perception method provided in the embodiments of this disclosure;
[0022] Figure 3 This is a schematic diagram of the logical architecture of the fusion processing in the embodiments of this disclosure;
[0023] Figure 4 A block diagram of a traffic participant sensing device provided in an embodiment of this disclosure;
[0024] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0025] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0027] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0028] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0029] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0030] To facilitate understanding of the embodiments of this disclosure, several concepts will be briefly explained below.
[0031] Bird's-Eye View (BEV): This refers to an intermediate feature representation that maps data from multiple roadside sensors (such as cameras, lidar, millimeter-wave radar, etc.) onto a two-dimensional top-down plane in the coordinate system of traffic participants through geometric transformation or feature projection. It is used to characterize the planar position, scale, and relative spatial relationship of targets in the surrounding environment and is one of the core representations in traffic perception tasks.
[0032] Evidence theory primarily refers to reasoning under uncertainty when information is incomplete or conflicting, such as the Dempster–Shafer evidence theory. By assigning confidence levels to different hypotheses and introducing uncertainty quality, it achieves the fusion of multi-source data and conflict resolution, thereby explicitly characterizing the uncertain parts while preserving valid information.
[0033] In traffic perception scenarios, multiple sensors are typically used for environmental perception, such as roadside cameras, LiDAR, and millimeter-wave radar. The collected multi-source data is mapped into the BEV space for tasks such as target detection. Related technologies often use a simple fusion method that stitches or adds the multi-source data together for task processing, lacking quality assessment and differentiation of the perceived data. However, in real traffic scenarios, the environment and weather are diverse. For example, in rain, fog, backlight, low light at night, and specular reflection, image noise increases significantly, brightness and contrast are distorted, and occlusion and blurring may occur, affecting the feature representation of the image data. Furthermore, if all modal perception data are fused equally, blurry or occluded perception data may actually affect the accuracy of perception, and thus the accuracy of task processing.
[0034] According to the traffic participant perception method of this disclosure, the target features and uncertainty quality corresponding to each modal perception data can be determined. Based on the uncertainty quality, the target features of each modal perception data are fused to obtain fused features. Then, perception task processing is performed based on the fused features to determine the perception result. In this way, the feature quality of the perception data can be displayed and quantified. Adaptive fusion based on uncertainty quality can distinguish the quality and effectiveness of different modal perception data, improve the accuracy of fused features, thereby improving the precision of perception task processing and realizing dynamic fault tolerance and robust output of multimodal perception data fusion.
[0035] The traffic participant perception method according to the embodiments of this disclosure can be executed by electronic devices such as terminal devices or servers. This disclosure does not limit this. The terminal device can be a multi-access edge computing (MEC), a roadside unit, a cloud device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in the memory.
[0036] Figure 1 A flowchart illustrating a traffic participant perception method provided in an embodiment of this disclosure. (Refer to...) Figure 1 The method includes:
[0037] Step S110: Acquire multimodal perception data of traffic participants.
[0038] In this embodiment of the disclosure, multimodal perception data represents multiple different types of perception data. Perception task processing is performed based on multimodal perception data, which can provide a more complete and accurate perception of the traffic environment compared to single perception data.
[0039] Multimodal perception data can be acquired through roadside sensing devices, such as image data collected by roadside cameras with multiple perspectives, or point cloud data collected by roadside lidar or millimeter-wave radar. This disclosure does not impose any limitations on these methods.
[0040] In one possible embodiment, after acquiring multimodal sensing data, time synchronization alignment and spatial coordinate system calibration are performed on each modal sensing data. Specifically, 1) Time dimension: Time synchronization is performed on each modal sensing data corresponding to the same frame scene according to the timestamp of the time service device, so that the time dimension of each modal sensing data remains consistent.
[0041] 2) Spatial Dimension: Based on the sensor's intrinsic parameters (representing the sensor's internal attribute parameters, such as the camera's focal length, distortion, etc.), extrinsic parameters (such as the installation position and angle of the camera and radar, etc.), and the transformation relationship of the sensor's coordinate system, the geometric mapping relationship between the position coordinates of each modal sensing data and the world coordinate system is uniformly modeled to map each modal sensing data into the world coordinate system. This provides a spatial geometric basis for subsequent projection in the BEV space and the generation and alignment of BEV features.
[0042] In the embodiments of this disclosure, the subsequent operations such as determining target features and uncertainty quality are all performed on each modal sensing data after time and space alignment, in order to ensure the accuracy of dimensionality and calculation.
[0043] Step S120: For each modal sensing data in the multimodal sensing data, perform feature extraction to obtain the target features and uncertainty quality corresponding to each modal sensing data.
[0044] During network model inference, not only can prediction results (e.g., category, bounding box, etc.) be obtained, but quantitative indicators used to characterize the data noise level or prediction reliability, such as variance, entropy, and evidence quality, can also be displayed and output, thereby enabling the modeling of the reliability of the prediction results. In this embodiment of the disclosure, when extracting features from each modality of sensing data, target features can be obtained and uncertainty assessment can be performed to obtain uncertainty quality. The larger the uncertainty quality value, the lower the reliability of the sensing data and the higher the probability of it being disturbed.
[0045] Step S130: Based on the uncertainty quality corresponding to each modal sensing data, fuse the target features corresponding to each modal sensing data to obtain fused features.
[0046] Step S140: Based on the fusion features, perform perception task processing and determine the perception result.
[0047] In this embodiment, the uncertainty quality and validity of each modal sensing data can be automatically identified and quantified. Then, adaptive fusion can be performed based on the uncertainty quality to obtain fusion features. Sensing task processing is performed based on the fusion features to determine the sensing result. By distinguishing sensing data of different quality, the weight of failed or highly uncertain modal sensing data can be reduced, improving the accuracy of fusion features. This achieves dynamic fault tolerance for multimodal sensing data fusion, reduces the problem of accuracy degradation caused by single sensor failure or environmental interference, and improves the accuracy of sensing task processing.
[0048] The traffic participant perception method of this disclosure will now be described in detail.
[0049] As mentioned above, the generation of target features in step S120 will be explained using multimodal sensing data, including image data and radar point cloud data, as an example.
[0050] In one possible embodiment, for image data, the target feature corresponding to the image data can be the image BEV feature. Specifically, feature extraction is performed on each modal sensing data in the multimodal sensing data to obtain the target feature corresponding to each modal sensing data, including:
[0051] S1. Input the image data into the image feature extraction network model to obtain a two-dimensional feature map.
[0052] For example, image data is input into an image feature extraction network model to obtain a two-dimensional feature map. C is the number of channels (also known as depth), H is the height (i.e., the vertical dimension of the two-dimensional feature map), and W is the width (i.e., the horizontal dimension of the two-dimensional feature map).
[0053] The image feature extraction network model can be a backbone network, such as ResNet, ConvNeXt, Transformer encoder, etc., and this disclosure does not limit it.
[0054] S2. Based on the two-dimensional feature map, determine the depth variance value of each pixel in the two-dimensional feature map. The depth variance value is used to characterize the geometric uncertainty of the pixel position.
[0055] In this embodiment of the disclosure, a depth prediction branch can be constructed based on the two-dimensional feature map. This depth prediction branch is used to perform parametric regression on each pixel in the two-dimensional feature map to obtain depth distribution information. For example, for any pixel coordinate in the two-dimensional feature map... Obtain the depth mean and depth variance Among them, depth variance can be used to characterize the geometric uncertainty of the pixel's position. The larger the depth variance value, the higher the probability that the pixel may be affected by factors such as raindrops, glare, and blur.
[0056] In another possible embodiment, to ensure the non-negativity of the depth variance value, the logarithmic variance can be obtained first. And the depth variance value is obtained by restoring it through exponential mapping, for example, specifically, .
[0057] S3. Determine the image BEV feature corresponding to the image data based on the depth variance value of each pixel in the two-dimensional feature map.
[0058] In this embodiment of the disclosure, a two-dimensional feature map is projected onto a three-dimensional space, and further aggregated into a BEV grid for the BEV space to obtain image BEV features. During projection and aggregation, a weight can be assigned to each projected sampling point according to the depth variance value of the pixel, which can reduce the interference of inaccurate or low-quality pixels on the image BEV features and improve the accuracy of image BEV feature generation.
[0059] For this step S3, this disclosure provides possible implementation methods, specifically:
[0060] 1) Based on the depth variance value of each pixel in the two-dimensional feature map, determine the confidence weight value of the projection of each pixel in the BEV space. The confidence weight value is negatively correlated with the depth variance value.
[0061] For example, the confidence weight value can be calculated as follows: ,in It is a monotonically decreasing function.
[0062] For example, the confidence weight value can also be calculated as follows: ,in, For a certain pixel (or sampling point). This represents the depth variance value. The confidence weight value is... To control the decay rate through smoothing hyperparameters.
[0063] In this embodiment of the disclosure, the confidence weight value is negatively correlated with the depth variance value, which can weaken the weight of image regions with high depth variance values and reduce their influence in the image BEV features.
[0064] 2) For any query point in the BEV space, determine the reference pixel point sampled in the two-dimensional feature map corresponding to the query point, and the attention weight value of the reference pixel point relative to the query point.
[0065] In this embodiment of the disclosure, a two-dimensional feature map can be projected onto a BEV space. The query point in the BEV space can be understood as a BEV grid point. For each BEV grid, its image features can be calculated by the corresponding pixel point in the two-dimensional feature map during projection. Multiple reference pixels can be sampled in the two-dimensional feature map through attention mechanisms and other methods, and attention weight values obtained through the attention network can be obtained.
[0066] 3) Determine the target weight value of the reference pixel based on the attention weight value and confidence weight value corresponding to the reference pixel.
[0067] For example, for a query point q in the BEV space, multiple reference pixels p corresponding to it are determined in the two-dimensional feature map, and the attention weight values are represented as follows: The confidence weight is Based on the attention weight value and confidence weight value, the target weight value of the reference pixel is obtained after normalization. The target weight value can be calculated as follows: .
[0068] 4) Based on the target weight value of each reference pixel, perform a weighted summation of the image features of each reference pixel to obtain the image features of the query point.
[0069] For example, the image features of query point q are: ,but .
[0070] in, This represents the image features of the reference pixel p.
[0071] 5) Based on the image features of each query point in the BEV space, obtain the image BEV features corresponding to the image data.
[0072] In this embodiment of the disclosure, after traversing each query point in the BEV space, the complete image BEV features can be obtained. For example, the image BEV features can be represented as... .
[0073] In this embodiment of the disclosure, by extracting the depth variance values of the pixels in the image data, a pixel-level uncertainty assessment is constructed. Based on the depth variance values, the image BEV features of the image data are determined. Features with high depth variance values (i.e., features that can be identified as noise or interference features) are automatically suppressed in the weighted summation, thereby realizing the uncertainty assessment of image features. This enables the automatic suppression of interference from low-quality features on image BEV features during projection into the BEV space, reducing BEV projection errors that may be caused by environmental degradation and improving target positioning accuracy.
[0074] In one possible embodiment, the depth variance value is used to evaluate the uncertainty of image data in this embodiment of the disclosure. Explicit regression variance can be used, or a Bayesian neural network or a Monte Carlo Dropout network can be used to obtain the uncertainty by calculating the dispersion of the output distribution through multiple forward inferences. This embodiment of the disclosure does not limit this approach.
[0075] In one possible implementation, for radar point cloud data, the target features corresponding to the point cloud data can be radar BEV features. Specifically, for each modal sensing data in the multimodal sensing data, feature extraction is performed to obtain the target features corresponding to each modal sensing data. This includes: voxelizing the point cloud data and obtaining the point cloud features corresponding to each voxel through feature extraction; and for the BEV space, aggregating the point cloud features corresponding to each voxel to generate the radar BEV features corresponding to the point cloud data.
[0076] For example, features can be extracted from point cloud data using point cloud convolution or sparse convolution, and the point cloud features corresponding to each voxel can be aggregated to generate radar BEV features with the same BEV feature dimension as the image. .
[0077] Regarding the generation of uncertainty quality in step S120 above, this disclosure also provides possible implementation methods. Specifically, for each modal sensing data in the multimodal sensing data, feature extraction is performed to obtain the uncertainty quality corresponding to each modal sensing data, which may include:
[0078] 1) For each modal sensing data in the multimodal sensing data, input any modal sensing data into the corresponding evidence network model, encode any modal sensing data to obtain a K-dimensional evidence vector, which includes K evidence values, where K is an integer greater than 1.
[0079] In this embodiment of the disclosure, for each modal sensing data, a corresponding lightweight evidence network model can be constructed. This evidence network model is then used for encoding and activation function processing to output a non-negative evidence vector. For example, for any modal sensing data m, its evidence vector can be represented as... .
[0080] Here, K can be understood as the number of categories or the number of hypothetical states in the classification task. This represents the evidence value of modality m for the k-th category or state.
[0081] For example, K can characterize the number of target categories in traffic environment perception, such as motor vehicles, pedestrians, non-motor vehicles, road guardrails, traffic signs, and other target categories that BEV perception needs to identify. As another example, K can characterize the number of working states of a sensor, such as the states of a camera or radar, such as "normal operation, slight degradation, severe degradation, and complete failure". It can be used for the quantitative evaluation of sensor modal reliability.
[0082] 2) Determine the total evidence value based on each evidence value in the evidence vector, and determine the uncertainty quality of any modal sensing data based on the ratio of K to the total evidence value. The uncertainty quality is negatively correlated with the total evidence value.
[0083] In this embodiment of the disclosure, each evidence value in the evidence vector can be summed and combined with K to determine the total evidence value. For example, the total evidence value can be expressed as: , .
[0084] Uncertainty quality can be expressed as Therefore, the uncertainty quality can be calculated as follows: .
[0085] in, The closer the uncertainty quality value is to 1, the more unreliable and lower the quality of the modal sensing data. For example, if the camera is severely obstructed, the effective evidence vector extracted from the image data obtained by the camera tends to 0, and the uncertainty quality obtained at this time tends to 1. In this way, the quality of the sensing data can be explicitly mathematically quantified.
[0086] In one possible embodiment, regarding step S130 above, which involves fusing the target features corresponding to each modal sensing data based on the uncertainty quality corresponding to each modal sensing data to obtain fused features, this disclosure provides possible implementation methods:
[0087] Based on the uncertainty quality corresponding to each modal sensing data, a fusion weight value corresponding to each modal sensing data is determined, wherein the fusion weight value is negatively correlated with the corresponding uncertainty quality; the target feature corresponding to each modal sensing data is multiplied by its corresponding fusion weight value to obtain multiple products, and the multiple products are summed to obtain the fusion feature.
[0088] For example, taking multimodal sensing data, which includes image data and point cloud data, as an example, the uncertain quality of image data is represented as: The uncertainty quality of point cloud data is represented as The BEV features of the image are Radar BEV characteristics are .
[0089] Will The input consists of a gated network comprising a multilayer perceptron (MLP) and activation functions, which generates image channel coefficient vectors that match the feature channel dimension C. and radar channel coefficient vector Its calculation method can be expressed as:
[0090]
[0091]
[0092] To ensure the stability of the feature scale after fusion, the coefficients obtained above can be normalized to obtain the fusion weight values corresponding to each modality of sensing data. The specific calculation method can be expressed as follows:
[0093]
[0094]
[0095] in, This is a preset parameter, and its value is a positive number greater than 0. This can avoid the denominator being 0 during normalization.
[0096] Then, based on the fusion weight values corresponding to the image data The fusion weight value of point cloud data The image BEV features and radar BEV features are weighted and summed to obtain the fused features. For example, the fused features can be represented as follows: : .
[0097] In this embodiment of the disclosure, target feature fusion is performed in combination with uncertainty quality. For target features with high uncertainty quality, their fusion weight value can be reduced to improve the accuracy and quality of fused features, thereby achieving adaptive fusion and fault tolerance in modal failure scenarios.
[0098] In this embodiment, environmental perception and traffic participant tracking can be performed in real time at intersections and road segments. Typically, detection can be performed frame by frame, followed by cross-frame association. However, related technologies usually rely only on the estimation results of single-frame detection. But in scenarios where traffic participants, such as motor vehicles traveling at high speeds or making sharp turns, are involved, the detected target may be occluded or temporarily disappear. Therefore, the single-frame detection results may have large deviations or false detections. In cross-frame association tracking, trajectory misassociation may also occur. Therefore, this embodiment also provides a modeling and utilization mechanism for position uncertainty. This allows the confidence level of the target position estimation to be explicitly considered when performing state updates and trajectory association during tracking processing. Position alignment is performed on multiple frames, and a smoother update strategy is adopted for perception data with high uncertainty observations. This can improve the temporal consistency and stability of dynamic target tracking and increase accuracy.
[0099] Specifically, for multi-frame perception task processing, this disclosure also provides possible implementation methods, including acquiring multimodal perception data collected in each frame or at each time step, extracting fusion features from the multimodal perception data in each frame or at each time step, and then aligning the fusion features of multiple frames for decoding to obtain the perception result. Therefore, the step S140 above, which involves processing the perception task based on the fusion features to determine the perception result, may include:
[0100] S141: Obtain motion estimation information of traffic participants, and determine the pose transformation matrix of traffic participants based on the motion estimation information.
[0101] For example, motion estimation information of traffic participants may include speed, heading angle, attitude angle, acceleration, mileage data, etc. This disclosure does not impose any limitations. The pose transformation matrix from the previous frame to the current frame can be calculated based on the motion estimation information corresponding to the previous frame (time t-1) and the current frame (time t).
[0102] S142: Based on the fusion features of the previous frame and the pose transformation matrix, map the fusion features of the previous frame to the target coordinate system of the current frame to obtain the mapped fusion features of the previous frame.
[0103] For example, the fusion features of the previous frame are Based on the pose transformation matrix, a spatial affine transformation is performed on the fused features of the previous frame, mapping them to the coordinate system of the traffic participants at the current time t, resulting in the coarsely aligned fused features of the previous frame after mapping. .
[0104] S143: Jointly process the mapped fusion features of the previous frame and the fusion features of the current frame to obtain the residual offset value.
[0105] In this embodiment of the disclosure, considering that alignment using only the pose transformation matrix may produce errors due to the motion of the dynamic target itself, this embodiment of the disclosure can also estimate the target's displacement between frames to further correct the displacement error caused by motion.
[0106] For example, a convolutional network can be used to fuse features from the previous frame after mapping. Features fused with the current frame Joint processing is performed to obtain the residual offset value in the BEV space. The residual offset value can be understood as a displacement parameter at the pixel or BEV grid level, which can characterize the deviation between the historical fusion features after coarse alignment and the actual position.
[0107] S144: Based on the residual offset value, adjust the sampling position corresponding to the fusion feature of the previous frame so that the sampling position corresponding to the fusion feature of the previous frame is aligned with the sampling position of the current frame in the BEV space, and obtain the aligned multi-frame fusion feature sequence, which includes the aligned fusion feature of the previous frame and the fusion feature of the current frame.
[0108] For example, the residual offset value is input into the deformable convolutional layer to adaptively adjust the sampling position of the coarsely aligned historical fusion features, so as to complete the second precise temporal alignment and obtain the temporally aligned multi-frame fusion feature sequence.
[0109] S145: Input the multi-frame fused feature sequence into the multi-sensory task detection network for decoding to obtain the perception results corresponding to multiple perception tasks. The perception results are used for downstream task decision-making.
[0110] In this embodiment of the disclosure, a multi-branch detection head can be constructed for decoding outputs of different perception tasks. For example, the multi-branch detection head includes a target detection branch and a segmentation or occupancy prediction branch. Decoding can be performed through the target detection branch to output the target's three-dimensional bounding box (including center coordinates, size, rotation angle, etc.) and category. Decoding can be performed through the segmentation or occupancy prediction branch to output lane line segmentation masks, BEV space occupancy network probabilities, etc. The output results of each branch are then integrated to obtain the final multimodal perception result for use in downstream task decision-making. Downstream tasks may include behavior prediction, decision scale, etc., which are not limited in this embodiment of the disclosure.
[0111] The following explanation uses a specific application scenario, taking multimodal sensing data, including image data and radar point cloud data, as an example. (See attached document.) Figure 2 The diagram shown is a system architecture diagram of the traffic participant perception method provided in this embodiment of the disclosure. Figure 2 As shown, the system architecture includes an input and synchronization module, a feature extraction module, a fusion module, and a post-processing output module.
[0112] 1) Input and synchronization module.
[0113] In this embodiment of the disclosure, the input module is mainly used to acquire multimodal sensing data collected by multiple sensors, and to perform time synchronization and coordinate system calibration (i.e., spatial dimension alignment) on the multimodal sensing data.
[0114] like Figure 2 As shown, it can receive image data from cameras at multiple perspectives on the road and point cloud data from radar, and then perform timestamp alignment and coordinate system calibration on the image data and point cloud data.
[0115] 2) Feature extraction module.
[0116] like Figure 2 As shown in this embodiment, in addition to conventional feature extraction, an uncertainty feature extraction branch is added to the feature extraction module, which can evaluate the uncertainty of image depth estimation and the quality of each modal perception data.
[0117] like Figure 2As shown, for image data processing, using temporally and spatially aligned image data as input, a two-dimensional feature map is extracted. Based on the two-dimensional feature map, uncertainty is evaluated for each pixel or feature location in the feature map, determining the depth variance and depth mean for each pixel or feature location. The depth variance can be used to characterize geometric uncertainty. Furthermore, for the image data, the uncertainty quality is determined through its corresponding evidence network model.
[0118] For radar point cloud data processing, feature extraction is performed using temporally and spatially aligned point cloud data as input to obtain point cloud features. Furthermore, the uncertainty quality of the point cloud data is determined through its corresponding evidence network model.
[0119] 3) Integration module.
[0120] In this embodiment of the disclosure, the fusion module is mainly used to spatially aggregate and modally fuse image features and point cloud features in the BEV space, and based on the depth variance value obtained by the feature extraction model and the weight allocation of the uncertainty quality control fusion, it can realize the differentiation and construction of perceived data quality.
[0121] For image data, using the two-dimensional feature map of the image data and the depth variance of the pixels as input, the two-dimensional feature map can be projected onto the BEV space based on the depth variance value to determine the image BEV features corresponding to the image data.
[0122] For point cloud data, the point cloud features corresponding to each voxel can be aggregated and processed in the BEV space to generate radar BEV features.
[0123] Therefore, in this embodiment of the disclosure, image BEV features and radar BEV features can be fused by combining uncertainty quality to obtain fused features.
[0124] See Figure 3 The diagram shown is a schematic representation of the logical architecture of the fusion processing in an embodiment of this disclosure. Figure 3 As shown, based on the uncertainty quality I corresponding to the image data, the fusion weight value I corresponding to the image data is determined, and the fusion weight value I is multiplied with the image BEV feature to obtain the weighted image BEV feature.
[0125] Based on the uncertainty quality L corresponding to the point cloud data, the fusion weight value L corresponding to the point cloud data is determined, and the fusion weight value L is multiplied with the radar BEV feature to obtain the weighted radar BEV feature.
[0126] Then, the weighted image BEV features and the weighted radar BEV features are summed to obtain the final fused features.
[0127] 4) Post-processing output module.
[0128] In this embodiment of the disclosure, the post-processing output module is mainly used to decode the fused features and output various perception results.
[0129] In this embodiment of the disclosure, the fusion features of multiple frames can be temporally aligned, and a multi-branch detection head can be set to decode the temporally aligned fusion features of multiple frames through the multi-branch detection head to obtain the perception results corresponding to multiple perception tasks. For example, the perception results may include three-dimensional bounding box coordinates, target category, lane line segmentation, occupancy grid, etc. This embodiment of the disclosure does not impose any restrictions on these, and then the results are integrated and output.
[0130] In this embodiment, uncertainty assessment is performed on multimodal sensing data to determine the uncertainty quality and depth variance of image data. This allows for quantitative analysis of the uncertainty of multimodal sensing data. Furthermore, based on the uncertainty quality, weighted fusion of multimodal sensing data is performed to obtain fusion features. This enables pixel-level modal complementarity at the feature channel level. In cases where a sensor fails or the quality of the acquired sensing data is low, the impact on sensing is reduced. Thus, sensing task processing is performed based on the fusion features, improving sensing accuracy.
[0131] In one possible embodiment, based on the above embodiments, the present disclosure mainly involves fusing the target features corresponding to the multimodal perception data and then decoding based on the fused features to obtain the perception result. Alternatively, in another embodiment, multiple perception results can be obtained by first decoding the target features corresponding to the multimodal perception results respectively, and then the multiple perception results can be fused to obtain the final perception result.
[0132] In one possible embodiment, the models involved in the embodiments of this disclosure can be obtained through pre-training. For example, multiple models can be set up and trained separately, such as an image feature extraction network model for image feature extraction, a model for point cloud data feature extraction, and an evidence network model for uncertainty quality extraction, etc., which are trained separately. Alternatively, multiple models can be trained jointly, or they can be trained jointly as modules of an overall perception model. This training can enable the perception model to have uncertainty assessment capabilities, improving its robustness and reliability in complex scenarios.
[0133] For example, taking multimodal perception data, including image data and radar point cloud data, as an example, this disclosure provides a possible model training method.
[0134] 1) Obtain the training dataset, which includes image data and point cloud data, and perform time and space alignment preprocessing on the image data and point cloud data.
[0135] By randomly setting the image data or point cloud data in the training dataset to an all-zero matrix (to simulate the case of complete sensor failure) or adding noise (to simulate the case of sensor degradation) with a preset probability, the model can be trained based on the processed training dataset. This can enhance the model's ability to automatically adjust when a certain modality of sensing data is missing or when a certain modality of sensing data is interfered with.
[0136] 2) Model forward inference.
[0137] The above-processed training dataset is input into the model, and the reasoning processes involved in the traffic participant perception method in this embodiment of the present disclosure, such as uncertainty quality extraction, target feature extraction, fusion processing, and decoding processing, are executed in sequence to obtain the model's predicted perception result output, as well as intermediate outputs such as uncertainty quality, depth variance value, and evidence vector corresponding to each modality perception data.
[0138] 3) Calculation of joint loss function.
[0139] Based on the predicted perception results output by the model and the labeled real perception results, the total loss value is calculated. Specifically, it can include the loss values of multiple optimization objectives to achieve multi-objective optimization.
[0140] For example, it can include task loss, depth variance regularization loss, and uncertainty loss of the evidence network. The total loss value can then be expressed as: , .
[0141] in, The task loss value can be represented by methods such as Focal Loss and L1 Loss, which measure the deviation between the model's predicted perception results and the actual perception results. The depth variance regularization loss value can be represented by the weights. Adjustments can constrain the predictive validity of depth variance; The uncertainty loss value representing the evidence network can be represented by the evidence cross-entropy loss based on the Dirichlet distribution. This loss penalizes the model for assigning high confidence to incorrect categories, thereby guiding the model to output high uncertainty quality for degraded data. This can be achieved through weighting. Adjustments were made.
[0142] 4) Parameter backpropagation and model optimization.
[0143] Based on the calculated total loss value, backpropagation is performed to each layer of the entire model network. Through optimization algorithms such as gradient descent, the values of the model's trainable parameters are updated, such as the trainable parameters involved in the deep variance prediction branch, evidence network, and fusion module.
[0144] 5) Closed-loop training and convergence.
[0145] Through continuous iterative training, until the model's total loss function converges or the preset iteration termination condition is met, the training of the entire model is completed, and a final model that can be deployed and used is obtained.
[0146] The preset iteration termination condition can be that the sensing accuracy and anti-failure capability reach a preset index threshold, or it can be that a preset number of iterations is reached, etc., and this embodiment does not impose any restrictions.
[0147] In this embodiment of the disclosure, the multimodal sensing data in the training dataset is randomly subjected to failure or noise processing to simulate sensor failure and degradation scenarios, thereby iteratively training the model to improve the model training accuracy and efficiency. In addition, in this embodiment of the disclosure, the uncertainty index output by the model can be used to screen out high uncertainty samples (i.e., poor quality samples) as key samples. Only high uncertainty samples are manually labeled, and retraining can be performed based on these labeled samples, thereby reducing the cost of manual labeling and improving the model's resistance to failure and its generalization ability for complex scenarios.
[0148] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0149] In addition, this disclosure also provides traffic participant sensing devices, electronic devices, computer-readable storage media, and computer program products, all of which can be used to implement any of the traffic participant sensing methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding descriptions in the method section and will not be repeated here.
[0150] Figure 4 This is a block diagram of a traffic participant sensing device provided in an embodiment of the present disclosure.
[0151] Reference Figure 4 This disclosure provides a traffic participant sensing device, which includes:
[0152] Acquisition module 41 is used to acquire multimodal perception data of traffic participants;
[0153] The first determining module 42 is used to extract features from each modal sensing data in the multimodal sensing data to obtain the target features and uncertainty quality corresponding to each modal sensing data.
[0154] The fusion module 43 is used to fuse the target features corresponding to each modal sensing data according to the uncertainty quality corresponding to each modal sensing data to obtain fused features;
[0155] The second determining module 44 is used to perform perception task processing based on the fusion features and determine the perception result.
[0156] In one possible embodiment, the multimodal sensing data includes image data, and the target features include BEV features from a bird's-eye view of the image; when performing feature extraction on each modal sensing data in the multimodal sensing data to obtain the target features corresponding to each modal sensing data, the first determining module 42 is used to:
[0157] The image data is input into an image feature extraction network model to obtain a two-dimensional feature map;
[0158] Based on the two-dimensional feature map, the depth variance value of each pixel in the two-dimensional feature map is determined, and the depth variance value is used to characterize the geometric uncertainty of the pixel position;
[0159] Based on the depth variance value of each pixel in the two-dimensional feature map, the image BEV feature corresponding to the image data is determined.
[0160] In one possible embodiment, when determining the image BEV feature corresponding to the image data based on the depth variance value of each pixel in the two-dimensional feature map, the first determining module 42 is used to:
[0161] Based on the depth variance value of each pixel in the two-dimensional feature map, a confidence weight value for the projection of each pixel in the BEV space is determined, and the confidence weight value is negatively correlated with the depth variance value.
[0162] For any query point in the BEV space, determine the reference pixel point sampled by the query point in the two-dimensional feature map, and the attention weight value of the reference pixel point relative to the query point;
[0163] The target weight value of the reference pixel is determined based on the attention weight value and confidence weight value corresponding to the reference pixel.
[0164] Based on the target weight value of each reference pixel, the image features of each reference pixel are weighted and summed to obtain the image features of the query point.
[0165] Based on the image features of each query point in the BEV space, the image BEV features corresponding to the image data are obtained.
[0166] In one possible embodiment, the multimodal sensing data includes radar point cloud data, and the target features include radar BEV features; when performing feature extraction on each modal sensing data in the multimodal sensing data to obtain the target features corresponding to each modal sensing data, the first determining module 42 is used to:
[0167] The point cloud data is voxelized, and the point cloud features corresponding to each voxel are obtained through feature extraction.
[0168] For the BEV space, the point cloud features corresponding to each voxel are aggregated to generate the radar BEV features corresponding to the point cloud data.
[0169] In one possible embodiment, when performing feature extraction on each modal sensing data in the multimodal sensing data to obtain the uncertainty quality corresponding to each modal sensing data, the first determining module 42 is used to:
[0170] For each modal sensing data in the multimodal sensing data, any modal sensing data is input into the corresponding evidence network model, and the any modal sensing data is encoded to obtain a K-dimensional evidence vector, wherein the evidence vector includes K evidence values, and K is an integer greater than 1;
[0171] Based on each evidence value in the evidence vector, a total evidence value is determined, and based on the ratio of K to the total evidence value, the uncertainty quality of any modal sensing data is determined, wherein the uncertainty quality is negatively correlated with the total evidence value.
[0172] In one possible embodiment, when fusing the target features corresponding to each modal sensing data according to the uncertainty quality corresponding to each modal sensing data to obtain fused features, the fusion module 43 is used to:
[0173] Based on the uncertainty quality corresponding to each modal sensing data, a fusion weight value corresponding to each modal sensing data is determined, wherein the fusion weight value is negatively correlated with the corresponding uncertainty quality;
[0174] The target feature corresponding to each modality sensing data is multiplied by its corresponding fusion weight value to obtain multiple products, and the multiple products are summed to obtain the fusion feature.
[0175] In one possible embodiment, the multimodal sensing data is acquired in the current frame, the fusion feature represents the fusion feature of the current frame, and when performing sensing task processing based on the fusion feature to determine the sensing result, the second determining module 44 is used to:
[0176] Obtain motion estimation information of the traffic participants, and determine the pose transformation matrix of the traffic participants based on the motion estimation information;
[0177] Based on the fusion features of the previous frame and the pose transformation matrix, the fusion features of the previous frame are mapped to the target coordinate system of the current frame to obtain the mapped fusion features of the previous frame.
[0178] The mapped fused features of the previous frame and the fused features of the current frame are jointly processed to obtain the residual offset value;
[0179] Based on the residual offset value, the sampling position corresponding to the previous frame fusion feature is adjusted so that the sampling position corresponding to the previous frame fusion feature is aligned with the current frame sampling position in the BEV space, thereby obtaining an aligned multi-frame fusion feature sequence, which includes the aligned previous frame fusion feature and the current frame fusion feature.
[0180] The multi-frame fused feature sequence is input into a multi-sensory task detection network for decoding to obtain the perception results corresponding to multiple perception tasks. The perception results are used for downstream task decision-making.
[0181] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0182] Reference Figure 5 This disclosure provides an electronic device comprising: at least one processor 501; at least one memory 502; and one or more I / O interfaces 503; wherein the memory 502 stores one or more computer programs executable by the at least one processor 501, the one or more computer programs being executed by the at least one processor 501 to enable the at least one processor 501 to perform the traffic participant perception method described above.
[0183] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the traffic participant perception method described above. The computer-readable storage medium may be volatile or non-volatile.
[0184] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the above-described traffic participant perception method.
[0185] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0186] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0187] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0188] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0189] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0190] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0191] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0192] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0193] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0194] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A traffic participant perception method, characterized in that, include: Acquire multimodal perception data of traffic participants; For each modal sensing data in the multimodal sensing data, feature extraction is performed to obtain the target features and uncertainty quality corresponding to each modal sensing data; Based on the uncertainty quality corresponding to each modal sensing data, the target features corresponding to each modal sensing data are fused to obtain fused features; Based on the fusion features, a perception task is processed to determine the perception result.
2. The method of claim 1, wherein, The multimodal perception data includes image data, and the target features include BEV features from a bird's-eye view of the image; the step of extracting features from each modal perception data in the multimodal perception data to obtain the target features corresponding to each modal perception data includes: The image data is input into an image feature extraction network model to obtain a two-dimensional feature map; Based on the two-dimensional feature map, the depth variance value of each pixel in the two-dimensional feature map is determined, and the depth variance value is used to characterize the geometric uncertainty of the pixel position; Based on the depth variance value of each pixel in the two-dimensional feature map, the image BEV feature corresponding to the image data is determined.
3. The method according to claim 2, characterized in that, The step of determining the image BEV feature corresponding to the image data based on the depth variance value of each pixel in the two-dimensional feature map includes: Based on the depth variance value of each pixel in the two-dimensional feature map, a confidence weight value for the projection of each pixel in the BEV space is determined, and the confidence weight value is negatively correlated with the depth variance value. For any query point in the BEV space, determine the reference pixel point sampled by the query point in the two-dimensional feature map, and the attention weight value of the reference pixel point relative to the query point; The target weight value of the reference pixel is determined based on the attention weight value and confidence weight value corresponding to the reference pixel. Based on the target weight value of each reference pixel, the image features of each reference pixel are weighted and summed to obtain the image features of the query point. Based on the image features of each query point in the BEV space, the image BEV features corresponding to the image data are obtained.
4. The method according to claim 1, characterized in that, The multimodal sensing data includes radar point cloud data, and the target features include radar BEV features; the step of extracting features from each modal sensing data in the multimodal sensing data to obtain the target features corresponding to each modal sensing data includes: The point cloud data is voxelized, and the point cloud features corresponding to each voxel are obtained through feature extraction. For the BEV space, the point cloud features corresponding to each voxel are aggregated to generate the radar BEV features corresponding to the point cloud data.
5. The method according to claim 1, characterized in that, The step of extracting features from each modal sensing data in the multimodal sensing data to obtain the uncertainty quality corresponding to each modal sensing data includes: For each modal sensing data in the multimodal sensing data, any modal sensing data is input into the corresponding evidence network model, and the any modal sensing data is encoded to obtain a K-dimensional evidence vector, wherein the evidence vector includes K evidence values, and K is an integer greater than 1; Based on each evidence value in the evidence vector, a total evidence value is determined, and based on the ratio of K to the total evidence value, the uncertainty quality of any modal sensing data is determined, wherein the uncertainty quality is negatively correlated with the total evidence value.
6. The method according to any one of claims 1-5, characterized in that, The step of fusing the target features corresponding to each modal sensing data according to the uncertainty quality corresponding to each modal sensing data to obtain fused features includes: Based on the uncertainty quality corresponding to each modal sensing data, a fusion weight value corresponding to each modal sensing data is determined, wherein the fusion weight value is negatively correlated with the corresponding uncertainty quality; The target feature corresponding to each modality sensing data is multiplied by its corresponding fusion weight value to obtain multiple products, and the multiple products are summed to obtain the fusion feature.
7. The method according to claim 1, characterized in that, The multimodal sensing data is acquired in the current frame, the fusion feature represents the fusion feature of the current frame, and the sensing task processing based on the fusion feature to determine the sensing result includes: Obtain motion estimation information of the traffic participants, and determine the pose transformation matrix of the traffic participants based on the motion estimation information; Based on the fusion features of the previous frame and the pose transformation matrix, the fusion features of the previous frame are mapped to the target coordinate system of the current frame to obtain the mapped fusion features of the previous frame. The mapped fused features of the previous frame and the fused features of the current frame are jointly processed to obtain the residual offset value; Based on the residual offset value, the sampling position corresponding to the previous frame fusion feature is adjusted so that the sampling position corresponding to the previous frame fusion feature is aligned with the current frame sampling position in the BEV space, thereby obtaining an aligned multi-frame fusion feature sequence, which includes the aligned previous frame fusion feature and the current frame fusion feature. The multi-frame fused feature sequence is input into a multi-sensory task detection network for decoding to obtain the perception results corresponding to multiple perception tasks. The perception results are used for downstream task decision-making.
8. A traffic participant sensing device, characterized in that, include: The acquisition module is used to acquire multimodal perception data of traffic participants; The first determining module is used to extract features from each modal sensing data in the multimodal sensing data to obtain the target features and uncertainty quality corresponding to each modal sensing data. The fusion module is used to fuse the target features corresponding to each modal sensing data according to the uncertainty quality corresponding to each modal sensing data to obtain fused features; The second determining module is used to perform perception task processing based on the fusion features and determine the perception result.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor to enable the at least one processor to perform the traffic participant perception method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the traffic participant perception method as described in any one of claims 1-7.