Multimodal fusion method, apparatus, and related device

CN122548607APending Publication Date: 2026-08-11BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]本申请实施例提供一种多模态融合方法、装置及相关设备,能够解决相关技术中环境感知精度不足、鲁棒性不高和动态场景下检测性能受限的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548607A_ABST
    Figure CN122548607A_ABST
Patent Text Reader

Abstract

This application discloses a multimodal fusion method, apparatus, and related equipment, belonging to the field of vehicle technology. The multimodal fusion method includes: acquiring radar data and visual data; performing spatiotemporal consistency processing and weighted fusion on the radar data and the visual data to obtain fused features. This application achieves deep multimodal fusion and is applicable to intelligent driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of vehicles, and in particular relates to a multimodal fusion method, apparatus and related equipment. Background Technology

[0002] Autonomous driving systems are placing increasingly higher demands on the accuracy, robustness, and real-time performance of environmental perception. Visual sensors provide rich semantic information, but their performance is unstable under complex lighting conditions; millimeter-wave radar performs excellently in adverse weather conditions, but its resolution and ability to detect non-metallic targets are limited. Most multimodal fusion methods employ simple cascading or weighting strategies, which suffer from insufficient environmental perception accuracy, low robustness, and limited detection performance in dynamic scenes. Summary of the Invention

[0003] This application provides a multimodal fusion method, apparatus, and related equipment, which can solve the problems of insufficient environmental perception accuracy, low robustness, and limited detection performance in dynamic scenes in related technologies.

[0004] Firstly, a multimodal fusion method is provided, executed by a vehicle, the method comprising: Acquire radar and visual data; The radar data and the visual data are subjected to spatiotemporal consistency processing and weighted fusion to obtain fused features.

[0005] In this embodiment, by combining the high-precision physical perception capability of radar with the rich semantic information of the vision module for multimodal feature fusion, the global perception capability and feature representation richness are enhanced. Spatiotemporal consistency processing of radar and vision data solves the spatiotemporal bias problem introduced by sampling differences among multiple sensors. Weighted fusion of the spatiotemporally consistent radar and vision data yields fused features. By combining millimeter-wave radar and vision sensors, and utilizing spatiotemporal consistency processing and weighted fusion techniques, the system's detection accuracy, robustness, and adaptability in fast-moving and complex dynamic scenarios are effectively improved, reliably supporting environmental perception for autonomous vehicles.

[0006] In some embodiments, the spatiotemporal consistency processing and weighted fusion of the radar data and the visual data to obtain fused features includes: The radar data and the visual data are subjected to spatiotemporal consistency processing to determine a dynamic target set and a first feature set, wherein the first feature set includes a spatiotemporally aligned first radar feature set and a first visual feature set; Based on the set of dynamic targets, the first feature set is updated, and the current dynamic scene and the fusion weights under the current dynamic scene are determined; Using an attention mechanism, the first radar feature set and the first visual feature set in the updated first feature set are weighted and fused according to the fusion weights to obtain fused features. The updated first feature set is then fused and filtered to obtain the second feature set.

[0007] In this embodiment, spatiotemporal consistency processing is performed on radar and visual data, including time alignment, spatial alignment, and time consistency processing, solving the spatiotemporal deviation problem introduced by sampling differences among multiple sensors. Based on the spatiotemporally consistent radar and visual data, an updated first feature set is determined, which includes dynamic attributes and the spatiotemporally consistent radar and visual data. According to the current dynamic scene, the fusion weights of millimeter-wave radar and visual features are dynamically adjusted to achieve optimized fusion of target features. Through an attention mechanism, feature fusion is performed on the updated first feature set (including a first radar feature set and a first visual feature set), fully utilizing the complementary characteristics between radar and visual data to enhance global perception capabilities and feature richness. By combining millimeter-wave radar and visual sensors, and utilizing spatiotemporal consistency processing, dynamic attribute adaptive weighting, and multimodal fusion technology, the detection accuracy, robustness, and adaptability of the system in fast-moving and complex dynamic scenes are effectively improved, reliably supporting environmental perception for autonomous vehicles.

[0008] In some embodiments, updating the first feature set based on the dynamic target object set and determining the current dynamic scene and the fusion weights under the current dynamic scene includes: Based on the set of dynamic target objects and the first feature set, a set of dynamic attributes is determined, and the first feature set is updated according to the set of dynamic attributes. Based on the set of dynamic target objects, the current dynamic scene is determined, and the fusion weight under the current dynamic scene is determined according to the set of dynamic attributes.

[0009] In this embodiment of the application, dynamic attribute adaptive weighting is used, wherein the fusion weight of millimeter-wave radar and visual features can be dynamically adjusted according to the target motion trend. After subsequent feature fusion, the target feature optimization fusion for different motion states can be achieved.

[0010] In some embodiments, performing spatiotemporal consistency processing on the radar data and the visual data to determine the dynamic target set and the first feature set includes: Using a time interpolation algorithm, a Kalman filter is used to align the radar data and the visual data in time. By using extrinsic parameter calibration technology, the time-aligned radar data and time-aligned visual data are unified into the vehicle coordinate system to obtain spatiotemporally aligned radar data and spatiotemporally aligned visual data. The visual feature points corresponding to the spatiotemporally aligned radar data and the spatiotemporally aligned visual data are determined, and the motion vectors of the visual feature points are calculated using the optical flow method to determine the dynamic target set. The spatiotemporally aligned radar data and spatiotemporally aligned visual data of multiple consecutive frames are input into the spatiotemporal encoder, and time consistency processing is performed to obtain the first radar feature set and the first visual feature set. Wherein, the first radar feature set is the radar data with spatiotemporal consistency, and the first visual feature set is the visual data with spatiotemporal consistency.

[0011] In this embodiment, spatiotemporal consistency processing is performed on radar and visual data. This processing includes time alignment, spatial alignment, and time consistency processing to address the spatiotemporal bias caused by sampling differences between multiple sensors and improve detection accuracy in dynamic scenes. Specifically, a Kalman filter is used for time alignment to correct biases and avoid distortion in dynamic attribute calculations. In spatial alignment, targets around the vehicle are determined to be either static or dynamic, identifying the set of dynamic targets in the radar and visual data to calculate their dynamic attributes. Subsequently, the fusion weights can be adjusted based on the target motion trends corresponding to the dynamic attributes, achieving optimized fusion of target features for different motion states.

[0012] In some embodiments, the step of using a time interpolation algorithm and a Kalman filter to perform time alignment of the radar data and the visual data includes: Based on the sampling frequencies of the radar data and the visual data, a unified set of timestamps for interpolation is determined; The radar data and the visual data are time-aligned by interpolation, wherein the radar data is real-time data or legacy data, and the visual data is real-time data or legacy data. If the current timestamp belongs to the set of timestamps, use a Kalman filter to predict the estimated value of the radar data and the estimated value of the visual data corresponding to the current timestamp, and update the residual data in the radar data to the estimated value of the radar data, and update the residual data in the visual data to the estimated value of the visual data.

[0013] In this embodiment, a time interpolation algorithm is used to perform time alignment of the radar data and the visual data using a Kalman filter, so as to ensure that the time alignment error is less than or equal to a preset threshold, thereby avoiding the radar and visual observations of the same target object corresponding to different states at different times, and further avoiding the distortion of subsequent dynamic attribute calculations.

[0014] In some embodiments, the dynamic attributes corresponding to the first radar feature set and the first visual feature set each include at least one of the following: basic motion attributes of multiple dynamic targets, spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes; the spatiotemporal correlation attributes include: motion trajectory fitting degree and / or life cycle state; the perception enhancement attributes include: multi-sensor consistency and / or dynamic confidence. The determination of the dynamic attribute set based on the dynamic target set and the first feature set includes at least one of the following: Based on the set of dynamic targets, the basic motion attributes in the set of dynamic attributes are determined according to the difference between the current frame data and the previous frame data in the first feature set. Based on the persistent state of the target object in the first feature set of the spatiotemporal encoder, the life cycle state in the dynamic attribute set is determined. By comparing the motion trajectories in the first radar feature set and the first visual feature set, the fitting degree of the motion trajectory in the dynamic attribute set is determined. Based on the first feature set, determine the interaction behavior between the vehicle and surrounding dynamic targets, and determine the interaction behavior features in the dynamic attribute set; The position information of the same target in the first radar feature set and the first visual feature set are compared, and the consistency of the multiple sensors in the dynamic attribute set is determined based on the deviation of the position information. The dynamic confidence level in the dynamic attribute set is determined based on the multi-sensor consistency and the stable state of the first feature set over multiple consecutive frames.

[0015] In some embodiments, the fusion weights include radar weights and visual weights; The step of determining the current dynamic scene based on the set of dynamic targets and determining the fusion weights under the current dynamic scene according to the set of dynamic attributes includes: Based on the set of dynamic targets, the density of dynamic targets around the vehicle is calculated, and based on the density of the dynamic targets, the current dynamic scene is determined; The spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes in the dynamic attributes corresponding to the first radar feature set and the first visual feature set are weighted and summed according to a preset ratio to determine the radar weight and the visual weight, respectively. The preset ratio is related to the current dynamic scene. Based on the basic motion attributes in the dynamic attributes corresponding to the first radar feature set and the first visual feature set, the interaction parameters between the vehicle and the surrounding dynamic targets are determined; if the interaction parameters corresponding to the first radar feature set are outside the safe threshold range, or if the interaction parameters corresponding to the first visual feature set are outside the safe threshold range, the radar weights and the visual weights are updated.

[0016] In this embodiment, the current dynamic scene is determined to be either a conventional scene or a high-interaction-density scene based on the set of dynamic targets. Using a preset ratio corresponding to the current dynamic scene, the spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes in the first and second dynamic attributes are weighted and summed to obtain radar weights and visual weights, respectively. Based on the basic motion attributes in the first and second dynamic attributes, it can be determined whether the interaction parameters (e.g., collision time) corresponding to the first and second dynamic attributes are outside the safety threshold range, and the radar weights and visual weights are updated accordingly. This achieves dynamic adjustment of the fusion weights based on dynamic attributes, meaning the fusion weights of millimeter-wave radar and visual features can be dynamically adjusted according to the target movement trend. Furthermore, considering multiple dynamic scenes, this effectively improves the robustness and adaptability of the system in fast-moving and complex dynamic scenes, making it suitable for environmental perception in autonomous vehicles.

[0017] In some embodiments, the second feature set includes a second radar feature set and a second visual feature set; the step of fusing the radar feature set and the visual feature set according to the fusion weights through an attention mechanism to obtain fused features, and then performing fusion filtering on the updated first feature set to obtain the second feature set, includes: Based on the fusion weights, the first radar feature set and the first visual feature set in the updated first feature set are weighted and concatenated to obtain the fusion input embedding matrix. The fused features are obtained by performing feature fusion on the fused input embedding matrix through an attention mechanism; The features that cannot be fused in the updated first feature set are removed respectively to obtain the second radar feature set and the second visual feature set.

[0018] In this embodiment, a fusion input embedding matrix is ​​obtained based on the fusion weights and the updated first feature set (including a first radar feature set and a first visual feature set). An attention mechanism is used to perform feature fusion on the fusion input embedding matrix, fully utilizing the complementary characteristics between radar and visual data to enhance global perception capabilities and feature richness. Furthermore, features that cannot be fused in the updated first radar and first visual feature sets are removed, resulting in a second radar feature set and a second visual feature set. These second radar and second visual feature sets represent fusionable features that can be used for subsequent vehicle target detection tasks, improving the detection accuracy, robustness, and adaptability of dynamic targets around vehicles in fast-moving and complex dynamic scenarios.

[0019] In some embodiments, the step of performing feature fusion on the fused input embedding matrix through an attention mechanism to determine the fused features includes: Using the accuracy of the fusion perception task as the loss function, the parameter matrix of the attention mechanism is iteratively updated through the backpropagation algorithm until the attention weights are adapted to the confidence of radar sensors and vision sensors in different scenarios. The fusion perception task includes target detection task and target tracking task. The fused input embedding matrix is ​​mapped to query features, key features, and value features respectively through the parameter matrix, and the attention weights are calculated through the attention mechanism and the values ​​are weighted and summed to obtain the fused features.

[0020] In this embodiment, the accuracy of the fusion perception task is used as the loss function, and the three parameter matrices are iteratively updated using a backpropagation algorithm. Based on these three parameter matrices and the fusion input embedding matrix, attention weights are calculated using an attention mechanism to obtain the fusion features. These fusion features are adapted to the confidence levels of different sensors (4D millimeter-wave radar and visual sensors) in the current dynamic scene, achieving deep fusion of multimodal features. These fusion features, used for fusion perception tasks in complex dynamic scenes, can improve the detection performance of autonomous vehicles.

[0021] In some embodiments, after fusing the radar feature set and the visual feature set according to the fusion weights using an attention mechanism to obtain fused features, and then performing fusion filtering on the updated first feature set to obtain a second feature set, the method further includes: Based on the dynamic target set, the second radar feature set and the second visual feature set in the second feature set are classified as dynamic targets respectively, and the category information of the dynamic targets is obtained. Based on the category information, a second radar feature set and a second visual feature set with common dynamic targets are extracted, and an initial third radar feature set and an initial third visual feature set are determined. Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the initial third visual feature set, and update the dynamic confidence of multiple dynamic targets based on the category relationship matrix; In the initial third radar feature set and the initial third visual feature set, dynamic targets with a dynamic confidence level less than or equal to a preset threshold are discarded, and the target detection information, the initial third radar feature set, and the initial third visual feature set are updated. Output at least one of the following: the fusion weight, the updated target detection information, the updated third radar feature set, and the updated third visual feature set; Wherein, the category information corresponding to the common dynamic target in the second radar feature set and the second visual feature set is the same target category.

[0022] In this embodiment, the second radar feature set and the second visual feature set are features that have undergone spatiotemporal consistency processing, dynamic attribute adaptive weighting, and multimodal fusion. The category information of dynamic targets extracted from the second radar feature set and the second visual feature set are all features of the target category, thus determining the third radar feature set and the third visual feature set. Prediction is performed on the third radar feature set and the third visual feature set to obtain the corresponding target detection information. Based on the similarity between the recognition results (i.e., target detection information) of different sensors (radar and vision) performing the target detection task, the dynamic confidence level corresponding to the multiple dynamic targets can be updated in real time. Targets with insufficient dynamic confidence are discarded, and the target detection information, the third radar feature set, and the third visual feature set are updated. By combining and considering the similarity between the results of different sensors performing the fusion perception task, discarding targets that do not meet the standards, the detection accuracy of different target categories is optimized, and efficient detection and classification of dynamic targets around the vehicle is achieved.

[0023] In some embodiments, determining the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the initial third visual feature set, and updating the dynamic confidence of the plurality of dynamic targets based on the category relationship matrix, includes: Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the target detection information corresponding to the initial third visual feature set; In the third radar feature set and the third visual feature set, the target objects corresponding to the maximum similarity in each similarity are removed, the third radar feature set and the third visual feature set are updated, and the dynamic confidence of the target object is updated according to the maximum similarity. Repeat the steps of determining the category relationship matrix between the target detection information corresponding to the third radar feature set and the target detection information corresponding to the third visual feature set, removing the target object corresponding to the maximum similarity in each similarity in the third radar feature set and the third visual feature set, updating the third radar feature set and the third visual feature set, and updating the dynamic confidence of the target object according to the maximum similarity, until the similarity of each dynamic target object in the category relationship matrix is ​​determined, and the dynamic confidence of each dynamic target object is updated according to the similarity of each dynamic target object.

[0024] In this embodiment, by determining the similarity between the target detection information of each target in the third radar feature set and the initial third visual feature set, the similarity between the results of different sensors performing the fusion perception task is judged, and the dynamic confidence is updated. Furthermore, target objects with unsatisfactory dynamic confidence can be eliminated to optimize the vehicle detection results.

[0025] Secondly, a multimodal fusion device is provided, comprising: The acquisition module is used to acquire radar data and visual data; The fusion module is used to perform spatiotemporal consistency processing and weighted fusion on the radar data and the visual data to obtain fused features.

[0026] Thirdly, an electronic device is provided, comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the multimodal fusion method as described in the first aspect.

[0027] Fourthly, a vehicle is provided that includes electronic equipment as described in the third aspect.

[0028] Fifthly, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the multimodal fusion method as described in the first aspect.

[0029] In a sixth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being used to run programs or instructions to implement the steps of the multimodal fusion method as described in the first aspect.

[0030] In a seventh aspect, a computer program / program product is provided, the computer program / program product being stored in a storage medium, the computer program / program product being executed by at least one processor to implement the steps of the multimodal fusion method as described in the first aspect. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a flowchart illustrating a multimodal fusion method provided in an embodiment of this application; Figure 2 This is a schematic diagram of time alignment provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a multimodal fusion device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0034] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0035] Autonomous driving systems are increasingly demanding higher accuracy, robustness, and real-time performance in environmental perception. Visual sensors provide rich semantic information, but their performance is unstable under complex lighting conditions; millimeter-wave radar performs excellently in adverse weather conditions, but its resolution and ability to detect non-metallic targets are limited. Most multimodal fusion methods employ simple cascading or weighting strategies, which suffer from low robustness, limited detection performance in dynamic scenes, and insufficient accuracy in environmental perception.

[0036] To address, or at least partially address, the aforementioned problems, embodiments of this application provide a multimodal fusion method. This method combines dynamic attribute adaptive weighting, spatiotemporal consistency processing, and deep fusion of multimodal features, improving detection performance in complex dynamic scenarios and providing reliable support for environmental perception of autonomous vehicles.

[0037] The multimodal fusion method, apparatus, and related equipment provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0038] Figure 1 This is a flowchart illustrating a multimodal fusion method provided in an embodiment of this application. Figure 1 As shown, this multimodal fusion method, applied to vehicles, includes: Step 100: Acquire radar data and visual data; The system acquires radar data obtained by a radar sensor detecting the environment surrounding the vehicle. The radar sensor is a millimeter-wave radar capable of acquiring spatial data. Optionally, the radar sensor includes a 4D millimeter-wave radar, which possesses information perception capabilities in four dimensions: distance, speed, azimuth, and vertical height. The 4D millimeter-wave radar detects the environment surrounding the vehicle and acquires radar features; the radar data includes the three-dimensional spatial position of the target object. Velocity vector and reflected signal strength Among them, the 4D millimeter-wave radar employs multi-polarization transmission technology to improve its detection performance against non-metallic targets.

[0039] Visual data is obtained by a vision sensor detecting the environment surrounding the vehicle. Optionally, the vision sensor includes a vision camera, such as a standard RGB camera, an infrared camera, and a depth camera. The vision sensor captures visual features to obtain visual data from detecting the environment surrounding the vehicle; the visual data includes image information about the vehicle's surroundings. The image information contains rich texture information w and semantic information y.

[0040] Step 200: Perform spatiotemporal consistency processing and weighted fusion on the radar data and the visual data to obtain fused features; In this embodiment, by combining the high-precision physical perception capability of radar with the rich semantic information of the vision module for multimodal feature fusion, the global perception capability and feature representation richness are enhanced. Spatiotemporal consistency processing of radar and vision data solves the spatiotemporal bias problem introduced by sampling differences among multiple sensors. Weighted fusion of the spatiotemporally consistent radar and vision data yields fused features. By combining millimeter-wave radar and vision sensors, and utilizing spatiotemporal consistency processing and weighted fusion techniques, the system's detection accuracy, robustness, and adaptability in fast-moving and complex dynamic scenarios are effectively improved, reliably supporting environmental perception for autonomous vehicles.

[0041] Optionally, step 200 includes steps 1001 to 1003; Step 1001: Perform spatiotemporal consistency processing on the radar data and the visual data to determine the dynamic target set and the first feature set; wherein, the first feature set includes the spatiotemporally aligned first radar feature set and the first visual feature set; Spatiotemporal consistency processing is performed on radar data and visual data, which includes time alignment, spatial alignment and time consistency processing.

[0042] In temporal alignment, the radar data and visual data are aligned along the time dimension. In spatial alignment, the 4D millimeter-wave radar and visual sensor are moved from their respective sensor coordinate systems to the vehicle coordinate system, and the radar data and visual data are aligned along the spatial dimension. During spatial alignment, potentially moving targets with characteristic points, i.e., dynamic targets, can also be identified in the radar and visual data. In temporal consistency processing, multiple consecutive frames of temporally and spatially aligned radar and visual data can be extracted to obtain radar and visual data with temporal and spatial consistency, i.e., a first radar feature set and a first visual feature set.

[0043] Step 1002: Based on the set of dynamic target objects, update the first feature set, and determine the current dynamic scene and the fusion weight under the current dynamic scene; Based on the set of dynamic targets, the first feature set is updated, and the current dynamic scene and the fusion weight under the current dynamic scene are determined based on the set of dynamic targets.

[0044] Understandably, in complex and dynamic scenarios, the fusion weights can be adjusted based on the dynamic set of targets. These fusion weights include radar weights and visual weights. The fusion weights can be used to represent the weight ratio of radar features and visual features in subsequent feature fusion.

[0045] Step 1003: Using an attention mechanism, based on the fusion weights, perform feature fusion on the radar feature set and the visual feature set to obtain fused features, and perform fusion filtering on the updated first feature set to obtain a second feature set.

[0046] The second feature set includes a second radar feature set and a second visual feature set. The updated first radar feature set includes the first radar feature set and the first dynamic attribute, and the updated first visual feature set includes the first visual feature set and the second dynamic attribute.

[0047] Using an attention mechanism, based on fusion weights, features are fused from the updated first radar feature set and the updated first visual feature set to obtain fused features. After feature fusion, features that cannot be fused from the updated first radar feature set and the updated first visual feature set are removed, resulting in a second radar feature set and a second visual feature set. It can be understood that the second radar feature set and the second visual feature set represent the features that can be fused from the first radar feature set and the first visual feature set, respectively.

[0048] In this embodiment, by combining the high-precision physical perception capability of radar with the rich semantic information of the vision module for multimodal feature fusion, the global perception capability and feature representation richness are enhanced. By performing spatiotemporal consistency processing on radar and vision data, including time alignment, spatial alignment, and time consistency processing, the spatiotemporal bias problem introduced by sampling differences from multiple sensors is solved. Based on the spatiotemporally consistent radar and vision data, dynamic attributes are calculated, and an updated first feature set is determined. This updated first feature set includes dynamic attributes and the spatiotemporally consistent radar and vision data. The fusion weights of millimeter-wave radar and vision features are dynamically adjusted to achieve optimized target feature fusion. Through an attention mechanism, feature fusion is performed on the updated first feature set (including a first radar feature set and a first vision feature set), fully utilizing the complementary characteristics between radar and vision data to enhance global perception capability and feature representation richness. By combining millimeter-wave radar and vision sensors, and utilizing spatiotemporal consistency processing, dynamic attribute adaptive weighting, and multimodal fusion technologies, the system's detection accuracy, robustness, and adaptability in fast-moving and complex dynamic scenarios are effectively improved, thus reliably supporting environmental perception for autonomous vehicles.

[0049] In some embodiments, updating the first feature set based on the dynamic target object set and determining the current dynamic scene and the fusion weights under the current dynamic scene includes: Based on the set of dynamic target objects and the first feature set, a set of dynamic attributes is determined, and the first feature set is updated according to the set of dynamic attributes. Based on the set of dynamic targets, the current dynamic scene is determined, and the fusion weight under the current dynamic scene is determined according to the set of dynamic attributes; The dynamic attribute set includes a first dynamic attribute and a second dynamic attribute. Based on the dynamic target set, multiple dynamic targets in a first radar feature set and a first visual feature set can be identified. Then, according to the first radar feature set and the first visual feature set, the first dynamic attribute corresponding to the first radar feature set and the second dynamic attribute corresponding to the first visual feature set are calculated respectively. Both the first dynamic attribute and the second dynamic attribute are attributes of the multiple dynamic targets.

[0050] Based on the dynamic attribute set, the first radar feature set and the first visual feature set are updated. It can be understood that the first radar feature set is updated to the first radar feature set plus the first dynamic attribute; the first visual feature set is updated to the first visual feature set plus the second dynamic attribute.

[0051] The current dynamic scene can be determined based on the set of dynamic targets. Then, the fusion weights for the current dynamic scene are determined based on the first dynamic attribute and the second dynamic attribute. It is understood that in complex dynamic scenes, the fusion weights can be adjusted in real time based on the dynamic attributes.

[0052] In this embodiment of the application, dynamic attribute adaptive weighting is used, wherein the fusion weight of millimeter-wave radar and visual features can be dynamically adjusted according to the target motion trend. After subsequent feature fusion, the target feature optimization fusion for different motion states can be achieved.

[0053] In some embodiments, performing spatiotemporal consistency processing on the radar data and the visual data to determine the dynamic target set and the first feature set includes: Using a time interpolation algorithm, a Kalman filter is used to align the radar data and the visual data in time. By using extrinsic parameter calibration technology, the time-aligned radar data and time-aligned visual data are unified into the vehicle coordinate system to obtain spatiotemporally aligned radar data and spatiotemporally aligned visual data. The visual feature points corresponding to the spatiotemporally aligned radar data and the spatiotemporally aligned visual data are determined, and the motion vectors of the visual feature points are calculated using the optical flow method to determine the dynamic target set. The spatiotemporally aligned radar data and spatiotemporally aligned visual data of multiple consecutive frames are input into the spatiotemporal encoder, and time consistency processing is performed to obtain the first radar feature set and the first visual feature set. Wherein, the first radar feature set is the radar data with spatiotemporal consistency, and the first visual feature set is the visual data with spatiotemporal consistency.

[0054] The radar data and visual data undergo spatiotemporal consistency processing, which includes time alignment, spatial alignment, and time consistency processing. It should be noted that this embodiment does not strictly limit the order of the steps.

[0055] The spatiotemporal consistency processing includes steps 300 to 303: Step 300: Using a time interpolation algorithm, the radar data and visual data are time-aligned using a Kalman filter; Optionally, step 300 includes steps 3000 to 3002: Step 3000: Based on the sampling frequency of the radar data and the visual data, determine a unified set of timestamps for interpolation; A unified set of timestamps for interpolation can be determined based on the sampling frequency difference between the 4D millimeter-wave radar and the vision sensor. This set of timestamps is then used for subsequent time interpolation of radar and vision data to achieve time alignment.

[0056] Step 3001: Time-align the radar data and the visual data by interpolation, wherein the radar data is real-time data or legacy data, and the visual data is real-time data or legacy data; Based on the timestamp set, the radar data and the visual data are time-aligned through interpolation.

[0057] Figure 2 This is a schematic diagram illustrating time alignment provided in an embodiment of this application. For example... Figure 2 As shown, each timestamp in this unified timestamp set corresponds to one radar data and one visual data; the radar data can be real-time or legacy data, and the visual data can also be real-time or legacy data. Legacy data refers to stored outdated data, not real-time acquired data. For example, the T+50MS timestamp corresponds to visual data 2 (real-time data) of target A' and radar data 1 (legacy data) of target A'.

[0058] The real-time data is radar or visual data acquired in real time, and the legacy data can be an estimate predicted using a Kalman filter to correct for biases.

[0059] Step 3002: If the current timestamp belongs to the set of timestamps, use a Kalman filter to predict the estimated value of the radar data and the estimated value of the visual data corresponding to the current timestamp, and update the residual data in the radar data to the estimated value of the radar data, and update the residual data in the visual data to the estimated value of the visual data.

[0060] Using the filtered state update formula, a Kalman filter is used to predict the estimated values ​​of the radar data and visual data corresponding to the current timestamp, and the legacy data is updated accordingly based on the estimated values.

[0061] Alternatively, the filter state update formula is as follows: ;and ; in This represents the current state estimate of legacy data. This represents the predicted value of legacy data. For Kalman gain, For the observation matrix, To measure the noise covariance matrix.

[0062] In this embodiment, a time interpolation algorithm is used to perform time alignment of the radar data and the visual data using a Kalman filter, so as to ensure that the time alignment error is less than or equal to a preset threshold, thereby avoiding the radar and visual observations of the same target object corresponding to different states at different times, and further avoiding the distortion of subsequent dynamic attribute calculations.

[0063] Step 301: Using extrinsic parameter calibration technology, the time-aligned radar data and time-aligned visual data are unified into the vehicle coordinate system to obtain spatiotemporally aligned radar data and spatiotemporally aligned visual data.

[0064] A unified coordinate system (vehicle coordinate system) is established for the 4D millimeter-wave radar and vision camera using extrinsic parameter calibration technology, and a spatial alignment function is determined. The spatial alignment function can be expressed as a transformation function of the sensor coordinate system relative to the vehicle coordinate system.

[0065] The spatiotemporal features of point cloud (e.g., distance, velocity, and azimuth) in the time-aligned radar data and the image semantic features (e.g., target outline and pixel position) in the time-aligned visual data are unified into a bird's-eye view (BEV) to ensure that the coordinate information input dimension of all targets is consistent. In other words, the time-aligned radar data and the visual data are unified into the vehicle coordinate system to obtain the spatiotemporally aligned radar data and the visual data.

[0066] Step 302: Determine the visual feature points corresponding to the spatiotemporally aligned radar data and the spatiotemporally aligned visual data, and calculate the motion vectors of the visual feature points using the optical flow method to determine the dynamic target set.

[0067] Visual feature points are key information that can characterize the local uniqueness of an image; visual feature points can be extracted from the input image.

[0068] The input images corresponding to the spatiotemporally aligned radar data and visual data are determined, and the corresponding visual feature points are identified. For example, visible light images output by ordinary RGB cameras, infrared images output by infrared cameras, depth images output by depth cameras, and depth images and bird's-eye views generated by LiDAR projection can all be used as input images to extract visual feature points.

[0069] By using optical flow, the positional changes of the same target object in two consecutive frames of data (radar data and visual data) are matched, and motion information is calculated to determine the motion vector of the corresponding visual feature points. If the motion vector of a certain area differs significantly from the motion vector of the surrounding background, it can be identified as a dynamic target; and targets with feature points that may move can be identified, thus determining the set of dynamic targets.

[0070] Step 303: Input the spatiotemporally aligned radar data and spatiotemporally aligned visual data of multiple consecutive frames into the spatiotemporal encoder, perform time consistency processing, and obtain the first radar feature set and the first visual feature set. Wherein, the first radar feature set is the radar data with spatiotemporal consistency, and the first visual feature set is the visual data with spatiotemporal consistency.

[0071] The spatiotemporal encoder records spatiotemporally aligned radar data and spatiotemporally aligned visual data from a temporal dimension. Multiple consecutive frames of spatiotemporally aligned radar data are input into the spatiotemporal encoder, and temporal consistency processing is performed to generate spatiotemporally consistent radar data (i.e., the first radar feature set). Similarly, multiple consecutive frames of spatiotemporally aligned visual data are input into the spatiotemporal encoder, and temporal consistency processing is performed to generate spatiotemporally consistent visual data (i.e., the first visual feature set).

[0072] For example, the spatiotemporal encoder, through time consistency processing, takes the spatiotemporally aligned radar data or the spatiotemporally aligned visual data as input, and generates corresponding spatiotemporally consistent radar data or spatiotemporally consistent visual data. ;

[0073] in, This can represent the spatiotemporally aligned radar data or the spatiotemporally aligned visual data corresponding to the current timestamp; For radar data or visual data that are spatiotemporally consistent.

[0074] In this embodiment, spatiotemporal consistency processing is performed on radar and visual data. This processing includes time alignment, spatial alignment, and time consistency processing to address the spatiotemporal bias caused by sampling differences between multiple sensors and improve detection accuracy in dynamic scenes. Specifically, a Kalman filter is used for time alignment to correct biases and avoid distortion in dynamic attribute calculations. In spatial alignment, targets around the vehicle are determined to be either static or dynamic, identifying the set of dynamic targets in the radar and visual data to calculate their dynamic attributes. Subsequently, the fusion weights can be adjusted based on the target motion trends corresponding to the dynamic attributes, achieving optimized fusion of target features for different motion states.

[0075] In some embodiments, the first dynamic attribute and the second dynamic attribute are determined based on the dynamic target set, according to the first radar feature set and the first visual feature set.

[0076] Optionally, the dynamic attributes corresponding to the first radar feature set and the first visual feature set each include at least one of the following: basic motion attributes of multiple dynamic targets, spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes.

[0077] The basic motion attributes include: radial relative velocity, lateral / longitudinal movement velocity, acceleration, accelerometer movement velocity, and target state.

[0078] The spatiotemporal correlation attributes include: motion trajectory fitting degree and / or life cycle status; Motion trajectory fit is used to represent the fit of the motion trajectory of a dynamic target under different sensors. Lifecycle state is used to represent the continuous state of data in the spatiotemporal encoder (which can represent the continuous state of the target).

[0079] The interactive behavior features refer to the relative motion relationship between the vehicle and multiple dynamic targets in the surrounding area (e.g., crossing, following other vehicles).

[0080] The perception enhancement attributes include: multi-sensor consistency and / or dynamic confidence; Multi-sensor consistency refers to the measure of whether the observation results from different sensors observing the same target or scene maintain consistency. Dynamic confidence refers to the measure of the credibility of current features or data.

[0081] Optionally, determining the dynamic attribute set based on the dynamic target set and the first feature set includes at least one of the following: Based on the set of dynamic targets, the basic motion attributes in the set of dynamic attributes are determined according to the difference between the current frame data and the previous frame data in the first feature set. Based on the persistent state of the target object in the first feature set of the spatiotemporal encoder, the life cycle state in the dynamic attribute set is determined. By comparing the motion trajectories in the first radar feature set and the first visual feature set, the fitting degree of the motion trajectory in the dynamic attribute set is determined. Based on the first feature set, determine the interaction behavior between the vehicle and surrounding dynamic targets, and determine the interaction behavior features in the dynamic attribute set; The position information of the same target in the first radar feature set and the first visual feature set are compared, and the consistency of the multiple sensors in the dynamic attribute set is determined based on the deviation of the position information. The dynamic confidence level in the dynamic attribute set is determined based on the multi-sensor consistency and the stable state of the first feature set over multiple consecutive frames.

[0082] The process of determining the dynamic attribute set includes at least one of the following: determining the basic motion attributes; determining the motion trajectory fitting degree; determining the life cycle state; determining the interactive behavior characteristics; determining the multi-sensor consistency; determining the multi-sensor consistency.

[0083] The process of determining basic motion attributes: The 4D millimeter-wave radar can acquire the velocity difference, three-dimensional spatial position difference, and change difference between the current frame data and the previous frame data, and calculate the basic motion attribute in the first dynamic attribute. The vision sensor can acquire the velocity difference, three-dimensional spatial position difference, and change difference between the current frame data and the previous frame data, and calculate the basic motion attribute in the second dynamic attribute.

[0084] The process of determining the motion trajectory fitting degree is as follows: The motion trajectories in the first radar feature set and the first visual feature set are observed and compared to determine the motion trajectory fitting degree in the dynamic attribute set. For example, the motion trajectories in radar data between time stamp T and time stamp T+1 and the motion trajectories in visual data between time stamp T and time stamp T+1 are observed and compared. The Euclidean distance and direction cosine of the two trajectory lines are calculated, and the similarity of the direction cosines is compared to determine the motion trajectory fitting degree.

[0085] The process of determining the lifecycle state involves obtaining the lifecycle state from the first dynamic attribute and the second dynamic attribute in the spatiotemporal encoder. The lifecycle state includes the emergence state, the stable state, and the disappearance state. For example, if a feature or target object has just appeared in the spatiotemporal encoder, it is considered a newly emerged state; if it continues to appear, it is considered a continuous state. Furthermore, if the encoder at time stamp T+1 does not contain the feature or target object, but the encoder at time stamp T contains the feature or target object, it is considered a disappearance state.

[0086] The determination of the spatiotemporal correlation attributes includes: the determination of the motion trajectory fitting degree, and / or the determination of the life cycle state.

[0087] The process of determining interactive behavior features: Based on radar feature data and visual feature data, determine the distance change trend between the current vehicle and surrounding targets, judge the interactive behavior between the surrounding targets and the vehicle, and determine the corresponding interactive behavior features in the first dynamic attribute and the second dynamic attribute.

[0088] The process of determining multi-sensor consistency: Based on the deviation between the position information of the same target in radar features and the position information of the same target in visual features, multi-sensor consistency can be determined.

[0089] The process of determining dynamic confidence is as follows: Combining multi-sensor consistency and the stability of multiple consecutive frames in the spatiotemporal encoder, a comprehensive weighted evaluation is performed to determine the dynamic confidence in the corresponding first and second dynamic attributes.

[0090] Optionally, the stability of multiple consecutive frames in the spatiotemporal encoder is related to the timestamp alignment position (whether the sensor data corresponding to the current timestamp is real-time data or legacy data). For example, if multiple consecutive frames of sensor data are legacy data, then the stability of the data across multiple consecutive frames is determined to be low.

[0091] The determination of the perception enhancement attribute includes: the determination of the multi-sensor consistency; and / or, the determination of the dynamic confidence level.

[0092] In some embodiments, the fusion weights include radar weights and visual weights; The step of determining the current dynamic scene based on the set of dynamic targets and determining the fusion weights under the current dynamic scene according to the set of dynamic attributes includes: Based on the set of dynamic targets, the density of dynamic targets around the vehicle is calculated, and based on the density of the dynamic targets, the current dynamic scene is determined; The spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes in the dynamic attributes corresponding to the first radar feature set and the first visual feature set are weighted and summed according to a preset ratio to determine the radar weight and the visual weight, respectively. The preset ratio is related to the current dynamic scene. Based on the basic motion attributes in the dynamic attributes corresponding to the first radar feature set and the first visual feature set, the interaction parameters between the vehicle and the surrounding dynamic targets are determined; if the interaction parameters corresponding to the first radar feature set are outside the safe threshold range, or if the interaction parameters corresponding to the first visual feature set are outside the safe threshold range, the radar weights and the visual weights are updated.

[0093] The current dynamic scene includes: a normal scene and a high-interaction-density scene. The density of dynamic objects can refer to the number of dynamic objects per unit square meter. The current dynamic scene can be determined as a normal scene or a high-interaction-density scene based on the density of dynamic objects around the vehicle.

[0094] In a normal scenario, the spatiotemporal correlation attribute, interactive behavior feature, and perception enhancement attribute in the first dynamic attribute are weighted and summed using a first preset ratio to determine the radar weight. Similarly, in a normal scenario, the spatiotemporal correlation attribute, interactive behavior feature, and perception enhancement attribute in the second dynamic attribute are weighted and summed using a first preset ratio to determine the visual weight. In this normal scenario, the stability of target association is prioritized. For example, in a normal scenario, the radar weight and visual weight are determined by summing 40% for the spatiotemporal key attribute, 35% for the perception enhancement attribute, and 25% for the interactive behavior feature.

[0095] In a high-interaction-density scenario, the spatiotemporal correlation attribute, interactive behavior features, and perception enhancement attribute in the first dynamic attribute are weighted and summed using a second preset ratio to determine the radar weight; similarly, the spatiotemporal correlation attribute, interactive behavior features, and perception enhancement attribute in the second dynamic attribute are weighted and summed using a second preset ratio to determine the visual weight. This high-interaction-density scenario emphasizes capturing the subject's decision-making intent. For example, in a high-interaction-density scenario (e.g., congested roads and densely populated areas), the radar weight and visual weight are determined by summing 40% of the interactive behavior features, 30% of the spatiotemporal attribute, and 30% of the perception attribute, respectively.

[0096] Based on the basic motion attributes of the first dynamic attribute, a first interaction parameter between the vehicle and surrounding dynamic targets is determined; and based on the basic motion attributes of the second dynamic attribute, a second interaction parameter between the vehicle and surrounding dynamic targets is determined. If either the first or second interaction parameter is outside a safe threshold range, the radar weight and the visual weight are updated; if both the first and second interaction parameters are within a safe threshold range, the radar weight and the visual weight remain unchanged.

[0097] The interaction parameters include collision time. For example, the collision time between the vehicle and surrounding dynamic targets can be calculated using the radial velocity and acceleration from the first and second dynamic attributes. If the collision time corresponding to the first dynamic attribute is less than or equal to a safety threshold, or if the interaction parameter corresponding to the second dynamic attribute is less than or equal to a safety threshold, the radar weight in the fusion weight is increased.

[0098] Optionally, if the dynamic confidence of a sensor (4D millimeter-wave radar or vision sensor) is less than or equal to a first threshold (e.g., 0.5), the fusion weight corresponding to the sensor is reduced.

[0099] If the dynamic confidence level of a sensor is less than or equal to a first threshold, the current observation value of that sensor is unreliable. This can be addressed by temporarily stopping the fusion of features corresponding to that sensor and using only data from another sensor for decision-making (i.e., setting the corresponding sensor value to 0), thus preventing abnormal data from interfering with the fusion results.

[0100] In this embodiment, the current dynamic scene is determined to be either a conventional scene or a high-interaction-density scene based on the set of dynamic targets. Using a preset ratio corresponding to the current dynamic scene, the spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes in the first and second dynamic attributes are weighted and summed to obtain radar weights and visual weights, respectively. Based on the basic motion attributes in the first and second dynamic attributes, it can be determined whether the interaction parameters (e.g., collision time) corresponding to the first and second dynamic attributes are outside the safety threshold range, and the radar weights and visual weights are updated accordingly. This achieves dynamic adjustment of the fusion weights based on dynamic attributes, meaning the fusion weights of millimeter-wave radar and visual features can be dynamically adjusted according to the target movement trend. Furthermore, considering multiple dynamic scenes, this effectively improves the robustness and adaptability of the system in fast-moving and complex dynamic scenes, making it suitable for environmental perception in autonomous vehicles.

[0101] In some embodiments, the fusion weights may also incorporate the sensor’s historical performance and scene characteristics preset.

[0102] By using labeled historical datasets, the model can automatically learn the optimal initial weight distribution for different scenarios through deep learning or statistical learning. For example, radar weights are higher in high-speed scenarios, and visual semantic feature weights are higher in urban scenarios, thus adapting to changes in complex and dynamic environments.

[0103] In some embodiments, the step of using an attention mechanism to perform weighted fusion of the first radar feature set and the first visual feature set in the updated first feature set according to the fusion weights to obtain fused features, and then performing fusion filtering on the updated first feature set to obtain a second feature set, includes: Based on the fusion weights, the first radar feature set and the first visual feature set in the updated first feature set are weighted and concatenated to obtain the fusion input embedding matrix. The fused features are obtained by performing feature fusion on the fused input embedding matrix through an attention mechanism; The features that cannot be fused in the updated first feature set are removed respectively to obtain the second radar feature set and the second visual feature set.

[0104] The updated first radar feature set and the updated first visual feature set are weighted and concatenated according to the fusion weights to obtain the fusion input embedding matrix. The fusion input embedding matrix is ​​then fused using an attention mechanism to obtain the fused features.

[0105] It is understandable that adjusting the fusion weights can affect the fusion features. Since the fusion weights can be adaptively adjusted according to dynamic attributes, the target features can be optimized for different motion states.

[0106] Optionally, the step of performing feature fusion on the fused input embedding matrix through an attention mechanism to determine the fused features includes: Using the accuracy of the fusion perception task as the loss function, the parameter matrix of the attention mechanism is iteratively updated through the backpropagation algorithm until the attention weights are adapted to the confidence of radar sensors and vision sensors in different scenarios. The fusion perception task includes target detection task and target tracking task. The fused input embedding matrix is ​​mapped to query features, key features, and value features respectively through the parameter matrix, and the attention weights are calculated through the attention mechanism and the values ​​are weighted and summed to obtain the fused features.

[0107] The parameter matrices are three learnable and pre-trainable parameter matrices. By using the accuracy of the fusion perception task as the loss function, the first, second, and third parameter matrices are iteratively updated through backpropagation to ensure that the attention weights are adapted to the confidence levels of 4D millimeter-wave radar and visual sensors in different scenarios.

[0108] The original input data (i.e., the fused input embedding matrix) is linearly transformed with the pre-trained first parameter matrix to obtain query features; the fused input embedding matrix is ​​linearly transformed with the second parameter matrix to obtain key features; and the fused input embedding matrix is ​​linearly transformed with the third parameter matrix to obtain value features. Through the attention mechanism, the key features that can answer the query features are fused with the value features to obtain the fused features.

[0109] In this embodiment, the fusion feature can be used for fusion perception tasks (including target detection tasks and target tracking tasks). It is understood that, because the fusion feature combines the high-precision physical perception capability of millimeter-wave radar with the rich semantic information of the vision module, the detection performance of performing fusion perception tasks using the fusion feature is higher than that of performing fusion perception tasks using only radar features or visual features.

[0110] After fusing radar and visual features to obtain fused features, features that cannot be fused in the updated first radar feature set and first visual feature set are removed, resulting in the fused features in the updated first radar feature set and first visual feature set, namely the second radar feature set and the second visual feature set. For example, the second radar feature set and the second visual feature set can be determined based on the trained parameter matrix, the first radar feature set, and the first visual feature set.

[0111] For example, a cross-modal attention mechanism is constructed based on the Transformer's multimodal feature extraction module, and the attention mechanism includes the following steps: 1. Input Construction: The updated first radar feature set and the updated set of first visual features The fused input embedding matrix X is obtained by weighting and concatenating the radar weights a and b according to their respective weights. ; 2. Pre-training: Using the accuracy of the fusion perception task (e.g., object detection and object tracking) as the loss function, the parameters of Wq, Wk, and Wv are iteratively updated through backpropagation, so that the attention weights calculated by Q and K can be automatically adapted to the confidence of the two types of sensors in different scenarios. 3. Mapping generation: Perform linear transformations on the fused input embedding matrix X with the three independent trainable parameter matrices Wq, Wk, and Wv respectively to obtain Q=X·Wq, K=X·Wk, and V=X·Wv; 4. Calculate the weighted features The output of the attention mechanism is obtained as the fusion feature. ; 5. Remove features from the updated first radar feature set. and the updated set of first visual features The features that cannot be fused in the first radar feature set are used to obtain the second radar feature set. Second visual feature set .

[0112] In this embodiment, a fusion input embedding matrix is ​​obtained based on the fusion weights and the updated first feature set (including a first radar feature set and a first visual feature set). An attention mechanism is used to perform feature fusion on the fusion input embedding matrix, fully utilizing the complementary characteristics between radar and visual data to enhance global perception capabilities and feature richness. Specifically, the accuracy of the fusion perception task is used as the loss function, and the three parameter matrices are iteratively updated using a backpropagation algorithm. Based on these three parameter matrices and the fusion input embedding matrix, attention weights are calculated using the attention mechanism to obtain the fusion features. These fusion features are adapted to the confidence levels of different sensors (4D millimeter-wave radar and visual sensors) in the current dynamic scene, achieving deep fusion of multimodal features. These fusion features are used for fusion perception tasks in complex dynamic scenes, improving the detection performance of autonomous vehicles. Furthermore, after obtaining the fused features, features that cannot be fused in the updated first radar feature set and first visual feature set are removed to obtain the second radar feature set and the second visual feature set. The second radar feature set and the second visual feature set represent fused features, which can be used for subsequent vehicle target detection tasks to improve the detection accuracy, robustness and adaptability of dynamic targets around the vehicle in fast-moving and complex dynamic scenes.

[0113] This multimodal fusion method includes data acquisition, spatiotemporal consistency processing, dynamic fusion, and multimodal feature fusion. After the aforementioned processes, this multimodal fusion method can also perform a multi-target fusion perception task and optimize and update the radar feature set and visual feature set to obtain the corresponding output results (including the detection results of the fusion perception task, the updated radar feature set, and the updated visual feature set).

[0114] In some embodiments, after weighted fusion of the first radar feature set and the first visual feature set in the updated first feature set using an attention mechanism and according to the fusion weights to obtain fused features, and after fusion filtering of the updated first feature set to obtain the second feature set, the method further includes: Based on the dynamic target set, the second radar feature set and the second visual feature set in the second feature set are classified as dynamic targets respectively, and the category information of the dynamic targets is obtained. Based on the category information, a second radar feature set and a second visual feature set with common dynamic targets are extracted, and an initial third radar feature set and an initial third visual feature set are determined. Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the initial third visual feature set, and update the dynamic confidence of multiple dynamic targets based on the category relationship matrix; In the initial third radar feature set and the initial third visual feature set, dynamic targets with a dynamic confidence level less than or equal to a preset threshold are discarded, and the target detection information, the initial third radar feature set, and the initial third visual feature set are updated. Output at least one of the following: the fusion weight, the updated target detection information, the updated third radar feature set, and the updated third visual feature set; Wherein, the category information corresponding to the common dynamic target in the second radar feature set and the second visual feature set is the same target category.

[0115] The process involves performing a multi-target fusion perception task and optimizing and updating the radar feature set and visual feature set to obtain the corresponding output results, including the following steps S1 to S5: Step S1: Based on the dynamic target set, classify the second radar feature set and the second visual feature set in the second feature set into dynamic targets respectively, and obtain the category information of the dynamic targets; The target is identified by the second radar feature set and the second visual feature set, and the target with dynamic attributes (i.e., dynamic target) is classified and distinguished to obtain the first category information (e.g., people and vehicles) corresponding to the second radar feature set, and the second category information corresponding to the second visual feature set is obtained.

[0116] Step S2: Based on the category information, extract the second radar feature set and the second visual feature set that have common dynamic targets, and determine the initial third radar feature set and the initial third visual feature set. The common dynamic target objects in the second radar feature set and the second visual feature set all correspond to the same target category (e.g., a person). Target objects whose first and second category information does not correspond to the target category can be removed from the second radar feature set and the second visual feature set, and then the third radar feature set and the third visual feature set can be determined and stored initially.

[0117] Step S3: Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the initial third visual feature set, and update the dynamic confidence of multiple dynamic targets based on the category relationship matrix; Optionally, step S3 includes steps S31 to S33: Step S31: Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the target detection information corresponding to the initial third visual feature set; Using a neural network algorithm, multiple detection heads are used to predict target detection information corresponding to the third radar feature set and the third visual feature set, respectively. The multiple detection heads are independent. The target detection information includes the position, category, and velocity information of multiple dynamic targets.

[0118] For example, for radar sensors, through an independent detection head Using convolutional neural network algorithms, the positions of multiple dynamic targets are predicted. ,category and speed ,for For vision sensors, an independent detection head is used. Using convolutional neural network algorithms, the positions of multiple dynamic targets are predicted. ,category and speed ,for .

[0119] A category relationship matrix is ​​determined between the target detection information corresponding to the third radar feature set and the target detection information corresponding to the third visual feature set. This category relationship matrix can establish the correlation between targets from different sensors.

[0120] For example, through and The similarity calculation of the category relationship matrix : ,in This is a similarity measurement function.

[0121] Step S32: In the third radar feature set and the third visual feature set, remove the target object corresponding to the maximum similarity value in each similarity, update the third radar feature set and the third visual feature set, and update the dynamic confidence of the target object according to the maximum similarity value.

[0122] The category relationship matrix can characterize the similarity of the multiple dynamic targets, and the dynamic confidence of the multiple dynamic targets can be updated based on the similarity of the multiple dynamic targets.

[0123] Sort the similarities in the category relationship matrix in descending order, and extract the target object c with the maximum current similarity in the category relationship matrix; Based on the current maximum similarity value corresponding to the target object c, increase the dynamic confidence of the target object c; Target c is removed from the third radar feature set and the third visual feature set, and the third radar feature set and the third visual feature set are updated; wherein, the updated third radar feature set and the third visual feature set include the remaining target objects after removing target c.

[0124] Step S33: Repeat the process of determining the category relationship matrix between the target detection information corresponding to the third radar feature set and the target detection information corresponding to the third visual feature set, and step S32, until the similarity of each dynamic target in the category relationship matrix is ​​determined, and the dynamic confidence of each dynamic target is updated according to the similarity of each dynamic target.

[0125] It is understood that, based on the real-time third radar feature set and the real-time third visual feature set, the process of determining the category relationship matrix between the target detection information corresponding to the third radar feature set and the target detection information corresponding to the third visual feature set, and step S32, can be repeated to update the dynamic confidence of each dynamic target in real time.

[0126] In this embodiment, by determining the similarity between the target detection information of each target in the third radar feature set and the initial third visual feature set, the similarity between the results of different sensors performing the fusion perception task is judged, and the dynamic confidence is updated. Furthermore, target objects with unsatisfactory dynamic confidence can be eliminated to optimize the vehicle detection results.

[0127] Step S4: In the initial third radar feature set and the initial third visual feature set, discard dynamic targets with a dynamic confidence level less than or equal to a preset threshold, and update the target detection information, the initial third radar feature set and the initial third visual feature set.

[0128] The dynamic confidence of each dynamic target is detected in real time. If the dynamic confidence of a dynamic target is less than or equal to a preset threshold, the dynamic target is discarded from the third radar feature set and the third visual feature set, and the initial third radar feature set and the initial third visual feature set are updated.

[0129] Step S5: Output at least one of the following: the fusion weight, the updated target detection information, the updated third radar feature set, and the updated third visual feature set.

[0130] It should be noted that, through the attention mechanism, based on the fusion weight, feature fusion is performed on the updated first radar feature set and the updated first visual feature set to obtain fused features. After the updated first feature set is fused and filtered to obtain the second feature set, at least one of the fused features, the second radar feature set, and the second visual feature set can be output to the user.

[0131] In this embodiment, the second radar feature set and the second visual feature set are features that have undergone spatiotemporal consistency processing, dynamic attribute adaptive weighting, and multimodal fusion. The category information of dynamic targets extracted from the second radar feature set and the second visual feature set are all features of the target category, thus determining the third radar feature set and the third visual feature set. Prediction is performed on the third radar feature set and the third visual feature set to obtain the corresponding target detection information. Based on the similarity between the recognition results (i.e., target detection information) of different sensors (radar and vision) performing the target detection task, the dynamic confidence level corresponding to the multiple dynamic targets can be updated in real time. Targets with insufficient dynamic confidence are discarded, and the target detection information, the third radar feature set, and the third visual feature set are updated. By combining and considering the similarity between the results of different sensors performing the fusion perception task, discarding targets that do not meet the standards, the detection accuracy of different target categories is optimized, and efficient detection and classification of dynamic targets around the vehicle is achieved.

[0132] The multimodal fusion method provided in this application can be executed by a multimodal fusion device. This application uses the execution of the multimodal fusion method by a multimodal fusion device as an example to illustrate the multimodal fusion device provided in this application.

[0133] Figure 3 This is a schematic diagram of the structure of a multimodal fusion device provided in an embodiment of this application. Figure 3 As shown, the multimodal fusion device 300 includes a data acquisition module 301 and a processing module 302. The multimodal fusion device 300 is a vehicle or a component in a vehicle.

[0134] Data acquisition module 301 is used to acquire radar data and visual data; The processing module 302 is used to perform spatiotemporal consistency processing and weighted fusion on the radar data and the visual data to obtain fused features.

[0135] Optionally, the processing module 302 is configured to: The radar data and the visual data are subjected to spatiotemporal consistency processing to determine a dynamic target set and a first feature set, wherein the first feature set includes a spatiotemporally aligned first radar feature set and a first visual feature set; Based on the set of dynamic targets, the first feature set is updated, and the current dynamic scene and the fusion weights under the current dynamic scene are determined; Using an attention mechanism, the first radar feature set and the first visual feature set in the updated first feature set are weighted and fused according to the fusion weights to obtain fused features. The updated first feature set is then fused and filtered to obtain the second feature set.

[0136] Optionally, the processing module 302 is configured to: Based on the set of dynamic target objects and the first feature set, a set of dynamic attributes is determined, and the first feature set is updated according to the set of dynamic attributes. Based on the set of dynamic target objects, the current dynamic scene is determined, and the fusion weight under the current dynamic scene is determined according to the set of dynamic attributes.

[0137] In this embodiment of the application, dynamic attribute adaptive weighting is used, wherein the fusion weight of millimeter-wave radar and visual features can be dynamically adjusted according to the target motion trend. After subsequent feature fusion, the target feature optimization fusion for different motion states can be achieved.

[0138] Optionally, the processing module 302 is configured to: Using a time interpolation algorithm, a Kalman filter is used to align the radar data and the visual data in time. By using extrinsic parameter calibration technology, the time-aligned radar data and time-aligned visual data are unified into the vehicle coordinate system to obtain spatiotemporally aligned radar data and spatiotemporally aligned visual data. The visual feature points corresponding to the spatiotemporally aligned radar data and the spatiotemporally aligned visual data are determined, and the motion vectors of the visual feature points are calculated using the optical flow method to determine the dynamic target set. The spatiotemporally aligned radar data and spatiotemporally aligned visual data of multiple consecutive frames are input into the spatiotemporal encoder, and time consistency processing is performed to obtain the first radar feature set and the first visual feature set. Wherein, the first radar feature set is the radar data with spatiotemporal consistency, and the first visual feature set is the visual data with spatiotemporal consistency.

[0139] Optionally, the processing module 302 is configured to: Based on the sampling frequencies of the radar data and the visual data, a unified set of timestamps for interpolation is determined; The radar data and the visual data are time-aligned by interpolation, wherein the radar data is real-time data or legacy data, and the visual data is real-time data or legacy data. If the current timestamp belongs to the set of timestamps, use a Kalman filter to predict the estimated value of the radar data and the estimated value of the visual data corresponding to the current timestamp, and update the residual data in the radar data to the estimated value of the radar data, and update the residual data in the visual data to the estimated value of the visual data.

[0140] Optionally, the dynamic attributes corresponding to the first radar feature set and the first visual feature set each include at least one of the following: basic motion attributes of multiple dynamic targets, spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes; the spatiotemporal correlation attributes include: motion trajectory fitting degree and / or life cycle state; the perception enhancement attributes include: multi-sensor consistency and / or dynamic confidence. The processing module 302 is used for at least one of the following: Based on the set of dynamic targets, the basic motion attributes in the set of dynamic attributes are determined according to the difference between the current frame data and the previous frame data in the first feature set. Based on the persistent state of the target object in the first feature set of the spatiotemporal encoder, the life cycle state in the dynamic attribute set is determined. By comparing the motion trajectories in the first radar feature set and the first visual feature set, the fitting degree of the motion trajectory in the dynamic attribute set is determined. Based on the first feature set, determine the interaction behavior between the vehicle and surrounding dynamic targets, and determine the interaction behavior features in the dynamic attribute set; The position information of the same target in the first radar feature set and the first visual feature set are compared, and the consistency of the multiple sensors in the dynamic attribute set is determined based on the deviation of the position information. The dynamic confidence level in the dynamic attribute set is determined based on the multi-sensor consistency and the stable state of the first feature set over multiple consecutive frames.

[0141] Optionally, the fusion weights include radar weights and visual weights; The processing module 302 is used for: Based on the set of dynamic targets, the density of dynamic targets around the vehicle is calculated, and based on the density of the dynamic targets, the current dynamic scene is determined; The spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes in the dynamic attributes corresponding to the first radar feature set and the first visual feature set are weighted and summed according to a preset ratio to determine the radar weight and the visual weight, respectively. The preset ratio is related to the current dynamic scene. Based on the basic motion attributes in the dynamic attributes corresponding to the first radar feature set and the first visual feature set, the interaction parameters between the vehicle and the surrounding dynamic targets are determined; if the interaction parameters corresponding to the first radar feature set are outside the safe threshold range, or if the interaction parameters corresponding to the first visual feature set are outside the safe threshold range, the radar weights and the visual weights are updated.

[0142] Optionally, the second feature set includes a second radar feature set and a second visual feature set; The processing module 302 is used for: Based on the fusion weights, the first radar feature set and the first visual feature set in the updated first feature set are weighted and concatenated to obtain the fusion input embedding matrix. The fused features are obtained by performing feature fusion on the fused input embedding matrix through an attention mechanism; The features that cannot be fused in the updated first feature set are removed respectively to obtain the second radar feature set and the second visual feature set.

[0143] Optionally, the processing module 302 is configured to: Using the accuracy of the fusion perception task as the loss function, the parameter matrix of the attention mechanism is iteratively updated through the backpropagation algorithm until the attention weights are adapted to the confidence of radar sensors and vision sensors in different scenarios. The fusion perception task includes target detection task and target tracking task. The fused input embedding matrix is ​​mapped to query features, key features, and value features respectively through the parameter matrix, and the attention weights are calculated through the attention mechanism and the values ​​are weighted and summed to obtain the fused features.

[0144] Optionally, the processing module 302 is further configured to: Based on the dynamic target set, the second radar feature set and the second visual feature set in the second feature set are classified as dynamic targets respectively, and the category information of the dynamic targets is obtained. Based on the category information, a second radar feature set and a second visual feature set with common dynamic targets are extracted, and an initial third radar feature set and an initial third visual feature set are determined. Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the initial third visual feature set, and update the dynamic confidence of multiple dynamic targets based on the category relationship matrix; In the initial third radar feature set and the initial third visual feature set, dynamic targets with a dynamic confidence level less than or equal to a preset threshold are discarded, and the target detection information, the initial third radar feature set, and the initial third visual feature set are updated. Output at least one of the following: the fusion weight, the updated target detection information, the updated third radar feature set, and the updated third visual feature set; Wherein, the category information corresponding to the common dynamic target in the second radar feature set and the second visual feature set is the same target category.

[0145] Optionally, the processing module 302 is further configured to: Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the target detection information corresponding to the initial third visual feature set; In the third radar feature set and the third visual feature set, the target objects corresponding to the maximum similarity in each similarity are removed, the third radar feature set and the third visual feature set are updated, and the dynamic confidence of the target object is updated according to the maximum similarity. Repeat the steps of determining the category relationship matrix between the target detection information corresponding to the third radar feature set and the target detection information corresponding to the third visual feature set, removing the target object corresponding to the maximum similarity in each similarity in the third radar feature set and the third visual feature set, updating the third radar feature set and the third visual feature set, and updating the dynamic confidence of the target object according to the maximum similarity, until the similarity of each dynamic target object in the category relationship matrix is ​​determined, and the dynamic confidence of each dynamic target object is updated according to the similarity of each dynamic target object.

[0146] The multimodal fusion device provided in this application embodiment can achieve... Figure 1The multimodal fusion method embodiments include at least one step and at least one of the embodiments, and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0147] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 4 As shown, this application embodiment also provides an electronic device 400, including a processor 401 and a memory 402. The memory 402 stores a program or instructions that can run on the processor 401. When the program or instructions are executed by the processor 401, they implement the various steps of the above-described multimodal fusion method embodiment on the access point side and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0148] This application embodiment also provides a vehicle, including as follows: Figure 4 The electronic device shown.

[0149] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described multimodal fusion method embodiments and achieve the same technical effects. To avoid repetition, these will not be described again here.

[0150] The processor mentioned above is either the processor in the electronic device described in the above embodiments or the processor in the vehicle. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.

[0151] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described multimodal fusion method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0152] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0153] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described multimodal fusion method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0154] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0155] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0156] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0157] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0158] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A multi-modal fusion method, characterized in that, Applied to vehicles, including: Acquire radar and visual data; The radar data and the visual data are subjected to spatiotemporal consistency processing and weighted fusion to obtain fused features.

2. The method according to claim 1, characterized in that, The process of performing spatiotemporal consistency processing and weighted fusion on the radar data and the visual data to obtain fused features includes: The radar data and the visual data are subjected to spatiotemporal consistency processing to determine a dynamic target set and a first feature set, wherein the first feature set includes a spatiotemporally aligned first radar feature set and a first visual feature set; Based on the set of dynamic targets, the first feature set is updated, and the current dynamic scene and the fusion weights under the current dynamic scene are determined; Using an attention mechanism, the first radar feature set and the first visual feature set in the updated first feature set are weighted and fused according to the fusion weights to obtain fused features. The updated first feature set is then fused and filtered to obtain the second feature set.

3. The method according to claim 2, characterized in that, The step of updating the first feature set based on the dynamic target object set and determining the current dynamic scene and the fusion weights under the current dynamic scene includes: Based on the set of dynamic target objects and the first feature set, a set of dynamic attributes is determined, and the first feature set is updated according to the set of dynamic attributes. Based on the set of dynamic target objects, the current dynamic scene is determined, and the fusion weight under the current dynamic scene is determined according to the set of dynamic attributes.

4. The method according to claim 2, characterized in that, The step of performing spatiotemporal consistency processing on the radar data and the visual data to determine the dynamic target set and the first feature set includes: Using a time interpolation algorithm, a Kalman filter is used to align the radar data and the visual data in time. By using extrinsic parameter calibration technology, the time-aligned radar data and time-aligned visual data are unified into the vehicle coordinate system to obtain spatiotemporally aligned radar data and spatiotemporally aligned visual data. The visual feature points corresponding to the spatiotemporally aligned radar data and the spatiotemporally aligned visual data are determined, and the motion vectors of the visual feature points are calculated using the optical flow method to determine the dynamic target set. The spatiotemporally aligned radar data and spatiotemporally aligned visual data of multiple consecutive frames are input into the spatiotemporal encoder, and time consistency processing is performed to obtain the first radar feature set and the first visual feature set. Wherein, the first radar feature set is the radar data with spatiotemporal consistency, and the first visual feature set is the visual data with spatiotemporal consistency.

5. The method according to claim 4, characterized in that, The step of using a time interpolation algorithm and a Kalman filter to perform time alignment of the radar data and the visual data includes: Based on the sampling frequencies of the radar data and the visual data, a unified set of timestamps for interpolation is determined; The radar data and the visual data are time-aligned by interpolation, wherein the radar data is real-time data or legacy data, and the visual data is real-time data or legacy data. If the current timestamp belongs to the set of timestamps, use a Kalman filter to predict the estimated value of the radar data and the estimated value of the visual data corresponding to the current timestamp, and update the residual data in the radar data to the estimated value of the radar data, and update the residual data in the visual data to the estimated value of the visual data.

6. The method according to claim 3, characterized in that, The dynamic attributes corresponding to the first radar feature set and the first visual feature set each include at least one of the following: basic motion attributes of multiple dynamic targets, spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes; The spatiotemporal correlation attributes include: motion trajectory fitting degree and / or life cycle state; the perception enhancement attributes include: multi-sensor consistency and / or dynamic confidence. The determination of the dynamic attribute set based on the dynamic target set and the first feature set includes at least one of the following: Based on the set of dynamic targets, the basic motion attributes in the set of dynamic attributes are determined according to the difference between the current frame data and the previous frame data in the first feature set. Based on the persistent state of the target object in the first feature set of the spatiotemporal encoder, the life cycle state in the dynamic attribute set is determined. By comparing the motion trajectories in the first radar feature set and the first visual feature set, the fitting degree of the motion trajectory in the dynamic attribute set is determined. Based on the first feature set, determine the interaction behavior between the vehicle and surrounding dynamic targets, and determine the interaction behavior features in the dynamic attribute set; The position information of the same target in the first radar feature set and the first visual feature set are compared, and the consistency of the multiple sensors in the dynamic attribute set is determined based on the deviation of the position information. The dynamic confidence level in the dynamic attribute set is determined based on the multi-sensor consistency and the stable state of the first feature set over multiple consecutive frames.

7. The method according to claim 3, characterized in that, The fusion weights include radar weights and visual weights; The step of determining the current dynamic scene based on the set of dynamic targets and determining the fusion weights under the current dynamic scene according to the set of dynamic attributes includes: Based on the set of dynamic targets, the density of dynamic targets around the vehicle is calculated, and based on the density of the dynamic targets, the current dynamic scene is determined; The spatiotemporal correlation attributes, interactive behavior features, and perception enhancement attributes in the dynamic attributes corresponding to the first radar feature set and the first visual feature set are weighted and summed according to a preset ratio to determine the radar weight and the visual weight, respectively. The preset ratio is related to the current dynamic scene. Based on the basic motion attributes in the dynamic attributes corresponding to the first radar feature set and the first visual feature set, the interaction parameters between the vehicle and the surrounding dynamic targets are determined; if the interaction parameters corresponding to the first radar feature set are outside the safe threshold range, or if the interaction parameters corresponding to the first visual feature set are outside the safe threshold range, the radar weights and the visual weights are updated.

8. The method according to claim 2, characterized in that, The second feature set includes a second radar feature set and a second visual feature set; The process involves using an attention mechanism to perform weighted fusion of the first radar feature set and the first visual feature set in the updated first feature set, based on the fusion weights, to obtain fused features. Then, the updated first feature set is further fused and filtered to obtain a second feature set, including: Based on the fusion weights, the first radar feature set and the first visual feature set in the updated first feature set are weighted and concatenated to obtain the fusion input embedding matrix. The fused features are obtained by performing feature fusion on the fused input embedding matrix through an attention mechanism; The features that cannot be fused in the updated first feature set are removed respectively to obtain the second radar feature set and the second visual feature set.

9. The method according to claim 8, characterized in that, The step of performing feature fusion on the fused input embedding matrix through an attention mechanism to determine the fused features includes: Using the accuracy of the fusion perception task as the loss function, the parameter matrix of the attention mechanism is iteratively updated through the backpropagation algorithm until the attention weights are adapted to the confidence of radar sensors and vision sensors in different scenarios. The fusion perception task includes target detection task and target tracking task. The fused input embedding matrix is ​​mapped to query features, key features, and value features respectively through the parameter matrix, and the attention weights are calculated through the attention mechanism and the values ​​are weighted and summed to obtain the fused features.

10. The method according to any one of claims 2-9, characterized in that, The method further includes, after obtaining fused features by weighting and fusing the first radar feature set and the first visual feature set in the updated first feature set using an attention mechanism and according to the fusion weights, and then performing fusion filtering on the updated first feature set to obtain the second feature set, the method further includes: Based on the dynamic target set, the second radar feature set and the second visual feature set in the second feature set are classified as dynamic targets respectively, and the category information of the dynamic targets is obtained. Based on the category information, a second radar feature set and a second visual feature set with common dynamic targets are extracted, and an initial third radar feature set and an initial third visual feature set are determined. Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the initial third visual feature set, and update the dynamic confidence of multiple dynamic targets based on the category relationship matrix; In the initial third radar feature set and the initial third visual feature set, dynamic targets with a dynamic confidence level less than or equal to a preset threshold are discarded, and the target detection information, the initial third radar feature set, and the initial third visual feature set are updated. Output at least one of the following: the fusion weight, the updated target detection information, the updated third radar feature set, and the updated third visual feature set; Wherein, the category information corresponding to the common dynamic target in the second radar feature set and the second visual feature set is the same target category.

11. The method according to claim 10, characterized in that, The step of determining the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the initial third visual feature set, and updating the dynamic confidence of the multiple dynamic targets based on the category relationship matrix, includes: Determine the category relationship matrix between the target detection information corresponding to the initial third radar feature set and the target detection information corresponding to the initial third visual feature set; In the third radar feature set and the third visual feature set, the target objects corresponding to the maximum similarity in each similarity are removed, the third radar feature set and the third visual feature set are updated, and the dynamic confidence of the target object is updated according to the maximum similarity. Repeat the steps of determining the category relationship matrix between the target detection information corresponding to the third radar feature set and the target detection information corresponding to the third visual feature set, removing the target object corresponding to the maximum similarity in each similarity in the third radar feature set and the third visual feature set, updating the third radar feature set and the third visual feature set, and updating the dynamic confidence of the target object according to the maximum similarity, until the similarity of each dynamic target object in the category relationship matrix is ​​determined, and the dynamic confidence of each dynamic target object is updated according to the similarity of each dynamic target object.

12. A multimodal fusion device, characterized in that, Applied to vehicles, including: The data acquisition module is used to acquire radar data and visual data; The processing module is used to perform spatiotemporal consistency processing and weighted fusion on the radar data and the visual data to obtain fused features.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the multimodal fusion method as described in any one of claims 1-11.

14. A vehicle, characterized in that, Including the electronic device as described in claim 13.