Dynamic object perception system and method based on multi-source sensor information fusion

CN122157240BActive Publication Date: 2026-08-21ZHEJIANG FUBAO INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610612148.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-08-21
Estimated Expiration
2046-05-07

AI Technical Summary

Technical Problem

现有融合框架难以实时、准确地评估各传感器数据流在此类场景下的可信度,无法对因环境或自身局限性而引入的感知不确定性进行有效量化,进而导致无法动态地调整各信息源在最终融合决策中的贡献比重

Benefits of technology

[0008]与现有技术相比,本申请提供的一种基于多源传感器信息融合的动态物体感知系统及方法,其通过并行处理RGB-D图像、3D点云及毫米波雷达等多源异构数据,并引入一个跨模态不一致性度量机制,通过主动分析和量化不同传感器特征之间的差异与冲突,来实时评估各传感模态在当前环境下的感知置信度。基于此,驱动一个自适应注意力机制动态生成各模态的融合权重,最终实现对各模态特征向量的智能加权融合,输出一个高鲁棒性的融合目标状态。这样,能够智能地抑制低质量或冲突的传感器信息,增强高质量信息的贡献,从而在复杂多变的环境中实现更为鲁棒和精确的动态物体感知。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157240B_ABST
    Figure CN122157240B_ABST
Patent Text Reader

Abstract

The application discloses a dynamic object perception system and method based on multi-source sensor information fusion, relates to the field of dynamic object perception, and through parallel processing of multi-source heterogeneous data such as RGB-D images, 3D point clouds and millimeter wave radars, and introduction of a cross-modal inconsistency measurement mechanism, differences and conflicts between different sensor features are actively analyzed and quantified to evaluate the perception confidence of each sensing mode under the current environment in real time. Based on this, an adaptive attention mechanism is driven to dynamically generate the fusion weight of each mode, and finally, intelligent weighted fusion of each mode feature vector is realized, and a high-robustness fusion target state is output. In this way, low-quality or conflicting sensor information can be intelligently suppressed, and the contribution of high-quality information can be enhanced, so that more robust and accurate dynamic object perception can be realized in a complex and changeable environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of dynamic object perception, and more specifically, to a dynamic object perception system and method based on multi-source sensor information fusion. Background Technology

[0002] With the increasing global trend of population aging, the demand for elderly care services is growing, and elderly care robots, as an important technological means to address labor shortages and improve the quality of life for the elderly, are receiving increasing attention. Elderly care robots typically need to coexist with people and provide services in unstructured environments such as homes, which places extremely high demands on their environmental perception capabilities, especially the accurate and reliable perception of dynamic objects. To ensure the safety and smoothness of the robot during movement and interaction, it must be able to accurately identify and track dynamic targets in the environment, such as the elderly, pets, and mobile assistive devices, and predict their behavioral intentions in real time. Therefore, building a robust and efficient dynamic object perception system is a core technological prerequisite for realizing safe and intelligent elderly care services.

[0003] In existing technologies, a common strategy is to combine the rich texture and color information provided by a camera (RGB-D) with the precise three-dimensional spatial information provided by a LiDAR sensor. This fusion, to some extent, compensates for the limitations of a single sensor and improves perception accuracy in ideal environments. However, these fusion methods often employ fixed or simple rule-based weighting mechanisms, lacking adaptability to sensor states in dynamically changing environments. For example, when lighting conditions change drastically (such as moving from a bright living room to a dimly lit bedroom), the information quality of visual sensors deteriorates significantly, or when the target object's surface is highly reflective or transparent, LiDAR data may exhibit severe bias. Existing fusion frameworks struggle to assess the reliability of each sensor's data stream in such scenarios in real time and accurately, failing to effectively quantify the perception uncertainties introduced by environmental or inherent limitations. Consequently, they cannot dynamically adjust the contribution weight of each information source in the final fusion decision. This "static" or "passive" fusion approach significantly reduces the stability and reliability of the perception results when facing complex and ever-changing real-world home environments, making it difficult to meet the stringent safety standards required in elderly care settings.

[0004] Therefore, there is an urgent need for an optimized dynamic object perception system and method based on multi-source sensor information fusion. Summary of the Invention

[0005] This application is made in order to solve the above-mentioned technical problems.

[0006] According to one aspect of this application, a dynamic object perception method based on multi-source sensor information fusion is provided, comprising: Acquire RGB-D images, 3D point clouds, and millimeter-wave radar data; Modality-specific feature extraction is performed on RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors. Cross-modal inconsistency measures were performed on visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors to obtain an inconsistency measure matrix; Based on the inconsistency metric matrix, an adaptive attention weight metric is performed on the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector to obtain the attention weight vector. Based on the attention weight vector, the visual feature vector, the LiDAR feature vector, and the millimeter-wave radar feature vector are weighted and fused, and the final state output is obtained to obtain the fused target state.

[0007] According to another aspect of this application, a dynamic object perception system based on multi-source sensor information fusion is provided, comprising: A multimodal data acquisition module is used to acquire RGB-D images, 3D point clouds, and millimeter-wave radar data; The modality-specific feature extraction module is used to extract modality-specific features from RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors. The cross-modal inconsistency measurement module is used to perform cross-modal inconsistency measurement on visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors to obtain an inconsistency measurement matrix; The adaptive attention weight measurement module is used to perform adaptive attention weight measurement on visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors based on the inconsistency measurement matrix to obtain attention weight vectors. The fusion and state output module is used to perform weighted fusion of visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors based on attention weight vectors and output the final state to obtain the fused target state.

[0008] Compared with existing technologies, this application provides a dynamic object perception system and method based on multi-source sensor information fusion. It processes heterogeneous data from multiple sources, such as RGB-D images, 3D point clouds, and millimeter-wave radar, in parallel and introduces a cross-modal inconsistency measurement mechanism. By actively analyzing and quantifying the differences and conflicts between features of different sensors, it evaluates the perception confidence of each sensing modality in the current environment in real time. Based on this, it drives an adaptive attention mechanism to dynamically generate fusion weights for each modality, ultimately achieving intelligent weighted fusion of feature vectors from each modality and outputting a highly robust fused target state. This intelligently suppresses low-quality or conflicting sensor information and enhances the contribution of high-quality information, thereby achieving more robust and accurate dynamic object perception in complex and changing environments. Attached Figure Description

[0009] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 This is a flowchart of a dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application.

[0011] Figure 2 This is a data flow diagram of a dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application.

[0012] Figure 3 This is a flowchart of sub-step S1 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application.

[0013] Figure 4 This is a flowchart of sub-step S2 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application.

[0014] Figure 5 This is a flowchart of sub-step S3 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application.

[0015] Figure 6 This is a flowchart of sub-step S35 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application.

[0016] Figure 7 This is a flowchart of sub-step S5 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application.

[0017] Figure 8 This is a block diagram of a dynamic object perception system based on multi-source sensor information fusion according to an embodiment of this application. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] To address the problems mentioned above, this application proposes a dynamic object perception method based on multi-source sensor information fusion. Figure 1 This is a flowchart of a dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application. Figure 2 This is a data flow diagram of a dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application. For example... Figure 1 and Figure 2 As shown, the dynamic object perception method based on multi-source sensor information fusion includes the following steps: S1, acquiring RGB-D images, 3D point clouds, and millimeter-wave radar data; S2, performing modal-specific feature extraction on the RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors; S3, performing cross-modal inconsistency measurement on the visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors to obtain an inconsistency measurement matrix; S4, based on the inconsistency measurement matrix, performing adaptive attention weight measurement on the visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors to obtain an attention weight vector; S5, based on the attention weight vector, performing weighted fusion and final state output on the visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors to obtain the fused target state.

[0020] In the aforementioned dynamic object perception method based on multi-source sensor information fusion, step S1 involves acquiring RGB-D images, 3D point clouds, and millimeter-wave radar data. It should be understood that due to the functional limitations of a single sensor, RGB images can only present color and texture information in a two-dimensional plane, failing to reflect the spatial depth relationship of the target. LiDAR signals are easily blocked or scattered in environments with strong light or dust, leading to data loss. While millimeter-wave radar possesses anti-interference capabilities, its low spatial resolution makes it difficult to accurately depict the morphological details of the target. Using any one of these three devices alone cannot achieve comprehensive perception of dynamic objects in the environment. Therefore, this application deploys three types of devices—RGB-D cameras, 3D LiDAR, and millimeter-wave radar—to simultaneously collect corresponding data, thereby integrating three core types of information: color and texture, spatial depth, and target motion state, to construct a multi-dimensional environmental perception data system. This ensures that complete and accurate environmental data can be acquired in different home scenarios such as living rooms and bedrooms, and under different lighting conditions such as daytime and nighttime, providing reliable data support for subsequent analysis and avoiding perception interruptions caused by the failure of a single sensor.

[0021] Specifically, in one possible embodiment, step S1 is implemented as follows: First, based on the environmental perception range and accuracy requirements, the installation positions, angles, and acquisition parameters of the three types of sensors are determined to ensure that the sensor acquisition range can cover the target monitoring area without significant blind spots. Next, RGB-D cameras are installed at the front and sides of the mobile platform, and the camera sampling frequency is set to match the real-time perception requirements. Environmental color information is acquired through the camera image sensors, and the distance information between the target and the camera is obtained using the built-in infrared ranging module. The color image and depth image are registered to form RGB-D image data. Then, a 3D LiDAR is installed at the center of the top of the mobile platform. The LiDAR is controlled to rotate 360 ​​degrees at a fixed speed, emitting laser pulses to the surrounding environment and receiving laser signals reflected from the target. The three-dimensional spatial coordinates of the target are calculated based on the signal transmission time, generating 3D point cloud data containing target position information. Finally, millimeter-wave radars are installed at the front and rear ends of the mobile platform, respectively. The radar operating frequency and detection distance threshold are set, and millimeter-wave signals are emitted through the radar antenna. The echo signals reflected from the target are received, and the target's distance, speed, and other motion parameters are calculated based on the frequency difference between the echo signal and the emitted signal to obtain millimeter-wave radar data.

[0022] In the aforementioned dynamic object perception method based on multi-source sensor information fusion, step S2 involves extracting modal-specific features from RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors. It should be understood that due to the fundamental differences in the modal structures of RGB-D images, 3D point clouds, and millimeter-wave radar data—RGB-D images are stored as pixel matrices, with core information concentrated on color distribution and depth gradient; 3D point clouds exist as a set of discrete points, with key information being the spatial relationships between points and the target surface morphology; and millimeter-wave radar data is presented as signal parameters, with core content being the target's motion characteristics and signal strength—directly fusing these three types of raw data would lead to complex fusion logic due to inconsistent data structures, easily introducing redundant information and failing to effectively extract the core value of each modality. Therefore, this application further targets the modal characteristics of the three types of data, extracting visual features from RGB-D images, extracting LiDAR features from 3D point clouds, and obtaining millimeter-wave radar features from millimeter-wave radar data, respectively. This is to accurately mine the core representation information of each modal data, remove invalid and redundant data, retain the key information of each modal data, avoid fusion errors caused by differences in data structure, reduce the computational complexity of subsequent multimodal data fusion, and improve the accuracy and stability of the fusion results.

[0023] In particular, in one specific embodiment, Figure 3 This is a flowchart of sub-step S2 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application. Figure 3 As shown, step S2 includes: S21, performing multimodal parallel initial detection on RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain a visual candidate object set, a lidar candidate object cluster set, and a millimeter-wave radar candidate point set; S22, based on the trajectory database of the previous moment, performing intramodal data association and state estimation on the visual candidate object set, lidar candidate object cluster set, and millimeter-wave radar candidate point set to obtain the trajectory database updated at the current moment; S23, performing modal-specific feature encapsulation on the trajectory database updated at the current moment to obtain visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors.

[0024] Specifically, step S21 involves performing multimodal parallel initial detection on RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain a visual candidate object set, a lidar candidate object cluster set, and a millimeter-wave radar candidate point set. It should be understood that because RGB-D images, 3D point clouds, and millimeter-wave radar data have different modal attributes, the dynamic object information they carry differs. Serial detection can easily lead to data lag, and single-modal detection is prone to missing or misjudging targets. Therefore, this application further designs adapted initial detection algorithms for each of the three types of data, performing detection operations in parallel. This accurately filters out candidate sets that may contain dynamic objects from each modal data, ensuring that core information of each modality is not lost and that detection efficiency matches the real-time perception requirements. This avoids the time loss of serial detection, while covering the three key information dimensions of dynamic objects: appearance, space, and motion. It reduces the blind spots of single-modal detection, providing a comprehensive and accurate candidate basis for subsequent data association, and ensuring the real-time performance and initial accuracy of dynamic object perception.

[0025] Specifically, in one possible embodiment, step S21 is implemented as follows: For RGB-D images, a pre-trained and domain-fine-tuned deep learning object detection network (e.g., YOLO series, Faster R-CNN, etc.) is used. This network is first pre-trained on a large public dataset (such as the COCO dataset) to learn general object features, and then fine-tuned using a domain-specific dataset containing specific targets such as home environments, elderly people, wheelchairs, and mobility aids, enabling it to accurately identify dynamic objects in elderly care scenarios. During detection, the feature map obtained by fusing the RGB image and the depth image is input. Through anchor box matching and classification regression, regions with confidence scores higher than a set threshold (which can be set to 0.5) are selected, marked as visual candidate objects, and a visual candidate object set is formed. For 3D point clouds, the data volume is first compressed through voxelization, and then a density-based clustering algorithm is used to aggregate points that are close in space into clusters. Clusters whose sizes do not conform to the range of dynamic objects are removed, resulting in a set of LiDAR candidate object clusters. For millimeter-wave radar data, the original signal is first denoised and peak extracted, and detection points whose signal strength and motion parameters match the characteristics of dynamic objects are selected to form a set of candidate points for millimeter-wave radar.

[0026] Specifically, in step S22, based on the trajectory database from the previous moment, intramodal data association and state estimation are performed on the visual candidate object set, the lidar candidate object cluster set, and the millimeter-wave radar candidate point set to obtain the trajectory database updated at the current moment. It should be understood that since multimodal parallel initial detection can only acquire candidate objects at the current moment, lacking temporal association with targets at historical moments, the same dynamic object is prone to being repeatedly marked or the trajectory is interrupted, making it impossible to form a continuous target tracking trajectory. Therefore, this application further calls the trajectory database stored at the previous moment, performing data association and state estimation for each modal candidate set to match the current moment's candidate objects with historical trajectories, determine the temporal correspondence of the same target, and update the target's position, velocity, and other state parameters to establish the temporal motion trajectory of the dynamic object. This avoids trajectory breaks or mismatches caused by the isolation of candidate objects, ensures the continuity and consistency of target trajectories in each modality, provides temporally stable trajectory data for subsequent cross-modal fusion, and improves the reliability of dynamic object tracking.

[0027] In particular, in one specific embodiment, Figure 4 This is a flowchart of sub-step S22 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application. Figure 4 As shown, step S22 includes: S221, extracting a set of lidar trajectories from the trajectory database of the previous moment; S222, inputting each lidar trajectory in the lidar trajectory set into a constant-speed Kalman filter for state prediction to obtain the predicted state and prediction covariance of all lidar trajectories at the current moment; S223, calculating the gated cost matrix between the predicted state and prediction covariance of all lidar trajectories at the current moment and the lidar candidate object cluster set; S224, determining successfully matched trajectory-cluster pairs based on the gated cost matrix; S225, updating the trajectory based on the successfully matched trajectory-cluster pairs to obtain the updated trajectory database at the current moment.

[0028] More specifically, step S221 involves extracting the LiDAR trajectory set from the trajectory database of the previous moment. It should be understood that since the current LiDAR candidate object cluster only contains static spatial information of the current environment and lacks association with the historical dynamic object trajectories, it is impossible to determine whether the objects in the cluster are continuously tracked targets, easily leading to repeated target identification or trajectory breakage. Therefore, this application further accesses the stored trajectory database of the previous moment, filters and extracts trajectory data related to the LiDAR mode, thereby obtaining the LiDAR trajectory information of historical dynamic objects, including trajectory identifiers, historical positions, motion parameters, and other core content. This provides basic data support for subsequently establishing an association between the current candidate cluster and historical targets, avoiding discontinuous target tracking caused by relying solely on current data, ensuring the temporal continuity of dynamic object trajectories, and improving the stability of target tracking.

[0029] Specifically, in one possible embodiment, step S221 is implemented as follows: First, the trajectory database access interface is activated, and the complete trajectory database from the previous moment is loaded. This database categorizes and stores trajectories according to modal type, including three modal trajectories: LiDAR, vision, and millimeter-wave radar. Then, through the modal identification filtering module, all trajectory data labeled as LiDAR modality are located and extracted to form a LiDAR trajectory set. Finally, the extracted trajectory set is validated to confirm that each trajectory contains complete historical location coordinates, movement speed, trajectory confidence level, and a unique trajectory ID, ensuring the integrity and availability of data in subsequent association processes.

[0030] More specifically, step S222 involves inputting each lidar trajectory in the lidar trajectory set into a constant-speed Kalman filter for state prediction to obtain the predicted state and prediction covariance of all lidar trajectories at the current moment. It should be understood that since the lidar trajectory at the previous moment only reflects the target state at a historical time, while the target's position and motion parameters change over time, directly matching the historical state with the current candidate cluster would result in significant deviations and would fail to accurately pinpoint the possible state range of the target at the current moment. Therefore, this application further inputs each trajectory in the lidar trajectory set individually into a constant-speed Kalman filter, and calculates the target motion trend through the filter's state prediction model, thereby obtaining the predicted position, predicted velocity, and other state parameters of each trajectory at the current moment, as well as the prediction covariance reflecting the prediction uncertainty. This provides an accurate prediction benchmark for matching the current candidate cluster with historical trajectories, narrowing the matching search range, reducing the probability of mismatch, and improving the accuracy of data association.

[0031] Specifically, in one possible embodiment, step S222 is implemented as follows: First, the state equation and observation equation of the constant-speed Kalman filter are initialized, the state vector is set to include position and velocity parameters, and the process noise and observation noise matrices are determined. Then, each trajectory in the lidar trajectory set is traversed, and the final state (position, velocity) and covariance matrix of the trajectory at the previous moment are extracted as the initial input of the filter. Subsequently, the prediction phase of the filter is started, the predicted state vector of the target at the current moment is calculated according to the constant-speed motion model, and the corresponding prediction covariance matrix is ​​calculated through the error propagation formula. Finally, the predicted state and prediction covariance of each trajectory are associated and stored according to the trajectory ID to form a prediction result set.

[0032] More specifically, step S223 involves calculating the gated cost matrix between the predicted state and prediction covariance of all LiDAR trajectories at the current moment and the LiDAR candidate clusters. It should be understood that since the current LiDAR candidate clusters contain multiple candidate clusters, and the predicted state of each trajectory is uncertain, directly calculating the matching cost between all trajectories and all candidate clusters would result in a large amount of invalid computation and might mistakenly associate trajectories that are too far apart with candidate clusters. Therefore, this application further filters out trajectory-cluster combinations that exceed the reasonable matching range through a gating mechanism, and then calculates the matching cost of the remaining valid combinations. This generates a gated cost matrix containing only valid matching combinations, with matrix elements representing the matching similarity between trajectories and candidate clusters. This significantly reduces the amount of invalid matching computation, lowers computational complexity, and eliminates obviously unreasonable matching combinations, laying the foundation for subsequent accurate matching.

[0033] Specifically, in one possible embodiment, step S223 is implemented as follows: First, a gating threshold is determined, calculated based on the prediction covariance and a preset confidence level, to define a reasonable range for the trajectory prediction state. Then, the prediction state and prediction covariance of each trajectory are iterated, and the spatial distance (e.g., centroid distance) between the trajectory and each candidate cluster in the lidar candidate object cluster set is calculated. It is then determined whether the distance is within the gating threshold range, filtering out candidate clusters that exceed the threshold. Subsequently, for the remaining valid trajectory-cluster combinations, Mahalanobis distance is used to calculate the matching cost. The Mahalanobis distance combined with the prediction covariance reflects the reliability of the matching; a smaller cost value indicates a higher matching degree. Finally, the cost values ​​of all valid combinations are arranged in the order of trajectory rows and candidate cluster columns to form a gating cost matrix. Positions that fail the gating are marked as invalid values.

[0034] More specifically, step S224 determines successfully matched trajectory-cluster pairs based on the gated cost matrix. It should be understood that since the gated cost matrix only provides matching cost information between trajectories and candidate clusters, without specifying the optimal matching relationship, without a matching decision, it is impossible to determine which trajectory corresponds to which candidate cluster, easily leading to one-to-many or many-to-one mismatch problems. Therefore, this application further employs an optimal matching algorithm to solve the gated cost matrix, thereby selecting the trajectory-cluster combination with the lowest cost that satisfies the matching criteria from the cost matrix, determining the unique candidate cluster corresponding to each trajectory or classifying it as unmatched. This establishes a precise correspondence between historical trajectories and current candidate clusters, ensuring that the trajectory of each dynamic target continues at the current moment, avoiding target identity confusion caused by mismatches, and improving the accuracy of trajectory tracking.

[0035] Specifically, in one possible embodiment, step S224 is implemented as follows: First, matching criteria are set, including a cost threshold (matches below the threshold are considered valid) and a uniqueness criterion (a trajectory matches only one candidate cluster, and a candidate cluster matches only one trajectory). Then, the gated cost matrix is ​​input into the Hungarian algorithm, which solves for the matching scheme that minimizes the total matching cost through the optimal allocation of matrix rows and columns. Subsequently, the matching results are verified against the preset criteria, eliminating matching combinations with cost values ​​higher than the threshold, and retaining trajectory-cluster combinations that meet the criteria. Finally, a list of successfully matched trajectory-cluster pairs is output, with each record containing a trajectory ID, candidate cluster identifier, and matching cost value, while also marking unmatched trajectories and unmatched candidate clusters.

[0036] More specifically, step S225 involves updating the trajectory based on the successfully matched trajectory-cluster pairs to obtain the updated trajectory database for the current moment. It should be understood that since the successfully matched trajectory-cluster pairs only determine the correspondence, the state parameters of the historical trajectories remain predicted values ​​and do not incorporate the actual observation information of the current candidate clusters. If the trajectory state is not updated, the deviation between the trajectory and the actual target position will gradually increase. Therefore, this application further utilizes the observation data of the successfully matched candidate clusters to correct and update the state parameters of the corresponding historical trajectories, thereby combining the predicted state with the actual observations to obtain the accurate state of the target at the current moment. It also integrates the processing results of unmatched trajectories and candidate clusters to form an updated trajectory database. This ensures that the trajectory state is consistent with the actual movement of the target, maintains the continuity and accuracy of the trajectory, and provides reliable historical data for trajectory prediction and association at the next moment.

[0037] Specifically, in one possible embodiment, step S225 is implemented as follows: First, the list of successfully matched trajectory-cluster pairs is traversed, and the actual observation information of each candidate cluster is extracted, such as the centroid position and size parameters of the cluster. Then, the observation information is input into the constant-speed Kalman filter of the corresponding trajectory, and the filter update process is started. The predicted state is corrected by the observation residuals to obtain the optimal estimated state (position, velocity) of the trajectory at the current moment and the updated covariance matrix, and the confidence of the trajectory is updated. Subsequently, for unmatched trajectories, it is determined whether to delete them based on the number of times the trajectory has been unmatched. If the number exceeds the threshold, they are deleted. For unmatched candidate clusters, they are identified as new targets and new lidar trajectories are created and assigned unique trajectory IDs. Finally, the updated matched trajectories, the retained unmatched trajectories, and the newly created trajectories are integrated, classified and stored by mode, to form the trajectory database updated at the current moment.

[0038] Specifically, step S23 involves encapsulating modality-specific features in the updated trajectory database at the current moment to obtain visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors. It should be understood that because the updated trajectory database contains diverse forms of trajectory data across modalities, including various types of raw information such as position, features, and state, and has inconsistent structures, directly using it for cross-modal fusion would lead to complex fusion logic and easily introduce redundant information, making it impossible to efficiently extract the core value of each modality's trajectory. Therefore, this application further refines and standardizes the trajectory data for each modality to transform the dispersed trajectory information into feature vectors with fixed dimensions and clear physical meaning, ensuring that the feature structures of each modality are unified and retaining key information. This eliminates the structural differences in trajectory data across modalities, reduces the computational complexity of subsequent cross-modal fusion, and strengthens the core representation capabilities of each modality through feature encapsulation, enabling the fusion process to more accurately utilize the advantages of each modality and improve the overall accuracy of dynamic object perception.

[0039] Specifically, in one possible embodiment, step S23 is implemented as follows: For visual trajectories, the appearance features (such as texture and color distribution), two-dimensional position coordinates, motion direction, and confidence level of visual candidate objects are extracted from the trajectory database. After standardization, these information are combined in a preset dimensional order to form a visual feature vector. For LiDAR trajectories, the three-dimensional centroid coordinates, point cloud density within the cluster, three-dimensional bounding box size, and spatial attitude parameters of the candidate object clusters are extracted and normalized before being encapsulated into LiDAR feature vectors. For millimeter-wave radar trajectories, the motion velocity, acceleration, signal strength, and distance change rate of candidate points are extracted. After standardization of these motion-related parameters, they are combined to form millimeter-wave radar feature vectors. All three types of vectors use the same data format and dimensional length to meet the subsequent fusion requirements.

[0040] In the aforementioned dynamic object perception method based on multi-source sensor information fusion, step S3 involves performing a cross-modal inconsistency measurement on the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector to obtain an inconsistency measurement matrix. It should be understood that since the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector originate from different modal data, there are fundamental differences in the data representation logic between the modalities. Direct fusion is prone to distortion of the fusion result due to the failure to quantify these differences. Therefore, this application further selects an appropriate measurement method to calculate the degree of inconsistency between pairs of modalities for the modal attributes of the three types of feature vectors, generating an inconsistency measurement matrix. This accurately quantifies the magnitude of differences in the description of the same target by different modalities, providing data support for subsequent dynamic adjustment of fusion weights. This avoids fusion deviations caused by ignoring modal differences, ensuring that the fusion process can allocate modal weights according to the degree of difference, improving the reliability and accuracy of dynamic object perception results, and adapting to perception needs in complex environments.

[0041] In particular, in one specific embodiment, Figure 5 This is a flowchart of sub-step S3 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application. Figure 5 As shown, step S3 includes: S31, performing validity checks and state extraction on the visual feature vector, lidar feature vector, and millimeter-wave radar feature vector respectively to obtain a set of modal target structured coding vectors; S32, calculating the inconsistency between any two modal target structured coding vectors in the set of modal target structured coding vectors to obtain an inconsistency metric matrix.

[0042] Specifically, step S31 involves performing validity checks and state extraction on the visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors to obtain a set of structured encoded vectors for the modal target. It should be understood that during the acquisition and extraction process, visual, LiDAR, and millimeter-wave radar feature vectors may generate noise or invalid data due to environmental interference (such as sudden changes in illumination or object occlusion), and the original feature vector structures are scattered and contain redundant information. Directly using them for subsequent processing can easily lead to computational redundancy and perception errors. Therefore, this application further performs validity checks on the three types of feature vectors, removes invalid data, and extracts key information representing the core state of the target, encoding it in a unified format to form a structured vector set. This is used to screen high-quality feature data, extract key target states, and achieve standardization and structuring of feature vectors. This removes noise and redundant information from the features, improves the quality of feature data, and eliminates inter-modal structural differences through a unified encoding format, reducing computational complexity for subsequent cross-modal inconsistency measurement and fusion processing, and ensuring the efficiency and accuracy of the perception process.

[0043] Specifically, in one possible embodiment, step S31 is implemented as follows: First, validity check criteria are set. Visual feature vectors check whether the confidence level of appearance features is higher than a preset threshold and there are no missing pixels. LiDAR feature vectors check whether the three-dimensional spatial parameters are within a physically reasonable range (such as target size and distance) and the point cloud clusters are complete. Millimeter-wave radar feature vectors check whether the motion parameters (velocity and acceleration) are stable and without abnormal jumps. Then, the three types of feature vectors of each target are checked one by one, and invalid features that fail the check are eliminated, while valid features are retained. Subsequently, core state information is extracted from the valid feature vectors. Visual features extract the target's two-dimensional bounding box and texture feature values. LiDAR features extract the three-dimensional centroid coordinates and bounding box size. Millimeter-wave radar features extract the motion speed and signal strength. Finally, the extracted state information is encoded according to a preset dimensional order and data format to form a fixed-length modal target structured encoding vector. The three types of structured vectors of all targets are integrated into a modal target structured encoding vector set.

[0044] Specifically, step S32 involves calculating the inconsistency between any two modal target structured encoding vectors in the modal target structured encoding vector set to obtain an inconsistency metric matrix. It should be understood that although the modal target structured encoding vectors have achieved a unified format, different modal vectors still characterize the target based on their respective modal characteristics, such as visual focusing on appearance, lidar focusing on space, and millimeter-wave radar focusing on motion. There are still descriptive differences between any two modalities. If these differences are not quantified, subsequent fusion cannot accurately balance the contributions of each modality, easily leading to a problem where a certain modality bias dominates the fusion result. Therefore, this application further traverses the modal target structured encoding vector set, calculates the degree of inconsistency between any two different modal vectors, and organizes them into an inconsistency metric matrix. This comprehensively covers the differences between all modal pairs, providing fine-grained evidence of differences for fusion decisions. In a specific example of this application, step S32 includes: calculating the inconsistency between any two modal target structured encoding vectors in the modal target structured encoding vector set using the following formula:

[0045] in, For the first 3D position estimation in modal target structured encoding vector For the first 3D position estimation in modal target structured encoding vector For the first Standard deviation of scalar position estimates in modal target structured encoding vectors For the first Standard deviation of scalar position estimates in modal target structured encoding vectors For smoothing terms, Indicates the first and the Inconsistencies between modal target structured encoding vectors. This allows for the accurate capture of deviations in the description of the same target by any two modalities, enabling subsequent fusion strategies to be adjusted based on the differences between specific modal pairs. This avoids the loss of fusion accuracy caused by treating differences uniformly, further improving the robustness of dynamic object perception.

[0046] In the aforementioned dynamic object perception method based on multi-source sensor information fusion, step S4 involves adaptively measuring the attention weights of visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors based on an inconsistency metric matrix to obtain an attention weight vector. It should be understood that the reliability of visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors dynamically changes with the environment. For example, sudden changes in illumination can reduce the effectiveness of visual features, and object occlusion can weaken the accuracy of LiDAR features. Since the inconsistency metric matrix quantifies the degree of difference between each modality, using fixed weights for fusion cannot adapt to the real-time changes in modal reliability, easily leading to interference from low-reliability modes in the fusion result. Therefore, this application further uses the inconsistency metric matrix as a basis to perform adaptive attention weight measurement on the three types of feature vectors, thereby dynamically allocating weight ratios according to the differences between modalities, so that modes with high consistency and strong reliability occupy higher weights in the fusion. This avoids the limitations of fixed weights, allowing the fusion process to accurately match the actual performance of each modality in the current environment, significantly improving the accuracy and robustness of the multimodal fusion result, and ensuring that dynamic object perception can still stably output reliable results in complex environments.

[0047] In particular, in one specific embodiment, Figure 6 This is a flowchart of sub-step S4 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application. Figure 6 As shown, step S4 includes: S41, performing feature concatenation on the visual feature vector, the lidar feature vector, and the millimeter-wave radar feature vector to obtain a multimodal concatenated feature vector; S42, extracting inconsistency features from the inconsistency metric matrix to obtain an inconsistency score vector; and S43, inputting the inconsistency score vector and the multimodal concatenated feature vector into a small multilayer perceptron to obtain the attention weight vector.

[0048] Specifically, in step S41, the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector are concatenated to obtain a multimodal concatenated feature vector. It should be understood that since the visual feature vector only carries the target's appearance texture and two-dimensional contour information, the LiDAR feature vector only focuses on three-dimensional spatial position and shape information, and the millimeter-wave radar feature vector only contains motion speed and signal strength information, each of these three types of feature vectors has information limitations when used individually, failing to fully represent the multi-dimensional attributes of dynamic objects and easily leading to the loss of key information during perception. Therefore, this application further concatenates the three types of feature vectors according to preset rules to integrate the three core information categories of appearance, space, and motion, forming a unified feature form containing multi-dimensional target representation. This eliminates the information limitations of single-modal features, providing comprehensive target information input for subsequent small multilayer perceptrons, ensuring that perception decisions can be based on complete multimodal information, and improving the comprehensiveness and accuracy of dynamic object attribute judgment.

[0049] Specifically, in one possible embodiment, step S41 is implemented as follows: First, the sequence and dimensionality adaptation rules for feature stitching are determined. The preset stitching order is visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector. If the dimensions of the three types of vectors differ, they are unified to the same dimension through a feature mapping layer. Then, the dimensions of the three types of feature vectors for each target are verified to ensure that all vectors meet the preset dimensionality requirements. If there are abnormal dimension vectors, they are adjusted to the standard dimension through interpolation or feature clipping. Subsequently, the three types of feature vectors are concatenated end-to-end in the preset order to form a single high-dimensional feature vector. This vector simultaneously contains visual appearance information, LiDAR spatial information, and millimeter-wave radar motion information. Finally, the validity of the stitched feature vector is verified to confirm that no information is lost or the dimension is incorrect, ultimately obtaining a multimodal stitched feature vector.

[0050] Specifically, step S42 involves extracting inconsistency features from the inconsistency metric matrix to obtain an inconsistency score vector. It should be understood that since the inconsistency metric matrix is ​​a two-dimensional structure containing inconsistency information for different modal pairs and targets, directly using it for subsequent calculations would increase computational complexity due to its high dimensionality, and the matrix contains redundant difference data, failing to efficiently reflect the overall inconsistency patterns of each modality. Therefore, this application further extracts inconsistency features from the inconsistency metric matrix to condense the difference information of the two-dimensional matrix into a low-dimensional score vector, retaining the core indicators representing modal inconsistency. This reduces the input dimensionality of the subsequent small multilayer perceptron, decreases computational resource consumption, and extracts key inconsistency features, ensuring that the multilayer perceptron can accurately capture modal difference patterns, providing an efficient and core basis for the calculation of attention weights.

[0051] Specifically, in one possible embodiment, step S42 is implemented as follows: First, the extraction indicators for inconsistency features are determined. The preset extraction dimensions include the average inconsistency value, maximum inconsistency value, and inconsistency variance for each modality pair (visual-LiDAR, visual-millimeter-wave radar, LiDAR-millimeter-wave radar), totaling three core indicators. Then, the inconsistency metric matrix is ​​traversed, and statistics are performed according to modality pairs. The average inconsistency value (arithmetic mean of all elements), maximum inconsistency value (maximum value in the column of that modality pair in the matrix), and inconsistency variance (reflecting the degree of fluctuation in the difference between the modality pairs) for each modality pair across all targets are calculated. Subsequently, all extracted indicators are combined in a preset order (average inconsistency, maximum inconsistency, variance) to form a fixed-length vector. Finally, the indicators in the vector are standardized to eliminate dimensional differences and ensure balanced weights for each indicator, resulting in an inconsistency score vector.

[0052] Specifically, in step S43, the inconsistency score vector and the multimodal concatenated feature vector are input into a small multilayer perceptron to obtain the attention weight vector. It should be understood that since the inconsistency score vector only reflects the degree of difference between modalities and lacks consideration of the content of each modal feature vector itself, while the multimodal concatenated feature vector only contains multidimensional information of the target and does not reflect modal reliability differences, calculating attention weights based solely on either vector cannot simultaneously consider modal differences and feature content, easily leading to weight allocation deviating from the actual needs of dynamic object perception. Therefore, this application further inputs the inconsistency score vector and the multimodal concatenated feature vector together into the small multilayer perceptron to learn the correlation between the two through the network, outputting an attention weight vector that matches both modal reliability and feature content. This allows the calculation of attention weights to simultaneously integrate difference information and target attribute information, making weight allocation more accurate and better suited to the perception needs of complex scenarios, ultimately improving the overall performance of multimodal fusion and the accuracy of dynamic object perception.

[0053] Specifically, in one possible embodiment, step S43 is implemented as follows: First, a small multilayer perceptron (MLP) network structure is constructed: the input layer dimension is the sum of the inconsistency score vector dimension and the multimodal concatenated feature vector dimension; the hidden layer is set to two layers, using ReLU as the activation function; the output layer dimension is consistent with the number of modalities, i.e., 3; the output value is normalized to obtain the attention weight vector. The weight parameters of this MLP are trained through end-to-end supervised learning. The training data uses multimodal labeled data adapted to elderly care scenarios, covering common home environments (such as living rooms and bedrooms), including different lighting, occlusion, and dynamic interference scenarios. Each data point simultaneously collects information from three types of sensors and labels the real state of dynamic objects. The training process is as follows: the entire fusion module is incorporated into the training framework, and a hybrid loss function is defined by combining kinematic state error and semantic state error. The inconsistency score vector and multimodal concatenated feature vector of the training data are input into the MLP to obtain the attention weights and complete the weighted fusion. Based on the error between the fusion result and the real label, an appropriate optimizer is used to update the MLP weights through backpropagation. Iterative training continues until the loss function converges, ensuring that the MLP can learn the mapping relationship between modal reliability and attention weights under different environments, thus adapting to the dynamic perception needs of elderly care scenarios.

[0054] In particular, in another possible preferred embodiment, step S4 includes: concatenating the feature vectors of each modality and their corresponding row vectors in the inconsistency metric matrix to generate modulation features corresponding to each modality, thereby using the inconsistency information between modalities to perform gating adjustment on the original feature vectors; combining the modulation features corresponding to each modality and calculating a modality attention matrix through a self-attention mechanism, wherein the modality attention matrix represents the degree of correlation attention between different modalities; fusing the modality attention matrix with the inconsistency metric matrix to generate a total modality-related weight, wherein the fusion is used to adjust the degree of correlation attention based on the degree of inconsistency between modalities; weighting the modulation features based on the total modality-related weight, and inputting the weighted features into a multilayer perceptron to generate an attention weight vector.

[0055] Here, for the visual modality, the inconsistency between it and the lidar modality, and its inconsistency with the millimeter-wave radar modality, reflects differences in different physical worlds. For example, the former might be contour differences, while the latter might be motion blur. Therefore, the inconsistency scores in the inconsistency metric matrix should be treated differently. Furthermore, inconsistency information can be used as prior knowledge to adjust or gate the feature distribution of the corresponding modality. For example, when visual features and lidar features are highly inconsistent, it is expected that this inconsistent signal will directly act on the visual feature vector, causing it to self-suppress or adjust its expression in subsequent weight competition. Thus, the weighted representation of each modality's features should include not only the isolated state of each modality but also the interactions between modalities; that is, the weight of a modality should depend on its own credibility and its correlation with other modalities.

[0056] First, the feature vectors of each mode and their corresponding row vectors in the inconsistency metric matrix are concatenated to generate the modulation features corresponding to each mode, thereby using the inconsistency information between modes to perform gating adjustment on the original feature vectors.

[0057] Specifically, for the features of each modality and the row vectors in the corresponding inconsistency metric matrix To perform cascading based on dissimilar spatial distribution gating, i.e.:

[0058] in, For cascading, For the first Modal inconsistency measure row vector, For the first The original feature vector specific to the modality. It is an L1 norm. It is the L2 norm. For dot product, For the first Modal feature-inconsistency concatenation vector.

[0059] Then it is modulated by a learnable weight matrix and activated with a sigmoid activation function:

[0060] in, It is the Sigmoid activation function. For gated modulation, learnable weight matrix, For matrix multiplication, For the first Modulation characteristics after gating activation of a mode.

[0061] In other words, it is possible to learn the features of a modality. Based on its modal inconsistency measure, its spatial distribution shows significant differences, therefore its composite distribution... It can inhibit The intensity of the modulated feature expression is used to modulate or gate the feature distribution of the corresponding modality, reflecting inconsistency information as prior knowledge. This allows the intensity of the modulated feature expression to be directly correlated with modality confidence, avoiding interference from low-confidence modality features in the fusion results. For example, in a dimly lit bedroom, the visual modality exhibits significantly increased inconsistency with other modalities due to insufficient light. Gating modulation can suppress noisy dimensions in the visual features, reducing their impact on the location perception results of older adults.

[0062] Next, the modulation features corresponding to each modality are combined, and a modality attention matrix is ​​calculated using a self-attention mechanism. The modality attention matrix represents the degree of correlation and attention between different modalities. That is, the modulation features of all modalities are combined. Introducing a self-attention mechanism as a set, firstly, three modulation features are... Two-dimensional arrangement to form the input matrix Its dimension is (n, L), where n=3 is the number of modes and L is the dimension of the modulation features. Then, the convolution matrix is ​​obtained from the input matrix E through two different two-dimensional convolutions. and and based on Obtain a 3x3 modal attention matrix, where, and The modal feature convolution matrix, For matrix transpose, This is the intermodal local correlation interest matrix. That is, its matrix elements... This represents the degree of attention given to the local association between modalities i and j based on two-dimensional convolutional correlation. This enables global reasoning about inter-modal associations, ensuring that the fusion process considers not only the credibility of individual modalities but also the cooperative or conflicting relationships between them. For example, lidar and millimeter-wave radar perceive the movement trajectories of the elderly with high consistency, but both deviate from the visual modality. The modality attention matrix clearly reflects the high correlation between lidar and millimeter-wave radar, and the low correlation between both and the visual modality, providing a global correlation basis for subsequent weight allocation.

[0063] Then, the modal attention matrix is ​​fused with the inconsistency metric matrix to generate total modal relevance weights. This fusion is used to adjust the degree of correlation attention based on the degree of inconsistency between modalities. ,in This is the total modal correlation weight matrix. This is a global intermodal inconsistency measurement matrix, which represents the total attention based on the degree of attention and inconsistency across all modalities (including itself). In other words, the higher the inconsistency, the greater the attention should be. This ensures that the total modal relevance weights simultaneously consider modal relationships and inconsistency differences, guaranteeing that the weight allocation aligns with intermodal synergy while reflecting the credibility differences caused by inconsistency. For example, when facing a highly reflective metal handrail, the inconsistency in the perception of the handrail edge between LiDAR and visual modalities increases significantly. Even if the two modal relevance levels are high, the fused total modal relevance weights will be adjusted according to the high inconsistency, avoiding the introduction of positional perception bias in the elderly due to over-attention to this associated combination.

[0064] Finally, the modulation features are weighted based on the total modality relevance weights, and the weighted features are input into a multilayer perceptron to generate an attention weight vector. In this way, by fully utilizing the inconsistency information in the inconsistency metric matrix, it can learn that inconsistencies with LiDAR (potentially contour problems) and inconsistencies with millimeter-wave radar (potentially motion estimation problems) are different situations, improving information fidelity. Furthermore, by establishing a strong nonlinear correlation with the spatial distribution gating between priors and features, inconsistency information can be dynamically filtered and reweighted to adjust the spatial dimensions of the original features. Moreover, the self-attention mechanism achieves context awareness. For example, even if a visual feature itself is weak, if it is highly inconsistent with all other modalities, and the other modalities are highly consistent with each other, the self-attention mechanism will capture this specific pattern, thereby increasing its weight and achieving global relational reasoning. For instance, in strong light environments, the credibility of visual modalities decreases significantly due to overexposure. After processing by the multilayer perceptron, the weighted visual modulation features output attention weights that are significantly lower than those of LiDAR and millimeter-wave radar, effectively avoiding the impact of visual noise on the perception of the elderly's motion state.

[0065] In the aforementioned dynamic object perception method based on multi-source sensor information fusion, step S5 involves weighted fusion of visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors based on an attention weight vector, resulting in a final state output to obtain the fused target state. It should be understood that although the adaptive attention weight vector assigns weights reflecting the current credibility of each modality, directly applying a uniform weighted average to all modal features still presents problems. On one hand, some modalities may output completely invalid or erroneous data at specific times, such as when a sensor is completely occluded. In this case, simply reducing its weight may not be sufficient to eliminate its negative impact, requiring a mechanism to completely exclude it from the fusion process. On the other hand, feature vectors from different modalities contain multiple attributes, such as kinematic states (position, velocity) and semantic states (target category, posture). These attributes have different properties, and each modality has different perception capabilities. Using a single weighted fusion strategy makes it difficult to achieve optimal fusion results for all attributes. Therefore, this application further uses an attention weight vector to first employ a fusion strategy combining weighted and selective approaches for each modality's features, intelligently combining features from effective modalities. This ensures that the fusion process can not only dynamically adjust the contribution of each modality, but also completely filter out invalid data sources and adopt the optimal fusion method for different information types, thereby generating a final target state that is more accurate and reliable in all dimensions.

[0066] In particular, in one specific embodiment, Figure 7 This is a flowchart of sub-step S5 of the dynamic object perception method based on multi-source sensor information fusion according to an embodiment of this application. Figure 7 As shown, step S5 includes: S51, validating the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector to obtain a list of valid modalities; S52, based on the attention weight vector and the list of valid modalities, performing multi-attribute weighted and selective fusion on the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector to obtain the fused kinematic state and the fused semantic state, wherein the fused kinematic state and the fused semantic state constitute the fused target state.

[0067] Specifically, step S51 involves validating the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector to obtain a list of valid modalities. It should be understood that in elderly care scenarios, visual feature vectors are easily affected by changes in lighting (such as darkness at night or direct sunlight), resulting in pixel loss or a sharp drop in confidence. LiDAR feature vectors may have incomplete point cloud clusters due to furniture obstruction or excessive target distance. Millimeter-wave radar feature vectors may experience abnormal fluctuations in motion parameters due to electromagnetic interference from indoor electrical appliances. Directly using these potentially invalid feature vectors for fusion would introduce erroneous data into subsequent processes, leading to distortion of the fused target state. Therefore, this application further performs validity checks on the three types of feature vectors—visual, LiDAR, and millimeter-wave radar—to filter out modalities with valid output data and form a list of valid modalities, thereby accurately eliminating the interference of invalid data on the fusion process. This ensures that subsequent fusion steps are based solely on reliable modal data, avoiding kinematic deviations (such as position and velocity) or semantic misjudgments (such as posture) caused by invalid information. This lays a reliable data foundation for the elderly care robot to accurately perceive dynamic objects such as the elderly and pets, ensuring the robot's safety and accuracy in interactive scenarios such as companionship and obstacle avoidance.

[0068] Specifically, in one possible embodiment, step S51 is implemented as follows: First, based on the technical characteristics of each modal sensor and the perception requirements of the elderly care scenario, a clear validity verification standard is preset. The verification standard for visual feature vectors is that the target detection category matches and the confidence level is not lower than a preset threshold, and the bounding box has no missing pixels, such as for people and pets. The verification standard for LiDAR feature vectors is that the number of effective points in the point cloud cluster meets the standard and the three-dimensional size is within the physical parameter range of the human body or common dynamic objects. The verification standard for millimeter-wave radar feature vectors is that the speed and distance parameters are within a reasonable range of motion and the signal strength is stable. Then, the confidence level and bounding box information of the visual feature vectors, the point cloud cluster parameters of the LiDAR feature vectors, and the motion and signal parameters of the millimeter-wave radar feature vectors are extracted one by one and compared with the corresponding verification standards to determine whether each modal feature vector is valid. Finally, the modal identifiers (such as visual and LiDAR) that pass the verification are summarized in a preset format to form a valid modal list. This list is directly passed to the subsequent multi-attribute fusion step to ensure that only the feature vectors of valid modalities participate in the fusion calculation and to ensure the reliability of the fusion process.

[0069] Specifically, in step S52, based on the attention weight vector and the effective modality list, multi-attribute weighted and selective fusion is performed on the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector to obtain the fused kinematic state and the fused semantic state. The fused kinematic state and the fused semantic state constitute the fused target state. It should be understood that since the visual, LiDAR, and millimeter-wave radar feature vectors contain two types of attributes—kinematic (e.g., 3D position, instantaneous velocity) and semantic (e.g., target posture)—kinematic attributes are continuous numerical types and each modality can output valid data, while semantic attributes are mostly discrete categorical types and only the visual modality can achieve accurate recognition, using a single fusion logic, such as weighted fusion for all, can lead to distortion of semantic attributes due to interference from non-professional modality data, or insufficient accuracy of kinematic attributes due to the lack of integration of multi-modal data advantages. Therefore, this application further performs multi-attribute weighted and selective fusion on the two types of attributes based on the attention weight vector and the effective modality list to adapt to the feature differences of different attributes and ensure the accuracy of kinematic attribute fusion and the effectiveness of semantic attribute fusion. In this way, the advantages of kinematic data from various effective modalities can be integrated through weighted fusion, reducing the impact of random errors in a single modality on the kinematic state. At the same time, selective fusion retains the high-confidence semantic information of the modality with the highest weight, avoiding misclassification caused by non-professional modal semantic data. Ultimately, this improves the adaptability of the fused target state to the decision-making needs of the elderly care robot, ensuring the safety and accuracy of the robot in interaction, navigation and other scenarios.

[0070] Specifically, in one possible embodiment, step S52 is implemented as follows: First, the kinematic attributes to be fused from the feature vectors of vision, lidar, and millimeter-wave radar are identified as three-dimensional position and instantaneous velocity, and the semantic attribute is the target posture. Next, based on the list of valid modalities, the modalities with currently valid output data are selected, while invalid data modalities are excluded. Then, for the kinematic attributes, the three-dimensional position parameters corresponding to each valid modality are weighted and calculated with the corresponding modal weights in the attention weight vector to obtain the fused three-dimensional position. Similarly, the instantaneous velocities of each valid modality are weighted and calculated to obtain the fused instantaneous velocity. The fused three-dimensional position and instantaneous velocity together constitute the fused kinematic state. Finally, for the semantic attributes, the valid modality with the highest weight in the attention weight vector is extracted, and posture information, such as standing or sitting posture, is read from the feature vector of this modality as the fused semantic state. Finally, the fused kinematic state and semantic state are integrated to form the fused target state, which is transmitted to the upper-level decision-making module of the elderly care robot for path planning and human-machine interaction control.

[0071] In summary, the dynamic object perception method based on multi-source sensor information fusion, as described in the embodiments of this application, is elucidated. It processes heterogeneous data from multiple sources, such as RGB-D images, 3D point clouds, and millimeter-wave radar, in parallel and introduces a cross-modal inconsistency measurement mechanism. By actively analyzing and quantifying the differences and conflicts between features of different sensors, it evaluates the perception confidence of each sensing modality in the current environment in real time. Based on this, an adaptive attention mechanism is driven to dynamically generate fusion weights for each modality, ultimately achieving intelligent weighted fusion of feature vectors from each modality and outputting a highly robust fused target state. This intelligently suppresses low-quality or conflicting sensor information and enhances the contribution of high-quality information, thereby achieving more robust and accurate dynamic object perception in complex and changing environments.

[0072] Figure 8 This is a block diagram of a dynamic object perception system based on multi-source sensor information fusion according to an embodiment of this application. Figure 8 As shown, the dynamic object perception system 100 based on multi-source sensor information fusion according to an embodiment of this application includes: a multimodal data acquisition module 110, used to acquire RGB-D images, 3D point clouds, and millimeter-wave radar data; a modality-specific feature extraction module 120, used to extract modality-specific features from the RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors; a cross-modal inconsistency measurement module 130, used to perform cross-modal inconsistency measurement on the visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors to obtain an inconsistency measurement matrix; an adaptive attention weight measurement module 140, used to perform adaptive attention weight measurement on the visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors based on the inconsistency measurement matrix to obtain an attention weight vector; and a fusion and state output module 150, used to perform weighted fusion and final state output on the visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors based on the attention weight vector to obtain the fused target state.

[0073] As described above, the dynamic object perception system 100 based on multi-source sensor information fusion according to the embodiments of this application can be implemented in various wireless terminals, such as servers with dynamic object perception algorithms based on multi-source sensor information fusion. In one possible implementation, the dynamic object perception system 100 based on multi-source sensor information fusion according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the dynamic object perception system 100 based on multi-source sensor information fusion can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the dynamic object perception system 100 based on multi-source sensor information fusion can also be one of many hardware modules of the wireless terminal.

[0074] Alternatively, in another example, the dynamic object perception system 100 based on multi-source sensor information fusion and the wireless terminal can also be separate devices, and the dynamic object perception system 100 based on multi-source sensor information fusion can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.

[0075] Here, those skilled in the art will understand that the specific operations of each step in the above-described dynamic object perception system based on multi-source sensor information fusion have been referenced above. Figures 1 to 7 The dynamic object perception method based on multi-source sensor information fusion has been described in detail in the previous section, and therefore, its repeated description will be omitted.

Claims

1. A dynamic object perception method based on multi-source sensor information fusion, characterized in that, include: Acquire RGB-D images, 3D point clouds, and millimeter-wave radar data; Modality-specific feature extraction is performed on RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors. Cross-modal inconsistency measurement is performed on visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors to obtain an inconsistency measurement matrix. This includes: performing validity checks and state extraction on visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors respectively to obtain a set of modal target structured coding vectors. Calculate the inconsistency between any two modal target structured coding vectors in the modal target structured coding vector set to obtain the inconsistency metric matrix; Based on the inconsistency metric matrix, an adaptive attention weight metric is performed on the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector to obtain the attention weight vector. Based on the attention weight vector, the visual feature vector, the LiDAR feature vector, and the millimeter-wave radar feature vector are weighted and fused, and the final state output is obtained to obtain the fused target state.

2. The dynamic object perception method based on multi-source sensor information fusion according to claim 1, characterized in that, Modality-specific feature extraction is performed on RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors, including: Multimodal parallel initial detection is performed on RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain a visual candidate object set, a lidar candidate object cluster set, and a millimeter-wave radar candidate point set; Based on the trajectory database of the previous moment, intramodal data association and state estimation are performed on the visual candidate object set, the lidar candidate object cluster set, and the millimeter-wave radar candidate point set to obtain the trajectory database updated at the current moment. Modality-specific feature encapsulation is performed on the trajectory database updated at the current moment to obtain visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors.

3. The dynamic object perception method based on multi-source sensor information fusion according to claim 2, characterized in that, Based on the trajectory database from the previous time step, intramodal data association and state estimation are performed on the visual candidate object set, the lidar candidate object cluster set, and the millimeter-wave radar candidate point set to obtain the trajectory database updated at the current time step, including: Extract the set of lidar trajectories from the trajectory database of the previous moment; Each lidar trajectory in the lidar trajectory set is input into a constant-speed Kalman filter for state prediction to obtain the predicted state and prediction covariance of all lidar trajectories at the current moment. Calculate the predicted state and prediction covariance of all lidar trajectories at the current time and the gated cost matrix between the lidar candidate clusters; Based on the post-gated cost matrix, successfully matched trajectory-cluster pairs are determined. The trajectory is updated based on the successfully matched trajectory-cluster pairs to obtain the updated trajectory database at the current moment.

4. The dynamic object perception method based on multi-source sensor information fusion according to claim 1, characterized in that, Each modal target structured encoding vector in the modal target structured encoding vector set includes a validity indicator for each modality, a three-dimensional position estimate, and a scalarized position estimate standard deviation.

5. The dynamic object perception method based on multi-source sensor information fusion according to claim 4, characterized in that, Calculating the inconsistency between any two modal target structured coding vectors in the modal target structured coding vector set to obtain an inconsistency metric matrix includes: calculating the inconsistency between any two modal target structured coding vectors in the modal target structured coding vector set using the following formula: ; in, For the first 3D position estimation in modal target structured encoding vector For the first 3D position estimation in modal target structured encoding vector For the first Standard deviation of scalar position estimates in modal target structured encoding vectors For the first Standard deviation of scalar position estimates in modal target structured encoding vectors For smoothing terms, Indicates the first and the Inconsistency between modal target structured encoding vectors.

6. The dynamic object perception method based on multi-source sensor information fusion according to claim 1, characterized in that, Based on the inconsistency metric matrix, an adaptive attention weight metric is applied to the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector to obtain the attention weight vector, including: Visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors are concatenated to obtain multimodal concatenated feature vectors. Inconsistency features are extracted from the inconsistency measure matrix to obtain the inconsistency score vector; The inconsistency score vector and the multimodal concatenated feature vector are input into a small multilayer perceptron to obtain the attention weight vector.

7. The dynamic object perception method based on multi-source sensor information fusion according to claim 1, characterized in that, Based on the attention weight vector, the visual feature vector, LiDAR feature vector, and millimeter-wave radar feature vector are weighted and fused, and the final state output is obtained to obtain the fused target state, including: The validity of visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors is validated to obtain a list of valid modes. Based on the attention weight vector and the effective modality list, the visual feature vector, the LiDAR feature vector, and the millimeter-wave radar feature vector are subjected to multi-attribute weighting and selective fusion to obtain the fused kinematic state and the fused semantic state. The fused kinematic state and the fused semantic state constitute the fused target state.

8. A dynamic object perception system based on multi-source sensor information fusion, characterized in that, include: A multimodal data acquisition module is used to acquire RGB-D images, 3D point clouds, and millimeter-wave radar data; The modality-specific feature extraction module is used to extract modality-specific features from RGB-D images, 3D point clouds, and millimeter-wave radar data to obtain visual feature vectors, lidar feature vectors, and millimeter-wave radar feature vectors. The cross-modal inconsistency measurement module is used to measure the cross-modal inconsistency of visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors to obtain an inconsistency measurement matrix. This includes: performing validity checks and state extraction on the visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors respectively to obtain a set of modal target structured coding vectors; and calculating the inconsistency between any two modal target structured coding vectors in the set of modal target structured coding vectors to obtain the inconsistency measurement matrix. The adaptive attention weight measurement module is used to perform adaptive attention weight measurement on visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors based on the inconsistency measurement matrix to obtain attention weight vectors. The fusion and state output module is used to perform weighted fusion of visual feature vectors, LiDAR feature vectors, and millimeter-wave radar feature vectors based on attention weight vectors and output the final state to obtain the fused target state.

Citation Information

Patent Citations

  • Mountain environment sensing system and method based on multi-sensor fusion

    CN121600412A