Autonomous driving multi-modal perception method based on privacy risk and sensor reliability evaluation
By introducing privacy risk indicators and sensor reliability assessment mechanisms, adaptive processing and fusion decision-making of multimodal data are performed, which solves the problems of insufficient privacy risk and reliability in multi-sensor fusion systems and achieves high-precision perception and robustness in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN UNIV OF TECH
- Filing Date
- 2026-02-11
- Publication Date
- 2026-06-05
AI Technical Summary
Existing multi-sensor fusion systems are insufficient in their dynamic response to privacy risks and adaptability to changes in sensor reliability, making it difficult to balance privacy protection and perception accuracy, and they also lack robustness in complex environments.
By introducing privacy risk indicators for adaptive fuzzy processing and combining sensor uncertainty and stability evaluation mechanisms, multimodal data is dynamically weighted and fused to achieve adaptive data processing and fusion decision-making.
While ensuring privacy protection, it improves the robustness and accuracy of the perception system, can maintain high-confidence perception results in complex environments, and has good real-time performance and engineering application potential.
Smart Images

Figure CN122153775A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent perception and autonomous driving technology, and in particular to a multimodal perception method based on dynamic assessment of privacy risks and fusion decision-making of sensor reliability. Background Technology
[0002] With the rapid development of autonomous driving technology, multi-sensor fusion systems have become a core component of vehicle environmental perception. Cameras, LiDAR, and millimeter-wave radar, among other sensors, leverage their respective advantages to jointly construct a multi-layered, multi-angle perception capability of the surrounding environment. Information complementarity enhances the accuracy of target detection and tracking. However, existing multi-sensor fusion methods still face two prominent challenges:
[0003] On the one hand, existing systems lack dynamic response mechanisms to privacy risks. Visual sensors such as cameras can easily collect sensitive information such as pedestrians and license plates. Current methods mostly use fixed-area or uniform-intensity blurring, which cannot adaptively adjust the protection intensity according to the real-time distribution (such as quantity and location density) of privacy targets in the scene, making it difficult to balance privacy protection and perception accuracy.
[0004] On the other hand, the multi-sensor fusion process is not adaptable enough to changes in its own reliability. Sensor performance is significantly affected by environmental interference (such as lighting, weather, and occlusion), while traditional fusion methods often use fixed weights or simple filtering algorithms, which fail to evaluate and respond to fluctuations in the confidence of each sensor in real time. When the performance of some sensors degrades, it can easily lead to a decrease in the reliability of the fusion results.
[0005] Therefore, there is an urgent need for a perception method that can simultaneously achieve quantitative assessment of privacy risks and dynamic measurement of sensor reliability, in order to support adaptive data processing and fusion decision-making, and improve the perception robustness and privacy security of the system in complex scenarios. Summary of the Invention
[0006] This invention proposes a multimodal perception method for autonomous driving based on privacy risk and sensor reliability assessment, aiming to simultaneously protect the privacy of traffic participants and ensure the robustness of the perception system. The overall approach is as follows: privacy risk indicators are introduced during the perception process, and camera data is adaptively blurred to reduce the risk of privacy leakage at the source; simultaneously, a sensor uncertainty and stability evaluation mechanism is combined to dynamically weight and fuse multimodal data, thereby obtaining more reliable environmental perception results. Specifically, the method includes the following steps: Step 1: Within the preset time window Within the system, raw information from M sensors, including cameras, LiDAR, and millimeter-wave radar, is acquired, and a target set for each frame is obtained using recognition and classification algorithms. Based on the similarity of target locations, similar targets detected at different times are grouped into the same target trajectory to form a target set. ; Step 2: Calculate the percentage of privacy-related targets and the percentage of regions within the time window based on the target set, to obtain the privacy risk vector for that time window. And based on this privacy risk vector, set the degree of blurring of the camera image for the next time window; Step 3: For each target Construct a sensor temporal belief distribution matrix, based on the privacy risk vector of the previous time window. Discount correction is applied to the camera belief distribution matrix; Step 4: Perform uncertainty and stability analysis on the belief distribution matrix formed by each sensor within the current time window to construct a score matrix and calculate the reliability factor. ; Step 5: Calculate the target importance weight based on the target region area and classification confidence score. Then, the belief distributions of each sensor are integrated and weighted and discounted according to the reliability factor and importance weight to obtain the updated belief distributions. Step 6: Perform evidence fusion on the multi-target belief distribution matrix of each sensor to obtain the final environmental perception and target classification results.
[0007] Furthermore, the privacy risk vector The expression is:
[0008] in This is an indicator of the percentage of privacy areas. The percentage of privacy-related objectives is expressed as follows:
[0009]
[0010] in For window size, Indicates the first Total area of the frame privacy region Represents the total area of the image. Indicates the first Number of frame privacy instances Indicates the first Total number of frame instances .
[0011] Furthermore, the degree of blur in the camera image is represented as follows, defining a blur mapping. , is represented as:
[0012] in These are manually set hyperparameters used to balance the impact of the privacy region area and the number of privacy instances on the image blurring algorithm; Furthermore, for each target The constructed sensor time belief distribution matrix is as follows:
[0013] in Represents the set of target categories. Indicates that sensor i is at time i Identified target Category The probability of support.
[0014] Furthermore, when applying discount corrections to the camera belief distribution matrix: Through monotonic mapping Get the discount factor At any time in the current window With the goal The camera belief distribution is subjected to Dempster-Shafer discounting, which returns a portion of the quality stream to the entire set. The distribution of beliefs after discounting :
[0015] Furthermore, uncertainty analysis is performed on the belief distribution matrix to construct the score matrix, using the following method: The average entropy of the belief distribution within the window is used as a metric, expressed as:
[0016] in Indicates sensor i For the target The uncertainty measure.
[0017] Furthermore, when constructing the score matrix for stability analysis of the belief distribution matrix: the time variance measure of the support probability of the primary discriminant class within the window is used, assuming... The main category is determined. The stability metric is expressed as follows:
[0018] in Indicates sensor i For the target Stability measure Indicates sensori For the target belong The average support level of the class within the current time window.
[0019] Furthermore, the scoring matrix and reliability factor are obtained by constructing the scoring matrix based on uncertainty and stability metrics. The two scores were then normalized.
[0020]
[0021] Set criterion vector Define sensor reliability score And normalized to a reliability factor. , is represented as:
[0022] Furthermore, the method for updating the importance weight and belief distribution matrix involves calculating the importance weight of each target using the region area and confidence level. After normalization and truncation, it is represented as:
[0023]
[0024] in This represents the average area of the target region within the window. Represents the total area of the image. This represents the minimum percentage of the target area; it's a hyperparameter used to prevent the weight from being zero if the camera doesn't detect the target. Weights are used in this context. Using the coefficients, a Dempster-Shafer discount is applied to each sensor to obtain the discounted belief distribution. , is represented as:
[0025] Furthermore, the evidence fusion method for the multi-target belief distribution matrix of each sensor is as follows: As M pieces of evidence, DS fusion is performed to obtain the final belief distribution. :
[0026] Finally passed To reach the final decision.
[0027] This invention discloses a multimodal perception method based on dynamic privacy risk assessment and sensor reliability fusion decision-making. This method is specifically applicable to scenarios such as autonomous vehicles. By fusing data from multiple sources such as cameras, LiDAR, and millimeter-wave radar, it achieves high-precision target detection while introducing privacy risk quantification and sensor reliability assessment mechanisms to adaptively adjust data preprocessing and fusion strategies, thereby balancing perception performance and privacy protection requirements.
[0028] This method dynamically adjusts the camera's blur level for the next moment based on the real-time distribution of sensitive targets in the scene using a privacy risk vector. This effectively avoids excessive loss of perception performance caused by fixed blur levels in traditional methods, achieving an adaptive balance between privacy protection and perception accuracy. Simultaneously, by dynamically adjusting the reliability factors of each sensor through real-time analysis of the uncertainty and stability of each sensor, it ensures that even when some sensors experience performance degradation or inaccuracy due to environmental interference, the perception decision still maintains high confidence, significantly improving the robustness and reliability of the system's perception in complex environments. Furthermore, this invention employs a sliding time window-based processing framework, which effectively smooths sensor noise in continuous time series, avoiding decision fluctuations caused by single-frame misidentification. The method is logically rigorous, computationally efficient, and can be directly integrated into existing autonomous driving perception frameworks, possessing excellent real-time performance and engineering application potential. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 The logic flowchart for the invention.
[0031] Figure 2 Different methods support different target categories in specific scenarios. A comparison chart of probability values. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] like Figure 1 As shown, this invention provides a multimodal perception method based on dynamic privacy risk assessment and sensor reliability fusion decision-making. Step 1: Collect multimodal raw data from cameras, LiDAR, and millimeter-wave radar within a preset time window, and extract the target set for each frame using recognition and classification algorithms. Based on the spatial similarity of the targets, similar targets detected at different times are grouped into a unified trajectory object, thereby constructing the target set within the time window; Step 2: Based on the proportion of privacy-related targets and their respective regions in the image, a privacy risk vector is constructed within the time window. Based on this vector, the system generates the image blur coefficients for the next time window using a linear weighting method, achieving adaptive adjustment of the privacy protection strength. Step 3: For each detected target, construct its temporal belief distribution matrix under different sensors. This matrix records the probability support distribution of each sensor for the target category at different times. Based on the privacy risk assessment results of the previous time window, apply probability discounting to the camera's belief distribution to eliminate classification bias caused by privacy protection processing. Step 4: Calculate the information entropy of the belief distribution and the variance of the support probability of the main discriminant class within the time window, and evaluate the uncertainty and stability of each sensor's recognition result for each target. Based on these two indicators, construct a comprehensive evaluation matrix for the sensors, and after normalization and combining it with preset weight coefficients, calculate the reliability factor of each sensor. Step 5: Taking into account the importance weight of the target (determined by the target area and classification confidence) and the reliability factor of the sensor, the belief distribution of each sensor is weighted and discounted to obtain the updated belief distribution; Step 6: Using the updated belief distributions of each sensor as evidence, a comprehensive decision is made using the evidence theory fusion method to obtain the final probability distribution of each target, and the category with the highest confidence is selected as the final classification result.
[0035] Example 1 This invention provides a multimodal perception method for autonomous driving based on privacy risk and sensor reliability assessment: This invention is applied to autonomous vehicles equipped with cameras, LiDAR, and millimeter-wave radar. The system first sets a time window size W = 3 frames. Within a given time window, for each frame, it identifies pedestrians, vehicles, and other traffic participants using object detection algorithms (such as YOLO or CenterNet) to obtain a target set. Then, based on location similarity and trajectory association algorithms (such as Hungarian matching), the detection results of the same object within the window are merged into trajectories, forming the target set. .
[0036] Based on the target recognition results within the current time window, the percentage of privacy regions in each frame is 0.03, 0.05, and 0.04, respectively, and the percentage of privacy instances is 0.33, 0.43, and 0.33, respectively. The calculations show... =0.04, The blur intensity of the next window image is... coefficient here Setting it to 0.6 allows for Gaussian blurring of images, with its variance set to [value missing]. In the form of.
[0037] For one of the detection targets Considering only single elements and the entire set, the temporal belief distribution matrices under camera, lidar, and millimeter-wave radar are shown in Tables 1, 2, and 3: Table 1. Belief distribution of Sensor 1 (camera)
[0038] Table 2. Belief Distribution of Sensor 2 (LiDAR)
[0039] Table 3. Belief Distribution of Sensor 3 (Millimeter-wave Radar)
[0040] Based on the privacy risk assessment results of the previous time window, the belief distribution of cameras is corrected by probability discounting to eliminate classification bias caused by privacy protection processing.
[0041] Based on the risk vector obtained from the previous time window The discount factor is obtained through monotonic mapping. ,in The values were set to 0.1 and 0.7. The camera belief distribution was then discounted and corrected, and the corrected belief distribution is shown in Table 4. Table 4. Belief distribution of Sensor 1 (camera) after privacy risk correction.
[0042] Then, the time-averaged mass distribution is calculated based on the detection results of each sensor. The quality allocation for the main category (pedestrians) is as follows: , , Then, the results are derived from the uncertainty calculation formula and the stability calculation formula. , After normalization, we obtain , Take the criterion weights The final reliability factor is obtained through weighting and normalization. .
[0043] This window contains a set of seven targets, whose attributes are shown in Table 5: Table 5 Target Attributes
[0044] According to the weighting formula, the minimum target proportion Set as ,get Then calculate the overall reliability weight. Then, the time-averaged mass distribution of each sensor was discounted and corrected according to its weight. The corrected belief distribution is shown in Table 5. Table 6. Corrected time-averaged mass distribution for each sensor
[0045] Finally, M-1 DS fusion operations are performed on the multi-target belief distribution matrices of each sensor. First, the multi-target belief distribution matrices of each sensor are calculated. With conflict and according to Normalization yields ; then use With conflict The final fusion quality allocation is obtained. As shown in Table 6: Table 7. Results of Fusion Quality Allocation
[0046] Based on the fusion quality allocation results, the target identification result is a pedestrian.
[0047] To further verify the effectiveness of the solution proposed in the embodiments of the present invention, the target recognition performance of the method of this patent, the method without privacy correction, and the method without time window correction were compared in scenarios with camera misalignment. The results are as follows: Figure 2 As shown. Figure 2 The paper demonstrates the classification accuracy of three methods for seven targets. It shows that the method of this patent guarantees classification accuracy while protecting privacy, and is also robust, with no decrease in recognition accuracy due to misidentification in a single frame.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal perception method for autonomous driving based on privacy risk and sensor reliability assessment, characterized in that, Includes the following steps: Step 1: Within the preset time window Within the system, raw information from M sensors, including cameras, LiDAR, and millimeter-wave radar, is acquired, and a target set for each frame is obtained using recognition and classification algorithms. Based on the similarity of target locations, similar targets detected at different times are grouped into the same target trajectory to form a target set. ; Step 2: Calculate the percentage of privacy-related targets and the percentage of regions within the time window based on the target set, to obtain the privacy risk vector for that time window. And based on this privacy risk vector, set the degree of blurring of the camera image for the next time window; Step 3: For each target Construct a sensor temporal belief distribution matrix, based on the privacy risk vector of the previous time window. Discount correction is applied to the camera belief distribution matrix; Step 4: Perform uncertainty and stability analysis on the belief distribution matrix formed by each sensor within the current time window to construct a score matrix and calculate the reliability factor. ; Step 5: Calculate the target importance weight based on the target region area and classification confidence score. Then, the belief distributions of each sensor are integrated and weighted and discounted according to the reliability factor and importance weight to obtain the updated belief distributions. Step 6: Perform evidence fusion on the multi-target belief distribution matrix of each sensor to obtain the final environmental perception and target classification results.
2. The autonomous driving multimodal perception method based on privacy risk and sensor reliability assessment according to claim 1, characterized in that, The privacy risk vector The expression is: in This is an indicator of the percentage of privacy areas. The expression for the percentage of privacy-related objectives is: in For window size, Indicates the first Total area of the frame privacy region Represents the total area of the image. Indicates the first Number of frame privacy instances Indicates the first Total number of frame instances .
3. The autonomous driving multimodal perception method based on privacy risk and sensor reliability assessment according to claim 1, characterized in that, The degree of blur in the camera image is represented as follows, defining a blur mapping. , is represented as: in These are manually set hyperparameters used to balance the impact of the privacy region area and the number of privacy instances on the image blurring algorithm.
4. The autonomous driving multimodal perception method based on privacy risk and sensor reliability assessment according to claim 1, characterized in that, For each target The constructed sensor time belief distribution matrix is as follows: in Represents the set of target categories. Indicates that sensor i is at time i Identified target Category The probability of support.
5. The autonomous driving multimodal perception method based on privacy risk and sensor reliability assessment according to claim 1, characterized in that, When discounting and correcting the camera belief distribution matrix: Through monotonic mapping Get the discount factor At any time in the current window With the goal The camera belief distribution is subjected to Dempster-Shafer discounting, which returns a portion of the quality stream to the entire set. The distribution of beliefs after discounting : 。 6. The autonomous driving multimodal perception method based on privacy risk and sensor reliability assessment according to claim 1, characterized in that, Uncertainty analysis is performed on the belief distribution matrix to construct the score matrix, using the following method: The average entropy of the belief distribution within the window is used as a metric, expressed as: in Indicates sensor i For the target The uncertainty measure.
7. The autonomous driving multimodal perception method based on privacy risk and sensor reliability assessment according to claim 1, characterized in that, When constructing the score matrix for stability analysis of the belief distribution matrix: the time variance measure of the support probability of the principal discriminant class within the window is used, let... The main categorization is... The stability metric is expressed as follows: in Indicates sensor i For the target Stability measure Indicates sensor i For the target belong The average support level of the class within the current time window.
8. The autonomous driving multimodal perception method based on privacy risk and sensor reliability assessment according to claims 6 to 7, characterized in that, The scoring matrix and reliability factor are obtained by constructing the scoring matrix based on uncertainty and stability metrics. The two scores were then normalized. Set criterion vector Define sensor reliability score And normalized to a reliability factor. , is represented as: 。 9. A multimodal perception method for autonomous driving based on privacy risk and sensor reliability assessment as described in claim 1, characterized in that, The method for updating the importance weight and belief distribution matrix is to calculate the importance weight of each target using the region area and confidence level. After normalization and truncation, it is represented as: in This represents the average area of the target region within the window. Represents the total area of the image. This represents the minimum percentage of the target area; it's a hyperparameter used to prevent the weight from being zero if the camera doesn't detect the target. Weights are used in this context. Using the coefficients, a Dempster-Shafer discount is applied to each sensor to obtain the discounted belief distribution. , is represented as: 。 10. The autonomous driving multimodal perception method based on privacy risk and sensor reliability assessment according to claim 1, characterized in that, The evidence fusion method for the multi-target belief distribution matrix of each sensor is as follows: As M pieces of evidence, DS fusion is performed to obtain the final belief distribution. : Finally passed To reach the final decision.