A screening processing system and method for intelligent driving long tail scene multi-modal data

CN122839152APending Publication Date: 2026-09-29SUZHOU KUSHUJU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611009883.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0002]智能驾驶技术的快速迭代,对多模态数据的质量、稳定性与可靠性提出极高要求,而城市路口、雨雪雾极端天气、隧道/桥梁特殊路况、乡村窄路、高速施工区等长尾场景数据,因采集环境复杂、传感器工况波动大、干扰因素多,普遍存在质量参差、模态不一致、时序紊乱、异常频发等问题,成为制约智能驾驶模型泛化能力与算法鲁棒性的关键瓶颈

Benefits of technology

[0010]本申请的有益效果在于:构建融合多维度信息的数据码本,通过相似度检索生成候选集;再以六项核心指标完成数据四级质量划分;结合时序对齐、帧率校准、运动检测与场景识别,感知边界并确定筛选区间;依托分级路由与缓存基准执行差异化筛选;对异常数据修复、低质量样本优化;最后经三级校验兜底,输出优质数据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839152A_ABST
    Figure CN122839152A_ABST
Patent Text Reader

Abstract

This invention provides a system and method for filtering and processing multimodal data in long-tail scenarios of intelligent driving, applicable to the field of data processing. Addressing the challenges of inconsistent data quality and difficult filtering in long-tail scenarios of intelligent driving, a comprehensive filtering and processing system is constructed. First, multimodal features are extracted and a data codebook integrating scene, target, spatiotemporal, operating condition, and interference information is constructed. A candidate set is generated through similarity retrieval. Then, data quality is quantified based on six core indicators, completing a four-level classification. By combining temporal alignment and frame rate calibration to perceive scene boundaries, the filtering interval and stability constraints are determined. A hierarchical routing decision is established based on quality levels and boundary information, using a cache library as a benchmark. Abnormal data is filtered and corrected, and features of low-quality samples are optimized. Through full-dimensional verification and multi-level fallback processing, continuous, stable, and high-fidelity high-quality multimodal data is output, effectively supporting the iterative optimization of intelligent driving models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a system and method for filtering and processing multimodal data in long-tail scenarios of intelligent driving. Background Technology

[0002] The rapid iteration of intelligent driving technology places extremely high demands on the quality, stability, and reliability of multimodal data. However, long-tail scenario data, such as urban intersections, extreme weather conditions like rain, snow, and fog, special road conditions like tunnels / bridges, narrow rural roads, and highway construction zones, generally suffer from problems such as inconsistent quality, modal inconsistencies, temporal disorder, and frequent anomalies due to complex collection environments, large fluctuations in sensor operating conditions, and numerous interference factors. These issues have become key bottlenecks restricting the generalization ability of intelligent driving models and the robustness of algorithms.

[0003] Existing technologies for processing multimodal data in long-tail scenarios have significant shortcomings: First, they lack a standardized, end-to-end data screening system, relying heavily on single quality indicators or manual experience, failing to consider temporal continuity, modal matching, and scene fidelity, and easily overlooking high-value long-tail data or misselecting high-quality samples; Second, for typical long-tail data problems such as temporal misalignment, inconsistent frame rates, sudden motion changes, and blurred scene boundaries, the processing logic is fragmented, lacking a unified perception and calibration scheme, making it difficult to accurately identify operational condition changes and scene boundaries; Third, anomaly identification is limited to a single dimension, only clustering... The surface anomalies, such as noise or missing data, are not combined with long-tail characteristics such as rarity and spatiotemporal deviation, making it impossible to distinguish between minor, moderate, and severe anomalies, resulting in insufficient differentiated processing capabilities. Fourth, the quality verification and fallback mechanisms are imperfect, mostly remaining at the single-frame or segment level, lacking batch-level comprehensive judgment and backtracking mechanisms, which can easily lead to the spread of local defects and the failure of the entire batch of data. Fifth, the data repair and optimization are not targeted enough, and no linkage rules between anomaly types and processing strategies have been established, which can easily lead to over-repair or under-repair, making it difficult to ensure the consistency between the repaired data and the original time series and modality.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0006] According to one aspect of this application, a method for filtering and processing multimodal data in long-tail scenarios of intelligent driving is provided, comprising: acquiring the original features of multimodal data; constructing a codebook for long-tail scenario multimodal data that integrates scene semantic features, target contour parameters, spatiotemporal location labels, operating condition information, and environmental interference parameters; performing feature similarity matching retrieval to generate a candidate set of time-series data for filtering; based on six core indicators—scene rarity, modal noise, data missing anomaly, spatiotemporal misalignment deviation, target occlusion interference, and operating condition matching error—completing sample-level data risk and quality normalization calculation, and achieving a four-level classification of invalid, inefficient, effective, and high-quality long-tail scenario data; and combining time-series frame alignment, sensor frame rate calibration, motion posture change detection, and scene switching recognition to achieve perception of long-tail scenario boundaries and operating condition switching, and determining the effective filtering interval and temporal continuity of multimodal data. The system establishes stable constraints for non-boundary data, including modality matching and scene fidelity. Based on data quality level and scene boundary information, it establishes hierarchical routing and screening decision rules, including direct retention, mild repair, modality completion, sample removal, and batch filtering. A data screening benchmark library is built based on a historical high-quality data caching module. Noise filtering, missing data interpolation, and spatiotemporal offset correction are performed on modal anomaly data. Feature weight fine-tuning, redundant information removal, interference parameter suppression, and multimodal consistency forced alignment are performed on low-quality long-tail samples. The system verifies the multi-dimensional output of multimodal data in terms of temporal continuity, scene authenticity, parameter matching degree, and effective information ratio. If the verification fails, multi-level conservative fallback processing, including single-frame correction, fragment replacement, and batch backtracking, is performed. Through risk quantification, hierarchical screening, modality repair, and temporal optimization, high-quality multimodal data for long-tail scenarios in intelligent driving is generated.

[0007] Another aspect of this application discloses a system for filtering and processing multimodal data in long-tail scenarios of intelligent driving, comprising: a distributed ultrasound image preprocessing module for performing size unification, type standardization, and pixel normalization processing on multi-center distributed cardiac ultrasound images, completing single-frame image segmentation, Gaussian filtering noise reduction, 3×3 Laplacian convolution texture scoring, and adaptive mask generation, outputting a normalized temporal frame sequence, mask matrix, and position index information; a privacy-compliant federated pre-training module for building a federated learning architecture, reconstructing mask image blocks based on MAE encoding and decoding, updating and pruning local gradients according to MSE loss; a central server removing offline clients, iteratively updating the global model by weighting and aggregating local parameters according to data volume, and realizing privacy-compliant feature extraction where data is usable but not visible; a sparse temporal feature encoding module for performing sparse sampling on cardiac ultrasound temporal frames, extracting single-frame spatial features using pre-trained ViT, and generating sine-cosine sparse temporal codes according to the acquisition index; and feature fusion. The data is then enhanced with Transformer multi-head self-attention for temporal correlation, outputting a frame-level temporal feature sequence. A dual-branch spatiotemporal feature fusion module constructs a dual-branch spatiotemporal fusion network, extracting global spatiotemporal features through 3D residual convolution combined with gated attention, and generating refined temporal features through 2D residual convolution fused with temporal encoding. The two types of features are then concatenated to obtain high-dimensional spatiotemporal fusion feature information. A temporal aggregation and classification prediction module performs attention-weighted temporal aggregation on the spatiotemporal fusion features, compressing them into video-level features. The input fully connected classification head is activated by softmax, outputting multi-class probability prediction results including cardiac amyloidosis. A robust enhancement and auxiliary screening module adapts to linear and convex matrix scanning modes through bidirectional polar coordinate transformation, simulating motion blur, Gaussian blur, and salt-and-pepper noise for joint data enhancement. By fusing adaptive masking, sparse temporal encoding, and dual-branch spatiotemporal fusion, the spatiotemporal features of cardiac ultrasound are accurately captured, generating intelligent auxiliary screening results for cardiac amyloidosis.

[0008] According to another aspect of this application, an electronic device is provided, on which a computer program is stored, which, when executed by a first processor, implements the above-described method for filtering and processing multimodal data in long-tail scenarios of intelligent driving.

[0009] According to another aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a second processor, implements the above-described method for filtering and processing multimodal data in long-tail scenarios of intelligent driving.

[0010] The beneficial effects of this application are as follows: constructing a data codebook that integrates multi-dimensional information, generating a candidate set through similarity retrieval; then completing the four-level quality division of data using six core indicators; combining temporal alignment, frame rate calibration, motion detection, and scene recognition to perceive boundaries and determine the screening interval; performing differentiated screening based on hierarchical routing and caching benchmarks; repairing abnormal data and optimizing low-quality samples; and finally, outputting high-quality data through three-level verification.

[0011] This application aims to address the issues of inconsistent screening standards and high reliance on manual intervention for long-tail data, achieving standardized and automated screening with significantly improved efficiency. It accurately perceives four types of boundaries: time sequence, frame rate, motion, and scene, eliminating temporal misalignment and ensuring data temporal and modal consistency. Furthermore, this application enables multi-dimensional anomaly identification and hierarchical repair, differentiated processing of abnormal data, improving data utilization and repair accuracy. A three-level verification system is employed as a fallback to prevent the spread of local defects, ensuring high reliability and high fidelity of output data and supporting improved model generalization capabilities.

[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0013] Figure 1 This document illustrates a flowchart of a method for filtering and processing multimodal data in long-tail scenarios of intelligent driving, provided in an embodiment of this application. Figure 2 This illustration shows a schematic diagram of a system for filtering and processing multimodal data in long-tail scenarios of intelligent driving, provided in an embodiment of this application. Detailed Implementation

[0014] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0015] The following is combined Figure 1 This application describes a method for filtering and processing multimodal data in long-tail scenarios of intelligent driving, based on exemplary embodiments thereof: S101: Obtain the original multimodal data features, construct a long-tail scene multimodal data codebook that integrates scene semantic features, target contour parameters, spatiotemporal location labels, working condition information, and environmental interference parameters, and perform feature similarity matching retrieval to generate a time-series data candidate filtering set.

[0016] In one implementation, to address the multimodal data filtering and processing needs of long-tail scenarios in intelligent driving, the first step is to acquire features from the raw multimodal data. Raw data from onboard multi-source sensors, including visual images, LiDAR point clouds, millimeter-wave radar, onboard audio, and vehicle CAN bus data, are accessed according to the sensor type and protocol, invalid empty packets are filtered, and the data is parsed into a standardized format. Then, through lightweight convolution, point cloud clustering, spectrum analysis, and numerical extraction, basic feature vector sequences such as image edge texture, point cloud shape distribution, radar distance and speed, audio temporal waveform, and vehicle operating condition values ​​are output, ensuring consistency in input for subsequent feature fusion and codebook construction. For example, if an onboard camera acquires 1080P color images, LiDAR outputs 16-line 3D point clouds, millimeter-wave radar outputs target signals, a microphone acquires temporal audio, and the CAN bus outputs vehicle speed and steering data, feature acquisition extracts image edge texture, point cloud shape distribution, radar distance and speed, audio temporal waveform, and vehicle operating condition numerical features, outputting a unified format of multimodal basic feature vector sequences to provide a unified foundation for codebook construction.

[0017] Multimodal basic features are fused to construct a long-tail scene multimodal data codebook. Extracted image, point cloud, radar, audio, and environmental condition features are processed using a Transformer encoder to extract scene semantic features such as roads, pedestrians, and obstacles. Target contour size and shape parameters are extracted from the image point cloud. Spatiotemporal location labels are generated based on GPS and time stamps, and environmental interference information such as vehicle status, illumination, rain, snow, and noise is fused. These features are then integrated into a unified fused feature set through feature concatenation and dimensionality compression. Clustering is then used to generate cluster centers with unique indices, constructing a data codebook that integrates scene semantics, target contours, spatiotemporal location, environmental condition status, and environmental interference information. This ensures consistency in input for subsequent similarity retrieval and candidate selection. For example, image semantic features, point cloud target parameters, GPS spatiotemporal labels, and illumination and environmental condition information are fused into a 512-dimensional feature vector. A codebook is generated through 1024 clustering classes, with each index corresponding to a fused feature class, forming a standardized long-tail scene multimodal data codebook that provides a unified benchmark for similarity retrieval.

[0018] Feature fusion employs an execution flow of "modal encoding + cross-modal attention concatenation + fully connected dimensionality compression": First, independent feature extraction networks encode scene semantic features into 128-dimensional vectors, target contour parameters into 64-dimensional vectors, spatiotemporal location labels into 64-dimensional vectors, working condition information into 64-dimensional vectors, and environmental interference parameters into 64-dimensional vectors; then, a cross-modal attention mechanism is used to calculate the association weights of the five types of features, and feature concatenation is performed according to the weights to obtain a 384-dimensional concatenated vector; finally, two fully connected networks are used for dimensionality compression to output a fixed 512-dimensional unified fused feature vector.

[0019] The codebook adopts an incremental clustering dynamic update mechanism: the initial codebook is generated by K-Means clustering based on historical long-tail scene samples, with a total of 1024 cluster centers, each center corresponding to a unique numerical index; every time 10,000 new high-quality samples are added, an incremental clustering update is performed, retaining 80% of the original high-frequency cluster centers and replacing 20% ​​of the low-frequency invalid centers. After the update, the codebook index size remains unchanged at 1024, and the mapping relationship between the index and features is updated synchronously.

[0020] The complete execution flow of feature similarity retrieval is as follows: Step 1, extract the 512-dimensional fused feature vector of the frame to be retrieved; Step 2, calculate the cosine similarity between this vector and all 1024 cluster center vectors of the codebook in parallel; Step 3, sort the similarity from high to low, and retain the top 5 codebook entries with similarity ≥ 0.6 as candidate matching items; Step 4, check the similarity difference between the candidate item of the current frame and the matching item of the previous frame. If the difference is > 0.3, it is judged as a mutation anomaly and the candidate item is removed; Step 5, bind the remaining valid candidate items with the frame number in chronological order to generate a time-series data candidate screening set.

[0021] A candidate set for time-series data is generated based on feature similarity matching retrieval. For the input multimodal fusion feature sequence to be processed and the constructed codebook, the feature matching degree is calculated frame-by-frame using the cosine similarity algorithm. A similarity threshold of 0.6 is set to filter highly matching codebook entries and sort them by similarity. Then, combined with temporal continuity constraints, abnormal results with inter-frame similarity mutations exceeding 0.3 are removed. Valid matching entries are integrated temporally to generate a candidate set for time-series data containing frame number, matching index, similarity, and type identifier, ensuring input consistency for subsequent graded screening and quality assessment. For example, if the similarity between the fusion feature of a certain time-series frame and codebook indices 12, 45, and 78 is 0.85, 0.72, and 0.68 respectively, all exceeding the threshold of 0.6, the result for index 45 with inter-frame mutations exceeding 0.3 is removed, while matching entries for indices 12 and 78 are retained. The candidate results for this frame are then integrated temporally to form a complete candidate set for time-series data, providing a unified basis for subsequent data quality grading.

[0022] S102, based on six core indicators—scene rarity, modal noise, data missing anomaly, spatiotemporal misalignment, target occlusion interference, and working condition matching error—completes sample-level data risk and quality normalization calculations, and achieves a four-level classification of invalid, inefficient, effective, and high-quality long-tail scene data.

[0023] In one implementation, to address the multimodal data filtering and processing needs of long-tail scenarios in intelligent driving, the first step is to acquire features from the raw multimodal data. For the raw multimodal data collected from the vehicle, a unified data access format is used, and basic features such as image edge texture, point cloud shape distribution, radar distance and speed, audio time-domain waveforms, and vehicle operating condition values ​​are extracted to form a unified feature vector sequence, ensuring consistent input for subsequent codebook construction. For example, the vehicle camera outputs 1080P color images, the LiDAR outputs 3D point cloud data, the millimeter-wave radar outputs target distance and speed information, the vehicle microphone outputs time-domain audio signals, and the vehicle CAN bus outputs operating condition data such as vehicle speed, steering angle, and throttle opening. The feature acquisition stage extracts basic features from each modal data, outputting a unified multimodal basic feature vector sequence, providing standardized input for subsequent feature fusion construction.

[0024] After acquiring basic features, a long-tail scene multimodal data codebook is constructed. Five types of features—image semantics, target contours, spatiotemporal location, operating conditions, and environmental interference—are fused and encoded. Feature concatenation and dimensionality compression are used to integrate them into a unified dimensional fusion feature. A fixed number of cluster centers are then generated using a clustering algorithm, with each cluster center corresponding to a unique index, forming a standardized multimodal data codebook to ensure consistent input for subsequent similarity retrieval. For example, semantic features of road pedestrians and obstacles, vehicle and pedestrian contour size parameters, GPS positioning and temporal location labels, vehicle speed and steering condition information, and illumination, rain, snow, and noise interference parameters are fused into a 512-dimensional feature vector. This vector is then used to generate a codebook through 1024 clustering classes, with each index corresponding to a fusion feature class. This constructs a long-tail scene multimodal data codebook that integrates multidimensional information, providing a unified benchmark for feature similarity matching.

[0025] Based on the constructed codebook, feature similarity matching retrieval is performed to generate a candidate set for time-series data. The similarity between the multimodal fusion features to be processed and the codebook features is calculated frame by frame. A fixed similarity threshold is set to filter high-matching entries. Then, combined with temporal continuity constraints, abnormal matching results are eliminated. Valid matching entries are integrated in chronological order to generate the candidate set for time-series data, ensuring consistency of input for subsequent graded screening. For example, if the similarity between the fusion features of a certain time-series frame and codebook indices 12, 45, and 78 is 0.85, 0.72, and 0.68 respectively, all exceeding the set threshold of 0.6, the result for index 45, where the inter-frame similarity mutation exceeds 0.3, is eliminated. Matching entries for indices 12 and 78 are retained, and candidate results for that frame are generated in chronological order, ultimately forming a complete candidate set for time-series data, providing a unified basis for subsequent data quality grading.

[0026] S103 combines temporal frame alignment, sensor frame rate calibration, motion posture change detection, and scene switching recognition to achieve long-tail scene boundary and working condition switching perception, and determines the effective screening interval of multimodal data and the stable constraints of non-boundary data that are temporally continuous, modally matched, and scene-fidelity preserved.

[0027] In one implementation, combining the boundary perception requirements of long-tail scenarios with the logic of operational condition switching recognition, a standardized perception scheme is established by introducing temporal alignment rules, frame rate calibration mechanisms, motion change detection, and scene switching judgment strategies, based on the requirements of boundary perception and operational condition switching in long-tail scenarios of intelligent driving. Firstly, a standardized perception scheme is developed to address the needs of boundary perception and operational condition switching in long-tail scenarios of intelligent driving. For four typical problems—multimodal data temporal misalignment, inconsistent frame rates, difficulty in recognizing sudden motion changes, and blurred scene boundaries—temporal alignment rules, frame rate calibration mechanisms, motion change detection algorithms, and scene switching judgment strategies are introduced. This establishes four core perception processes: temporal frame alignment, sensor frame rate calibration, motion attitude detection, and scene switching recognition, forming a reusable, quantifiable, and reproducible unified processing framework to ensure the standardization, consistency, and interpretability of boundary perception.

[0028] To address timing misalignment issues, the timing alignment rule employs an absolute timestamp synchronization strategy. Using a unified 10-millisecond time reference from the vehicle-mounted GNSS output as an anchor point, timestamp mapping and inter-frame registration are performed on the vision, LiDAR, millimeter-wave radar, vehicle audio, and CAN bus data respectively. For example, the original frame intervals for the vision sensor are 33 milliseconds, LiDAR 100 milliseconds, millimeter-wave radar 20 milliseconds, audio 1 millisecond, and CAN bus 10 milliseconds. These are all aligned to a 10-millisecond timing axis, and the timestamp deviation value Δt is calculated frame by frame using the formula Δt = |t_sensor t_reference|, when Δt>5 milliseconds, perform frame interpolation or discard processing to ensure that timing error is controlled within 5 milliseconds and eliminate modal timing misalignment.

[0029] To address the issue of inconsistent frame rates, the frame rate calibration mechanism employs a combination of dynamic interpolation and resampling. Using 50 frames per second as the unified standard timing, low frame rate data is padded with missing frames through linear interpolation, while high frame rate data is compressed by mean resampling to compress redundant frames. For example, the original frame rate of a LiDAR sensor is 10 frames per second, outputting one frame every 100 milliseconds, and then interpolated to reach 50 frames per second; the original frame rate of a millimeter-wave radar sensor is 50 frames per second, which is directly retained; and the original frame rate of a vision sensor is 30 frames per second, outputting one frame every 33 milliseconds, and then interpolated to reach 50 frames per second. Ultimately, the timing benchmark for all sensors is unified at 50 frames per second, ensuring timing continuity.

[0030] To address the difficulty in identifying sudden motion changes, the motion change detection algorithm employs a three-frame feature difference threshold method. It selects three core parameters—vehicle heading angle, speed, and lateral displacement—and calculates the inter-frame changes frame by frame. The formula is as follows: , , The system sets thresholds for speed abrupt changes at 20 km / h, heading angle abrupt changes at 15 degrees, and lateral displacement abrupt changes at 2 meters. For example, if the vehicle speed suddenly increases from 30 km / h to 55 km / h in three consecutive frames, ΔV = 25 km / h, exceeding the threshold, it is determined to be an acceleration abrupt change; if the heading angle suddenly increases from 0 degrees to 20 degrees, Δθ = 20 degrees, exceeding the threshold, it is determined to be a steering abrupt change, thus accurately identifying abrupt changes in the vehicle's dynamic operating conditions.

[0031] To address the issue of blurred scene boundaries, the scene transition determination strategy employs a dual-threshold fusion method combining semantic proportion and inter-frame similarity. It selects three semantic features: road type, environmental features, and target distribution. The difference in semantic proportion between consecutive frames is calculated, along with the feature cosine similarity, as shown in the formula: , A similarity threshold of 0.4 and a semantic proportion difference threshold of 30% are set. For example, if the semantic proportion of urban roads in the previous frame is 75% and that of rural roads in the next frame is 80%, the ΔP=55% exceeds the threshold, and the inter-frame similarity S=0.35 is below the threshold, it is determined to be a boundary between urban and rural scenes. If the semantic proportion of highways in the previous frame is 90% and that of highways in the next frame is 85%, the ΔP=5% and S=0.85 are determined to be non-boundary, thus accurately locating the boundary between long-tail scenes.

[0032] Timestamp synchronization and inter-frame registration are performed frame-by-frame on multimodal time-series data to complete time-series frame alignment and eliminate modal timing misalignment. This time-series frame alignment process aims to eliminate timing misalignment caused by inconsistent sampling times between different sensors, ensuring accurate correspondence between visual, LiDAR, millimeter-wave radar, vehicle audio, and vehicle CAN bus data at the same driving moment. This provides a consistent foundation of data for subsequent feature fusion, boundary perception, and data filtering. Time-series frame alignment uses the 10-millisecond absolute time output from the vehicle GNSS as a unified benchmark to establish a global time-series coordinate system. Timestamp parsing, deviation calculation, and inter-frame registration are performed frame-by-frame on the raw data from each sensor.

[0033] In practice, the physical timestamps of the original frames from each sensor are first extracted. The timestamp of each frame from the vision sensor is recorded as follows: The timestamp for each frame of the lidar is... The timestamp for each frame of the millimeter-wave radar is... The timestamp for each frame of the in-vehicle audio is as follows: The timestamp for each frame of the vehicle's CAN bus is as follows: Next, calculate the deviation of each frame relative to the unified time reference. The deviation calculation formula is as follows: ,in, For sensor frame timestamps, A unified timing reference point is defined at 10-millisecond intervals, and ΔT represents the timing deviation value.

[0034] For example, the original sampling frame rate of the visual sensor is 30 frames / second, with a single frame interval of approximately 33 milliseconds, and the original frame timestamps are 0 milliseconds, 33 milliseconds, and 66 milliseconds, respectively; the original sampling frame rate of the lidar is 10 frames / second, with a single frame interval of 100 milliseconds, and the original frame timestamps are 0 milliseconds, 100 milliseconds, and 200 milliseconds; the original sampling frame rate of the millimeter-wave radar is 50 frames / second, with a single frame interval of 20 milliseconds, and the original frame timestamps are 0 milliseconds, 20 milliseconds, and 40 milliseconds. The deviation of the frame timestamps of each sensor from the 10-millisecond unified benchmark is between 10 milliseconds and 50 milliseconds, with a maximum deviation of 50 milliseconds, indicating a significant timing misalignment.

[0035] To address the aforementioned timing discrepancies, a timestamp synchronization registration rule is adopted, setting a discrepancy threshold of 5 milliseconds. When ΔT ≤ 5 milliseconds, the frame is directly retained and mapped to the corresponding unified timing axis position; when ΔT > 5 milliseconds, a linear interpolation algorithm is used to generate a correction frame. The interpolation calculation formula is as follows: ,in , T_0 is the timestamp of the nearest valid frame before and after the deviation frame, and T_0 is the target reference timestamp. This is for correcting the frame timestamp.

[0036] After performing synchronous registration on the above example data, all sensor frames are aligned to a unified timing axis with 10-millisecond intervals of 0ms, 10ms, 20ms, 30ms, 40ms, and 50ms. The visual 33ms deviation frame is interpolated and corrected to the 30ms position, and the lidar 100ms deviation frame is interpolated and corrected to the 50ms and 100ms reference positions. The timing deviation of each sensor is controlled within 5ms, completely eliminating modal timing misalignment and achieving accurate correspondence of multimodal data at the same time. This provides timing-consistent input data for subsequent frame rate calibration, boundary perception, and other processing.

[0037] Dynamic interpolation and resampling calibration are performed based on the actual frame rate differences of each sensor to unify the timing benchmark for multimodal data. Sensor frame rate calibration for multimodal data aims to eliminate frame rate differences caused by different hardware designs and sampling periods of sensors such as vision, LiDAR, millimeter-wave radar, automotive audio, and vehicle CAN bus. This addresses issues such as discontinuous timing sequences, uneven data density, and asynchronous modal timing, providing a unified, continuous, and stable timing benchmark for subsequent timing alignment, feature fusion, and boundary perception.

[0038] Using the vehicle-mounted timing standard of 50 frames per second as the target benchmark, and based on the actual frame rate differences of each sensor, a combined calibration method is adopted, which involves dynamic interpolation to complete missing frames at low frame rates and average resampling to compress redundant frames at high frame rates. The frame rate of each modal data is normalized segment by segment to ensure that the frame rate of all sensor timing sequences is consistent, the intervals are uniform, and the timing is continuous after calibration.

[0039] When the sensor's original frame rate is below 50 frames per second, a linear interpolation algorithm is used to fill in the missing frames. The interpolation formula is as follows: ,in, , Features of adjacent valid frames , , This corresponds to the timestamp.

[0040] When the sensor's original frame rate is higher than 50 frames / second, a mean resampling algorithm is used. The mean of features is calculated at fixed intervals to generate the target frame. The resampling step size formula is as follows: ,in, The original frame rate of the sensor. =50 frames per second.

[0041] The original sensor frame rate differences are as follows: visual sensor: 30 frames / second, with a single frame interval of approximately 33 milliseconds; LiDAR: 10 frames / second, with a single frame interval of 100 milliseconds; millimeter-wave radar: 50 frames / second, with a single frame interval of 20 milliseconds; vehicle audio: 100 frames / second, with a single frame interval of 10 milliseconds.

[0042] The calibration process is as follows: Vision sensor (30 frames / second): Below 50 frames / second, interpolation is used to complete 20 missing frames / second, and one interpolated frame is inserted between the original frames every 33 milliseconds, resulting in a calibrated frame rate of 50 frames / second; LiDAR (10 frames / second): Below 50 frames / second, interpolation is used to complete 40 missing frames / second, and four interpolated frames are inserted between the original frames every 100 milliseconds, resulting in a calibrated frame rate of 50 frames / second; Millimeter-wave radar (50 frames / second): Equal to the target frame rate, directly retaining the original time-series frames without interpolation or resampling; In-vehicle audio (100 frames / second): Above 50 frames / second, the resampling step size is adjusted. =100 / 50=2, take the average of every 2 frames to generate 1 frame, compress redundant frames, and after calibration, the frame rate is 50 frames / second.

[0043] This study calculates changes in vehicle motion posture based on the differences in features across consecutive frames to identify abrupt changes in vehicle dynamic conditions. To accurately identify abrupt changes in dynamic conditions such as acceleration, deceleration, steering, and bumps during intelligent driving, and to avoid misjudgments of scene boundaries due to these changes, vehicle motion posture change detection is implemented. This detection is based on the multimodal feature differences across consecutive time-series frames, selecting three core motion parameters: vehicle heading angle, driving speed, and lateral displacement. By calculating the changes in these parameters frame by frame, the degree of motion posture fluctuation is quantified, accurately capturing key nodes of abrupt changes in conditions, and providing reliable motion feature data for scene boundary perception and condition switching identification. During detection, three consecutive frames of time-series data are selected as the calculation window. The heading angle θ, vehicle speed V, and lateral displacement X parameters are extracted from frames t-1 and t-2, respectively, and the inter-frame changes are calculated frame by frame. The core calculation formula is as follows: Vehicle speed change: Units are kilometers per hour; Change in heading angle: The unit is degrees; the change in lateral displacement is: The unit is meters.

[0044] Specific thresholds for abrupt changes are predefined: a speed change threshold of 20 km / h, a heading angle change threshold of 15 degrees, and a lateral displacement change threshold of 2 meters. When any parameter changes beyond its corresponding threshold, it is identified as a sudden change in the vehicle's dynamic operating condition, and the frame is marked as a key frame for motion attitude change. If multiple parameters exceed their limits simultaneously, it is identified as a strong operating condition change and is prioritized for inclusion in the boundary perception key analysis scope. For example, in three consecutive frames of time-series data, frame t-2 shows a speed of 30 km / h, a heading angle of 0 degrees, and a lateral displacement of 0 meters; frame t-1 shows a speed of 45 km / h, a heading angle of 8 degrees, and a lateral displacement of 1.2 meters; and frame t shows a speed of 55 km / h, a heading angle of 22 degrees, and a lateral displacement of 2.5 meters. The calculated values ​​of ΔV = 10 km / h, Δθ = 14 degrees, and ΔX = 1.3 meters were all within the limits, indicating a stable driving condition. In another set of three consecutive frames, the vehicle speed suddenly increased from 30 km / h to 55 km / h, with ΔV = 25 km / h, exceeding the 20 km / h threshold, indicating an acceleration change key frame. The heading angle suddenly changed from 0 degrees to 20 degrees, with Δθ = 20 degrees, exceeding the 15 degrees threshold, indicating a steering change key frame. The lateral displacement exceeded 2 meters, indicating a bumpy change key frame.

[0045] Through the above quantitative detection, different motion states such as smooth driving, slow speed change, sharp steering, and sudden bumps can be accurately distinguished, small normal fluctuations can be eliminated, and only sudden changes in real working conditions can be captured. This provides reliable motion feature support for subsequent accurate positioning of scene boundaries and data screening interval division, ensuring that the boundary perception results are highly matched with the actual operating state of the vehicle.

[0046] This method integrates scene semantic features with inter-frame similarity assessment to accurately locate the transition boundaries of long-tail scenes. It aims to address the problems of blurred boundaries and difficulty in identifying transition points in long-tail scenes such as urban areas, rural areas, highways, and parking lots, providing a reliable boundary basis for subsequent data screening, interval division, and stability constraint determination. The method uses scene semantic features as its core, combined with quantitative analysis of inter-frame similarity, and employs a dual-threshold fusion strategy to accurately identify scene transition critical points and distinguish between stable scene segments and transition boundary segments.

[0047] In practice, a pre-trained scene semantic extraction model is first used to parse multimodal data such as images and point clouds frame by frame, outputting three types of scene semantic features: road type, environmental features, and target distribution. These features are then quantified into normalized percentage values, ranging from 0 to 1. Simultaneously, fused feature vectors from adjacent frames are extracted, and the cosine similarity algorithm is used to calculate the inter-frame feature similarity. The calculation formula is as follows: ,in, For the current frame fusion features, The value of S is the fusion feature of the previous frame. The value of S ranges from 0 to 1. The lower the value, the greater the difference in features between frames.

[0048] The preset dual judgment thresholds are as follows: a scene semantic proportion difference threshold of 30% and an inter-frame similarity threshold of 0.4. When the scene semantic proportion difference between adjacent frames is ≥30% and the inter-frame similarity S≤0.4, it is judged as a scene transition boundary frame; if neither condition is met, it is judged as a stable scene frame. For example, in the scene semantic features of the previous frame, urban roads and buildings account for 65%, pedestrians for 20%, and rural vegetation for 15%; in the next frame, urban roads and buildings account for 10%, pedestrians for 5%, and rural vegetation for 85%. The calculated scene semantic proportion difference is 85%. 15% = 70%, which is greater than the 30% threshold; at the same time, the similarity of the fused features between the two frames is calculated as S = 0.35, which is lower than the 0.4 threshold. Therefore, this frame is determined to be a boundary frame for the transition between urban and rural scenes, and the boundary position is marked. For example, if the semantic content of highways in the previous frame is 90% and that in the next frame is 88%, the difference in semantic content is 2%, which is less than 30%. The inter-frame similarity is S = 0.82, which is higher than 0.4. Therefore, this frame is determined to be a stable highway scene frame, and the boundary is not marked.

[0049] By integrating timing alignment results, frame rate calibration parameters, motion attitude change information, and scene boundary signals, and considering both vehicle operating status and scene switching characteristics, the effective screening interval for multimodal data is precisely defined. Three types of stability constraints for non-boundary data—timing continuity, modal matching, and scene fidelity—are clearly defined, providing a clear scope and judgment criteria for subsequent hierarchical routing screening and ensuring the effectiveness and stability of data screening. Regarding interval division, based on the identified scene switching boundary frames, a boundary expansion and isolation strategy is adopted, setting a fixed number of transition buffers to avoid interference from unstable data near the boundary. Specifically, 5 frames are extended before and after the scene switching boundary frame, forming an 11-frame transition interval. Data within the transition interval is excluded from the effective screening range due to risks of scene ambiguity and modal fluctuations. Only time segments outside the transition interval with at least 10 consecutive frames are selected as stable and effective screening intervals, ensuring that the data within these intervals has a single scene and stable timing.

[0050] Regarding the setting of stability constraints, three quantitative judgment indicators are established, and threshold standards are defined to verify the quality of non-boundary data. Specifically, for the temporal continuity constraint: the similarity of fused features between adjacent frames is calculated using the following formula: To ensure smooth inter-frame temporal transitions without drastic jumps, a threshold of ≥0.7 is set for modal matching: the consistency ratio of features from the three modalities (computer vision, LiDAR, and millimeter-wave radar) is set to a threshold of ≥0.8 to ensure multimodal data synchronization without significant deviations. To ensure scene fidelity: the proportion of dominant scene semantics within a compute frame is set to a threshold of ≥0.9 to ensure scene uniformity without significant distortion or mixing.

[0051] For example, in a time series data segment, frame 30 is the boundary frame for the transition from urban roads to rural roads. According to the rules, frames 25 to 35 are designated as the transition interval and this segment of data is removed. Frames 1 to 24 are stable urban road segments, and frame-by-frame verification is performed: inter-frame feature similarity 0.82 ≥ 0.7, modality matching degree 0.89 ≥ 0.8, and scene semantic proportion 0.94 ≥ 0.9. All of these constraints are met, and the segment is determined to be a valid screening interval. Similarly, frames 36 to 60 are stable rural road segments, with verification indicators of 0.78, 0.85, and 0.91, respectively, all of which meet the standards and are included in the valid interval.

[0052] S104 establishes hierarchical routing and filtering decision rules based on data quality level and scenario boundary information, including direct retention, mild repair, modal completion, sample removal, and batch filtering, and builds a data filtering benchmark library based on the historical high-quality data caching module.

[0053] In one implementation, a hierarchical routing and filtering decision framework is constructed based on data quality levels and scenario boundary information. This framework aims to address the problems of inconsistent filtering standards and chaotic decision-making caused by the complex data types, significant quality differences, and variable boundary conditions in long-tail scenarios of intelligent driving. It establishes a standardized, hierarchical, and reusable filtering process, ensuring that precise processing strategies can be matched to data of different scenarios and quality levels, thereby improving filtering efficiency and result reliability. The framework adopts a hierarchical, interconnected structure, consisting of a data classification and judgment layer, a boundary constraint verification layer, a routing rule mapping layer, and a cache benchmark support layer. These four modules are sequentially connected and work together to form a complete link from data evaluation to result output. The data classification and judgment layer is responsible for quantifying data quality and classifying levels; the boundary constraint verification layer verifies the compliance of scenario boundaries; the routing rule mapping layer matches processing strategies; and the cache benchmark support layer provides high-quality reference data. The outputs of each module are inputs to each other, ensuring consistent process loop.

[0054] The data grading and judgment layer is based on six core indicators: scene rarity, modal noise, data missing anomaly, spatiotemporal misalignment, target occlusion interference, and working condition matching error. A normalized weighted algorithm is used to calculate the comprehensive quality score of the samples, with the formula: S = 0.2 × R + 0.2 × N + 0.2 × L + 0.2 × D + 0.1 × O + 0.1 × E, where R is rarity, N is noise, L is missing data, D is misalignment, O is occlusion, and E is error, with a score range of 0-1. Preset grade thresholds are: S ≥ 0.8 for excellent, 0.6 ≤ S < 0.8 for effective, 0.4 ≤ S < 0.6 for inefficient, and S < 0.4 for invalid, completing the four-level quality classification. The individual quantitative calculation rules for the six core indicators are as follows: Scene Rarity R: Based on historical full-data scene database statistics, calculate the frequency ratio of the current scene type in the full set of scenes. The rarity formula is R = 1. The current scene's frequency of occurrence / the total frequency of all scenes, ranging from 0 to 1. A higher value indicates a rarer scene. The rarity of ordinary urban road scenes is 0.2 by default, while long-tail scenes such as extreme weather (rain, snow, fog), tunnel construction areas, and narrow rural roads have a rarity of ≥0.6.

[0055] Modal noise level N: The noise values ​​for the four modes—image, point cloud, radar, and audio—are calculated separately, and the weighted average is taken. The formula is N = 0.4 × N_img + 0.3 × N_lidar + 0.2 × N_radar + 0.1 × N_audio. Where N_img is the normalized value of the image pixel noise variance, N_lidar is the normalized value of the effective point density deviation of the lidar, N_radar is the normalized value of the reciprocal of the signal-to-noise ratio of the millimeter-wave radar, and N_audio is the normalized value of the audio waveform distortion rate, ranging from 0 to 1; the higher the value, the more severe the noise.

[0056] Data Missing Anomaly Level L: Calculates the proportion of frames with missing data in the current sample time series frames out of the total number of frames. The formula is L = number of missing frames / total number of frames, with a value ranging from 0 to 1. Higher values ​​indicate more severe data loss. Single-modal missing data is calculated using a coefficient of 0.5, while multi-modal synchronous missing data is calculated using a coefficient of 1.0.

[0057] Spatiotemporal misalignment deviation D: Calculated by weighting time and spatial deviations, using the formula D = 0.5 × ΔT_norm + 0.5 × ΔD_norm. Where ΔT_norm is the normalized value of the maximum timestamp deviation among the multimodal components, with a reference deviation of 100ms; ΔD_norm is the normalized value of the maximum offset of the multimodal target spatial coordinates, with a reference offset of 10m, and its value ranges from 0 to 1. Higher values ​​indicate more severe misalignment.

[0058] Target occlusion interference degree O: This is the percentage of the total target area where key targets (pedestrians, vehicles, obstacles) are occluded. The formula is O = Total occluded area of ​​occluded targets / Total area of ​​key targets. The value ranges from 0 to 1, with higher values ​​indicating more severe occlusion. For mild occlusion (occlusion percentage <30%), the normalized value is ≤0.3; for severe occlusion (occlusion percentage ≥70%), the normalized value is ≥0.7.

[0059] The operating condition matching error degree E is calculated as follows: This is the deviation rate between the vehicle operating condition parameters (vehicle speed, heading angle, acceleration) in the current frame and the typical operating condition parameters of the scene. The formula is E = (vehicle speed deviation rate + heading angle deviation rate + acceleration deviation rate) / 3. The deviation rate is the absolute value of the difference between the actual value and the typical value divided by the typical value. The value ranges from 0 to 1, with higher values ​​indicating worse operating condition matching. For example, for a rare road condition sample during a rainstorm: rarity 0.9, noise level 0.7, missing value 0.1, misalignment 0.2, occlusion 0.3, and error level 0.2, the calculated S = 0.72, which is considered a valid level.

[0060] The boundary constraint verification layer takes temporal alignment results, frame rate calibration parameters, motion abrupt change signals, and scene boundary labels as input to verify boundary compliance. Boundary constraint rules are set as follows: temporal continuity ≥ 0.7, modal matching ≥ 0.75, and scene semantic fluctuation ≤ 20% within 5 frames before and after the boundary; otherwise, it is marked as a boundary anomaly. For example, the temporal continuity of the boundary frame before the transition from urban road to highway is 0.82, modal matching is 0.85, and semantic fluctuation is 12%, which is compliant; the temporal continuity of the boundary frame in extreme weather is 0.58, and it is marked as an abnormal boundary.

[0061] The routing rule mapping layer establishes a four-dimensional mapping rule of "quality level - boundary type - anomaly type - processing strategy", and presets five processing paths: high-quality + non-boundary + no anomalies -- direct retention; effective + non-boundary + slight noise -- mild repair; inefficient + boundary + short-term missing -- modal completion; invalid + boundary + severe misalignment -- sample removal; batch inefficient anomalies -- batch filtering. For example, inefficient level, urban / village boundary, and short-term missing samples are matched with the modal completion strategy; invalid level, extreme weather boundary, and severe misalignment samples are removed.

[0062] The caching benchmark support layer relies on a historical high-quality data cache library to store verified high-quality sample features, quality labels, and boundary attributes. It uses a cosine similarity algorithm to calculate the matching degree between the sample to be processed and the cache benchmark, with the formula as follows: A matching degree ≥ 0.7 is considered a high match, and the baseline processing strategy can be directly reused; 0.5 ≤ M < 0.7 is considered a medium match, and the strategy needs to be fine-tuned; M < 0.5 is considered a low match, and new processing rules need to be added. For example, the matching degree between the rainstorm sample to be processed and the cached rainstorm baseline is 0.78, and a mild repair strategy can be reused; the matching degree of the rare tunnel sample is 0.42, and a new special repair rule needs to be added.

[0063] This framework enables the unified execution of quality grading, boundary verification, rule matching, and benchmark comparison for various long-tail scenarios, such as urban intersections, extreme weather conditions like rain and snow, rare road conditions in tunnels, narrow rural roads, and highway construction. This avoids confusion in screening logic and inconsistent standards due to scenario differences, achieving standardized and regulated screening across the entire process and providing accurate decision-making basis for subsequent data repair and optimization.

[0064] Conducting multi-dimensional compliance verification of samples is a crucial quality control step before hierarchical routing screening. It aims to comprehensively verify sample compliance across four core dimensions: quality, boundary, time series, and modality. This precise verification accurately identifies data defects, providing accurate quantitative evidence for subsequent anomaly identification and routing decisions, and preventing unqualified data from entering subsequent processing flows. The verification covers four key dimensions: quality level, boundary type, time series continuity, and modality matching. It employs a combination of quantitative scoring and threshold judgment, analyzing three types of time series features—modal integrity, feature credibility, and operational condition adaptability—frame by frame to achieve refined data quality verification.

[0065] Based on six core indicators—scene rarity, modal noise, data missing anomaly, spatiotemporal misalignment, target occlusion interference, and working condition matching error—a comprehensive sample quality score is calculated, categorized into four levels: invalid, inefficient, valid, and excellent. The scoring formula is: S = 0.2 × R + 0.2 × N + 0.2 × L + 0.2 × D + 0.1 × O + 0.1 × E, where R is rarity, N is noise, L is missing data, D is misalignment, O is occlusion, and E is error, ranging from 0 to 1. S < 0.4 is invalid, 0.4 ≤ S < 0.6 is inefficient, 0.6 ≤ S < 0.8 is valid, and S ≥ 0.8 is excellent.

[0066] To determine whether a sample's temporal interval falls within a scene transition boundary, three boundary types are distinguished: transition boundary, transition interval, and stable region. Boundary determination is based on a semantic difference between frames ≥30% and a feature similarity ≤0.4. The first five frames before and after the boundary are considered the transition interval, and the rest are considered the stable region. The cosine similarity of fused features between adjacent frames is calculated using the following formula: , For the current frame features, The feature value is from the previous frame, ranging from 0 to 1. A value ≥0.7 indicates temporal continuity, and <0.7 indicates temporal anomalousness. The consistency ratio of features from the three modalities (computer vision, LiDAR, and millimeter-wave radar) is also considered. A value ≥0.8 indicates modal matching, and <0.8 indicates modal mismatch.

[0067] Frame-by-frame analysis of modal integrity, feature reliability, and operating condition adaptability: Modal integrity checks whether there are missing or broken frames in each modal data, and a completeness of ≥0.9 is considered acceptable; Feature reliability checks the feature signal-to-noise ratio and clarity, and a reliability of ≥0.8 is considered reliable; Operating condition adaptability checks the matching degree between the data and vehicle speed, steering angle, and other operating conditions, and an adaptability of ≥0.8 is considered a match.

[0068] During the verification process, each modal data is checked frame by frame to ensure its completeness, feature reliability, and operational condition matching. A frame-by-frame verification report is generated to accurately capture detailed data quality issues, providing a precise basis for subsequent anomaly identification and routing decisions. For example, a time-series sample of a city road has a comprehensive quality score of 0.72, classifying it as inefficient; the boundary type is a non-switching stable region; the average frame-by-frame temporal continuity is 0.8, ≥0.7, indicating temporal compliance; the average modal matching degree is 0.7, <0.8, indicating slight modal mismatch. Frame-by-frame analysis reveals: image modal completeness of 0.95, no missing data; point cloud feature reliability of 0.9, low noise interference; vehicle operational condition fit of 0.8, matching vehicle speed and steering, accurately capturing slight modal mismatch defects and providing a basis for subsequent minor repair routing decisions.

[0069] Anomalies, including those with temporal misalignment, modal missingness, and scene distortion, are identified using rarity weights, noise interference levels, and spatiotemporal bias values. Rarity weights measure the rarity of a sample in a long-tailed scene, ranging from 0 to 1; higher values ​​indicate rarer scenes and higher anomaly risk. Noise interference levels quantify the noise pollution level of modal data such as images, point clouds, and radar data, also ranging from 0 to 1; higher values ​​indicate more severe noise. Spatiotemporal bias values ​​characterize the degree of deviation in timestamp synchronization and spatial coordinate matching of multimodal data, ranging from 0 to 1; higher values ​​indicate more significant temporal misalignment and spatial mismatch.

[0070] Anomaly identification employs weighted quantization calculation, with the comprehensive anomaly score formula as: A = 0.4 × R + 0.3 × N + 0.3 × D, where R is the rarity weight, N is the noise interference level, and D is the spatiotemporal deviation value. Anomalies are categorized into levels based on the comprehensive score: 0.7 ≤ A ≤ 1.0 indicates severe anomalies, 0.4 ≤ A < 0.7 indicates moderate anomalies, and 0 ≤ A < 0.4 indicates minor anomalies. Simultaneously, individual indicator thresholds are used to determine anomaly types: D ≥ 0.5 indicates temporal misalignment, N ≥ 0.5 indicates modality loss, and R ≥ 0.5 and D < 0.5 indicates scene distortion.

[0071] For example, a rare road condition sample during a rainstorm has a rarity weight of 0.9, noise interference of 0.7, and spatiotemporal deviation of 0.8. Substituting these values ​​into the formula, the anomaly score A = 0.4 × 0.9 + 0.3 × 0.7 + 0.3 × 0.8 = 0.75, which is considered a severe anomaly. The noise interference of 0.7 ≥ 0.5 and the spatiotemporal deviation of 0.8 ≥ 0.5 indicate a combined anomaly of temporal misalignment and modal loss. Another sample of a normal urban road condition has a rarity weight of 0.8, noise interference of 0.2, and spatiotemporal deviation of 0.1. The anomaly score A = 0.4 × 0.8 + 0.3 × 0.2 + 0.3 × 0.1 = 0.41, which is considered a moderate anomaly. The rarity weight of 0.8 ≥ 0.5 and the spatiotemporal deviation of 0.1 < 0.5 indicate a slight scene distortion anomaly.

[0072] By quantifying and classifying three indicators, different types and degrees of anomalies can be accurately distinguished, avoiding misjudgment based on a single indicator. This provides quantitative support for differentiated routing decisions such as subsequent minor repair, modal completion, and sample removal, ensuring the accuracy of anomaly identification and the targeted nature of processing.

[0073] Based on a historical high-quality data caching module, a standardized screening benchmark library is constructed to provide a unified comparison basis for hierarchical routing screening. The benchmark library stores high-quality multimodal samples that have passed end-to-end verification. Each sample is bound to four types of standardized labels: first, quality labels, indicating four levels: invalid, inefficient, effective, and high-quality; second, boundary labels, indicating stable areas, transition areas, and switching boundaries; third, anomaly labels, indicating the anomaly types and degrees such as temporal misalignment, modality loss, and scene distortion; and fourth, feature labels, storing information such as fused feature vectors and key parameter values ​​to ensure that the benchmark data is traceable, comparable, and reusable.

[0074] When constructing the benchmark library, it is categorized and stored according to scene type, operating conditions, and interference level, covering long-tail scenes such as urban areas, rural areas, highways, rain and snow, and tunnels. Each scene category stores no less than 500 high-quality samples to ensure comprehensive coverage and strong representativeness of the benchmark. At the same time, a dynamic update mechanism is established to regularly add new and qualified samples and remove outdated and invalid data to maintain the timeliness and reliability of the benchmark library.

[0075] The benchmark library's dynamic update and capacity management rules are as follows: The total capacity of the benchmark library is fixed at 100,000 high-quality samples. Storage quotas are allocated according to scenario type: 30% for urban roads, 20% for highways, 20% for rural / mountain roads, 15% for extreme weather, and 15% for special working conditions. A full update is performed every 7 days to add new high-quality samples that have passed end-to-end verification. At the same time, the "first-in, first-out + low-matching priority elimination" rule is used to eliminate existing data: old samples that have been in the library for more than 180 days are eliminated first, and then low-value samples with a historical matching call frequency lower than 30% of the average level are eliminated to ensure that the total capacity does not exceed the limit and the representativeness of the samples is continuously optimized.

[0076] The complete batch filtering process is as follows: Step 1, calculate the percentage of invalid and inefficient samples in a single batch (default 100 frames per batch); Step 2, if the percentage is ≥30%, trigger batch filtering, mark the entire batch of data as unqualified, and do not proceed to the subsequent repair and optimization process; Step 3, automatically trace back the corresponding collection section, collection time, and sensor operating condition information for the batch, and generate a batch unqualified reason report; Step 4, return the unqualified batch to the collection stage, mark the scenarios and operating condition requirements that need to be recollected, and avoid the recurrence of similar problems.

[0077] In the routing decision-making stage, four key pieces of information are integrated: data quality level, scenario boundary information, anomaly type, and cache matching degree. A multi-dimensional weighted decision-making algorithm is used to generate the filtering results. The cache matching degree is calculated using cosine similarity, and the formula is: The value ranges from 0 to 1. A higher value indicates a higher similarity and better match between the sample to be processed and the baseline sample. Preset matching degree levels: M ≥ 0.8 is a high match, 0.5 ≤ M < 0.8 is a medium match, and M < 0.5 is a low match.

[0078] The hierarchical routing rules specify five handling paths: high-quality, unstable boundary, no anomalies, and high match (M≥0.8) samples are directly retained; effective / inefficient, stable boundary, slight noise / distortion, and medium match (0.5≤M<0.8) samples undergo mild repair; inefficient, transition zone, short-term missing / slight misalignment, and medium match samples undergo modal completion; invalid, switching boundary, severe missing / misalignment, and low match (M<0.5) samples are removed; and if the proportion of inefficient anomalies in the same batch exceeds 30%, the batch is filtered.

[0079] For example, urban stable road segment samples, with a quality level of excellent, non-boundary, no anomalies, and a cache match rate of 0.9, meet the direct retention rule and are judged as qualified samples, directly entering subsequent processing; rural transition zone samples, with a quality level of inefficient, boundary, slight noise, and a cache match rate of 0.7, are matched with a light repair path and noise filtering optimization is performed; rainstorm transition boundary samples, with a quality level of invalid, boundary, severe modal missing, and a cache match rate of 0.2, are judged as unqualified and samples are removed; 35 frames out of 100 tunnel samples are inefficient and anomaly-triggered, triggering batch filtering rules, and the entire batch is removed and re-collected. This decision logic can accurately adapt to long-tail samples with different quality, boundary, and anomaly levels, avoiding misjudgment based on a single standard, ensuring the targeting and effectiveness of screening, and providing accurate decision support for subsequent data repair and optimization.

[0080] S105 performs noise filtering and removal, missing data interpolation and completion, and spatiotemporal offset correction operations on modal anomaly data. For low-quality long-tailed samples, it performs feature weight fine-tuning, redundant information removal, interference parameter suppression, and multimodal consistency forced alignment processing.

[0081] In one implementation, a multimodal anomaly repair fusion mechanism is constructed based on modal anomaly identification, missing data localization, and spatiotemporal deviation calculation. Noise detection results are prioritized, while missing data localization and spatiotemporal deviation are used for collaborative correction of complex anomaly scenarios. Addressing the multimodal data anomaly correction requirements of intelligent driving, this mechanism first constructs a multimodal anomaly repair fusion mechanism. This mechanism uses modal anomaly identification, missing data localization, and spatiotemporal deviation calculation as its three fundamental supporting modules. It clarifies the processing logic of prioritizing noise detection and the collaborative correction of missing data localization and spatiotemporal deviation, establishing a standardized, hierarchical, and collaborative framework for complex anomaly correction. This ensures that both single and complex anomalies can be processed in an orderly, efficient, and accurate manner, avoiding correction failure due to anomaly superposition.

[0082] The modal anomaly identification module focuses on noise types in multimodal data such as images, point clouds, radar, and audio. It employs pixel-level variance analysis, point cloud density statistics, and signal-to-noise ratio (SNR) calculation methods to quantify the degree of noise interference. Noise judgment thresholds are set: image pixel variance ≥ 30, point cloud density deviation ≥ 25%, and radar SNR ≤ 15dB are considered modal noise anomalies. The module outputs the noise type and interference level value as the basis for priority processing.

[0083] The missing data localization module addresses issues such as frame breaks, sensor packet loss, and localized data gaps by comparing the theoretical sampling amount with the actual data amount frame by frame to calculate the missing rate: L = (1 Actual data volume / theoretical sampling volume) × 100%. Missing data is categorized as follows: 5% ≤ L < 20% is considered short-term missing data, and L ≥ 20% is considered severe missing data. This accurately locates the start and end frames and the missing region, providing a positional basis for completion.

[0084] The spatiotemporal deviation calculation module uses a unified time base and spatial coordinate system as a reference to calculate the timestamp deviation and spatial coordinate offset of multimodal data. The time deviation formula is as follows: Spatial deviation formula: Set deviation thresholds: 5ms≤ΔT<50ms and 1m≤ΔD<5m are considered mild deviations, while ΔT≥50ms and ΔD≥5m are considered severe deviations, quantifying the degree of spatiotemporal misalignment.

[0085] Prioritize handling modal noise anomalies, then simultaneously perform missing data completion and spatiotemporal correction. Single anomalies are handled independently, while complex anomalies are handled collaboratively to avoid processing conflicts. For complex anomaly scenarios, a step-by-step collaborative correction algorithm is employed. The first step uses Gaussian filtering to remove noise, the second step uses linear interpolation to complete missing data, and the third step uses timestamp re-registration to correct spatiotemporal offsets. These three steps are linked and do not interfere with each other, ensuring consistent correction. For example, a complex anomaly may occur in vehicle-mounted multimodal data: the image contains salt-and-pepper noise (pixel variance 35, reaching the noise threshold), three consecutive short-term missing frames from millimeter-wave radar (missing rate 12%), and spatiotemporal offsets of 60ms and 2.5m between the image and point cloud. The repair mechanism prioritizes noise correction, using a 3×3 Gaussian filter to remove salt-and-pepper noise, reducing the noise interference level to below 8. Simultaneously, missing data completion is performed, using linear interpolation based on radar data from previous and subsequent frames to complete three frames of missing data, reducing the missing rate to 1%. At the same time, spatiotemporal correction is carried out, re-registering timestamps and correcting spatial coordinates at 10ms intervals, reducing the spatiotemporal deviation to 8ms and 0.3m, completing the collaborative correction of composite anomalies, ensuring that the data quality meets the standards, and providing reliable input for subsequent feature extraction and scene analysis.

[0086] Multi-dimensional anomaly labels are generated for each sample, marking noisy modal frames, missing data segments, spatiotemporal offset frames, and samples with distorted features as objects to be repaired, while the remaining samples are identified as low-quality optimization objects. Full-dimensional anomaly detection is performed on all multi-modal samples, generating multi-dimensional anomaly labels for each sample. This aims to accurately identify anomaly types and differentiate processing objects, providing a clear scope and basis for subsequent differentiated repair and optimization. Anomaly detection covers four core dimensions: modal noise, missing data, spatiotemporal offset, and distorted features. Quantitative judgment criteria are set for each item, and anomaly details are analyzed frame by frame and modally by modality to ensure that the labels are accurate, traceable, and reusable.

[0087] Modal noise detection targets four types of data: image, LiDAR, millimeter-wave radar, and vehicle audio. Quantization thresholds are set as follows: image pixel noise variance ≥ 30 is considered a noisy frame; LiDAR effective point density deviation ≥ 25% is considered a noisy frame; millimeter-wave radar signal-to-noise ratio ≤ 15dB is considered a noisy frame; and audio temporal waveform distortion rate ≥ 20% is considered a noisy frame. The degree of noise pollution is accurately quantified through calculations using pixel variance, point cloud density, signal-to-noise ratio, and waveform distortion rate.

[0088] Data missing detection compares the theoretical sample size with the actual data size frame by frame to calculate the missing rate: L = (1 Actual data volume / theoretical sampling volume) × 100%. Grading standards are set: 5% ≤ L < 20% is considered short-term missing data, and L ≥ 20% is considered severe missing data. The start and end frames and missing modes are precisely marked. Spatiotemporal offset detection uses the vehicle-mounted GNSS unified time reference and the world coordinate system as a reference. The time deviation formula is... Spatial deviation formula A grading standard is set: 5ms ≤ ΔT < 50ms or 1m ≤ ΔD < 5m indicates mild offset, and ΔT ≥ 50ms or ΔD ≥ 5m indicates severe offset, quantifying the degree of spatiotemporal misalignment. Feature distortion detection calculates feature fidelity: F = percentage of effective features / percentage of total features. A grading standard is set: 0.5 ≤ F < 0.8 indicates mild distortion, and F < 0.5 indicates severe distortion, evaluating feature completeness and reliability.

[0089] Anomaly labels use a standardized encoding format, containing six pieces of information: sample number, time frame range, anomaly type, anomaly severity, anomaly modality, and anomaly value, ensuring that the label information is complete, parsable, and traceable. Objects exhibiting any of the following anomalies—modal noise, short-term / severe missing data, mild / severe offset, or severe distortion—are marked as objects requiring repair; objects without the above anomalies but with a performance of 0.8 ≤ F < 0.9 are marked as low-quality optimization objects.

[0090] For low-quality optimization targets, the specific execution rules for the four types of optimization operations are as follows. Specifically, feature weight fine-tuning: adjust the weight coefficients of each dimension of features based on scene rarity. For scenes with higher rarity, increase the weight of scene semantic features and target contour features by 10%-20%; reduce the weight of non-core features such as environmental interference and operating condition fluctuations by 5%-10% to strengthen the contribution of core features in long-tail scenes.

[0091] Redundant information removal: Identify feature dimensions with a variance below 0.01 by calculating the feature variance. These dimensions are considered redundant and their feature data is directly removed. At the same time, remove highly similar inter-frame features (similarity ≥ 0.95) that appear repeatedly, retain the features of key temporal frames, reduce feature redundancy, and improve the efficiency of subsequent processing.

[0092] Interference parameter suppression: For non-core interference parameters such as changes in illumination and road surface bumps, a normalization constraint method is adopted to compress the range of numerical fluctuations of interference parameters to 50% of the original range, thereby weakening the impact of interference parameters on feature matching and quality judgment, and highlighting the core features of the scene and the target.

[0093] Multimodal consistency forced alignment: Based on the spatial coordinates of the lidar point cloud, the spatial positions of the image target detection box and the millimeter-wave radar target point are calibrated one by one. The target coordinates of all modes are unified to the vehicle body coordinate system through the coordinate transformation matrix. At the same time, the timestamps of all modes are aligned with a unified time base of 10ms to ensure that the target parameters of each mode correspond one-to-one at the same time and eliminate the spatiotemporal deviation between modes.

[0094] For example, sample number S020, time frame 10-15, image pixel noise variance 38 (≥30), LiDAR point density deviation 32% (≥25%), millimeter-wave radar data missing rate 18% (5%≤L<20%), image and radar spatiotemporal offset 65ms (≥50ms), anomaly labels are generated: S020, 10-15 frames, modal noise + short-term missing + severe offset, severe, image / radar, variance 38 / deviation 32% / missing 18% / offset 65ms, marked as an object to be repaired.

[0095] For example, sample number S021, time frame 1-50, no noise, no missing data, no offset, feature fidelity 0.85 (0.8≤F<0.9), generates anomaly labels: S021, 1-50 frames, no anomaly, slight, full modality, fidelity 0.85, and is marked as a low-quality optimization object.

[0096] Through full-dimensional quantitative detection and standardized label generation, it is possible to accurately distinguish between objects to be repaired and low-quality objects to be optimized, clarify different processing ranges, avoid ineffective processing or over-repair, and ensure the targeted and efficient processing of subsequent noise filtering, missing completion, offset correction, feature optimization and other processes.

[0097] A linkage rule was established between anomaly types and remediation strategies. Targeted corrections were performed on anomalous data, and feature optimization was implemented for low-quality samples. Remediation authority allocation and optimization constraints were solidified. Comprehensive anomaly detection was conducted on all multimodal samples, generating multi-dimensional anomaly labels for each sample. This aimed to accurately identify anomaly types and differentiate processing targets, providing a clear scope and basis for subsequent differentiated remediation and optimization. Anomaly detection covered four core dimensions: modal noise, missing data, spatiotemporal offset, and feature distortion. Quantitative judgment criteria were set for each item, and anomaly details were analyzed frame-by-frame and modally to ensure accurate, traceable, and reusable labels.

[0098] Modal noise detection targets four types of data: image, LiDAR, millimeter-wave radar, and vehicle audio. Quantization thresholds are set as follows: image pixel noise variance ≥ 30 is considered a noisy frame; LiDAR effective point density deviation ≥ 25% is considered a noisy frame; millimeter-wave radar signal-to-noise ratio ≤ 15dB is considered a noisy frame; and audio temporal waveform distortion rate ≥ 20% is considered a noisy frame. The degree of noise pollution is accurately quantified by calculating pixel variance, point cloud density, signal-to-noise ratio, and waveform distortion rate. Data missing detection compares the theoretical sampling amount with the actual data amount frame by frame to calculate the missing rate: L = (1 Actual data volume / theoretical sampling volume) × 100%. Set grading criteria: 5% ≤ L < 20% is short-term missing, L ≥ 20% is severe missing, and accurately mark the start and end frames and missing modalities.

[0099] Spatiotemporal offset detection uses the vehicle-mounted GNSS unified time reference and the world coordinate system as a reference. The time deviation formula is as follows: Spatial deviation formula The following grading standards are set: 5ms ≤ ΔT < 50ms or 1m ≤ ΔD < 5m indicates mild offset, and ΔT ≥ 50ms or ΔD ≥ 5m indicates severe offset, quantifying the degree of spatiotemporal misalignment. Feature distortion detection calculates feature fidelity: F = effective feature percentage / total feature percentage. A grading standard is set: 0.5 ≤ F < 0.8 indicates mild distortion, and F < 0.5 indicates severe distortion, assessing feature completeness and reliability.

[0100] Anomaly labels employ a standardized encoding format, containing six pieces of information: sample number, time frame range, anomaly type, anomaly severity, anomaly modality, and anomaly value. This ensures the label information is complete, parsable, and traceable. If any of the following conditions are present: modal noise, short-term / severe missing data, mild / severe offset, or severe distortion, the object is marked as requiring repair. If none of the above anomalies are present but 0.8 ≤ F < 0.9, the object is marked as low-quality requiring optimization. For example, sample number S020, time frame 10-15, image pixel noise variance 38 (≥30), LiDAR point density deviation 32% (≥25%), millimeter-wave radar data missing rate 18% (5% ≤ L < 20%), image and radar spatiotemporal offset 65ms (≥50ms), generates the following anomaly labels: S020, 10-15 frames, modal noise + short-term missing data + severe offset, severe, image / radar, variance 38 / deviation 32% / missing data 18% / offset 65ms, and is marked as requiring repair. For example, sample number S021, time frame 1-50, no noise, no missing data, no offset, feature fidelity 0.85 (0.8≤F<0.9), generates anomaly labels: S021, 1-50 frames, no anomaly, slight, full modality, fidelity 0.85, and is marked as a low-quality optimization object.

[0101] Through full-dimensional quantitative detection and standardized label generation, it is possible to accurately distinguish between objects to be repaired and low-quality objects to be optimized, clarify different processing ranges, avoid ineffective processing or over-repair, and ensure the targeted and efficient processing of subsequent noise filtering, missing completion, offset correction, feature optimization and other processes.

[0102] After anomaly correction and feature optimization are completed, the effectiveness of the repair must be verified to ensure that the processed data meets quality standards, is time-series consistent, and has reliable features. Verification focuses on three core indicators: modal integrity, spatiotemporal consistency, and feature fidelity. Simultaneously, time-series mapping and transition zone expansion are performed to ensure accurate matching and smooth transition between the repaired and original time series. Modal integrity verification checks whether the data from each sensor is complete and without omissions after repair. The modal integrity rate is calculated as: M = (Number of valid data frames / Theoretical total number of frames) × 100%. A threshold is set: M ≥ 95% is considered acceptable. For example, after missing data completion, a sample millimeter-wave radar has a theoretical frame count of 50 frames and a valid frame count of 49 frames, resulting in a modal integrity rate of 98%, which is considered complete and compliant.

[0103] Spatiotemporal consistency verification is used to verify the accuracy of temporal alignment and spatial registration. Time deviation: Spatial deviation: Qualification Standard: ≤10ms ≤0.5m. For example, after offset correction, the image and point cloud have a temporal deviation of 8ms and a spatial deviation of 0.3m, meeting the spatiotemporal consistency requirements. Feature fidelity verification is used to evaluate the similarity between the repaired features and the original valid features, using cosine similarity: A threshold F ≥ 0.85 is set as acceptable. For example, if the feature similarity of the image after noise correction is 0.92, it is determined that the features are faithful and there is no obvious distortion.

[0104] After verification, all correction results are mapped back to the original time-series frames, maintaining a one-to-one correspondence between frame numbers and timestamps to avoid timing misalignment. To prevent jumps or abrupt changes between the abnormal repair segment and the normal segment, a repair transition zone is formed by extending two frames before and after the abnormal segment. The data in the transition zone uses linear smooth interpolation to ensure temporal continuity and gradual feature transition. For example, if frames 10–15 of a certain time-series sample are a composite abnormal segment of noise and missing data, after repair, the modal integrity rate is 98% ≥ 95%, which is acceptable; the spatiotemporal deviation is ΔT = 7ms and ΔD = 0.2m, which is acceptable; and the feature fidelity is 0.90 ≥ 0.85, which is acceptable. Mapped to the original frames 10–15, two frames are extended before and after (frames 8–9 and 16–17) as transition zones, using linear interpolation for smooth transition, resulting in no temporal jumps and continuous modal matching.

[0105] By performing validity checks, time-series mapping, and transition zone expansion, we can ensure that the repaired data modal is complete, spatiotemporally aligned, and features are faithfully preserved. At the same time, it accurately matches and smoothly connects with the original time series, providing compliant, reliable, and time-series consistent input data for subsequent hierarchical screening and feature analysis.

[0106] The system integrates anomaly identification results, repair operation sequences, and optimization constraint rules to generate modal anomaly correction and low-quality sample optimization results information, including anomaly type, sample-level labels, repair execution strategies, and quality optimization parameters. The results information is in a structured format, containing nine items: unique sample identifier, temporal frame range, anomaly type, anomaly level, sample label, repair strategy, repair parameters, optimization constraints, and a comparison of key indicators before and after repair. This ensures complete information, consistent format, and ease of parsing and backtracking. The anomaly identification results section records the specific type, numerical value, and level of four anomalies: modal noise, missing data, spatiotemporal offset, and feature distortion. Anomaly levels are categorized into three levels: mild, moderate, and severe, based on the anomaly comprehensive score A = 0.4 × R + 0.3 × N + 0.3 × D.

[0107] The repair operation sequence section records the processing steps such as noise filtering, missing data interpolation, spatiotemporal correction, and feature optimization in the execution order, clearly defining the operation object, execution method, and processing scope of each step. The optimization constraint rules section solidifies the three quantitative constraints that must be met after repair: modal integrity rate, spatiotemporal deviation, and feature fidelity, ensuring that the processing results are compliant and effective.

[0108] For example, sample number S030, time frame 20-28, the anomaly identification result is a composite anomaly of image salt-and-pepper noise (pixel variance 36) + millimeter-wave radar short-term missing data (missing rate 15%), with a moderate anomaly level; the sample is labeled as an object to be repaired; the repair strategy is 3×3 Gaussian filtering for noise reduction + linear interpolation for completion; the repair parameters are a noise threshold of 0.3 and an interpolation step size of 1 frame; the optimization constraints are modal integrity ≥95%, time deviation ≤10ms, and feature fidelity ≥0.85; before repair, the noise interference is 0.7, the missing rate is 15%, and the feature similarity is 0.62; after repair, the noise interference is 0.2, the missing rate is 1%, and the feature similarity is 0.91. This standardized result information can be directly stored in the data management system, supporting retrieval and backtracking by sample number, anomaly type, repair strategy, etc., providing a reliable quality archive for subsequent multimodal data screening and model training, ensuring that the entire data processing process is traceable, verifiable, and reusable.

[0109] S106 performs verification processing on the multi-dimensional output of multimodal data in terms of temporal continuity, scene authenticity, parameter matching degree, and proportion of effective information. If the verification fails, multi-level conservative fallback processing such as single frame correction, fragment replacement, and batch backtracking is performed. After risk quantification, hierarchical screening, modal repair, and temporal optimization, high-quality multimodal data of intelligent driving long-tail scenarios is generated.

[0110] In one implementation, combining the requirements for comprehensive output verification and quality assurance, a three-level control mechanism of frame-level verification, segment-level replacement, and batch-level backtracking is introduced. This establishes standardized handling schemes for temporal continuity verification, scene authenticity verification, parameter matching degree comparison, effective information ratio assessment, single-frame deviation correction, abnormal segment replacement, and historical batch backtracking. Addressing the comprehensive output verification and quality assurance requirements for multimodal data in intelligent driving, this three-level control mechanism first introduces frame-level verification, segment-level replacement, and batch-level backtracking. It establishes four verification standards: temporal continuity verification, scene authenticity verification, parameter matching degree comparison, and effective information ratio assessment, as well as three fallback handling measures: single-frame deviation correction, abnormal segment replacement, and historical batch backtracking. This forms a standardized, layered, and implementable end-to-end quality control solution, ensuring compliant data output, controllable anomalies, and verifiable safeguards, preventing the amplification of local defects or overall quality failure.

[0111] Frame-level verification uses a single frame of data as the smallest unit, prioritizing the checking of two core indicators: temporal continuity and parameter matching. It identifies non-compliant frames due to temporal jumps, modal mismatches, and feature distortion. Temporal continuity is calculated using the cosine similarity of fused features from adjacent frames. The threshold S is set to ≥ 0.7. Parameter matching degree is calculated based on the consistency ratio of visual, point cloud, and radar features, with a threshold set to ≥ 0.8. If a frame has S = 0.55 and parameter matching degree = 0.62, it is judged as a non-compliant frame due to temporal jump and modal mismatch. Single-frame timestamp re-registration and feature fine-tuning correction are performed. After correction, S = 0.81 and matching degree = 0.85, meeting the qualification standard.

[0112] Segment-level replacement operates in units of 10-20 consecutive frames, verifying scene authenticity and information validity, locating consecutive abnormal segments, and matching them with a cached benchmark library to complete the replacement. Scene authenticity is evaluated by the proportion of dominant semantics within the frame, with a threshold of ≥0.85; the proportion of effective information = number of effective features / total number of features × 100%, with a threshold of ≥0.6. For example, if 15 consecutive frames have a semantic proportion of only 0.5 and an effective information proportion of 0.35, it is determined to be an abnormal segment with scene distortion. A high-quality segment with the same working conditions and scene is retrieved from the cache library for replacement. After replacement, the semantic proportion is 0.92 and the effective information proportion is 0.78, completely eliminating local quality defects.

[0113] Batch-level backtracking performs a comprehensive quality assessment on the entire batch of data, integrating a weighted score from four indicators: temporal continuity, scenario authenticity, parameter matching degree, and percentage of effective information. Set a passing threshold Q ≥ 0.7. For example, for a certain batch... =0.52、 =0.48、 =0.45、 =0.41, and the calculated Q=0.47 is lower than the threshold, indicating that the overall quality is substandard. This triggers batch backtracking, which involves re-screening from high-quality batches under the same historical conditions to ensure the reliability of the entire batch of data.

[0114] A three-tiered control mechanism provides layered handling and tiered fallback: frame-level processing of single-point anomalies, segment-level control of local distortions, and batch-level fallback for overall failures, forming a closed-loop "point-segment-batch" end-to-end system to prevent single-check failures, the spread of local defects, and overall quality collapse. For example, timing jumps are corrected at the frame level, local scene distortions are replaced by segments, and insufficient batch data quality is addressed through batch backtracking. Different anomalies are precisely matched to corresponding handling levels, ensuring high reliability, high consistency, and high fidelity of output data.

[0115] For frame-level verification, priority is given to checking temporal continuity and parameter matching. Unqualified frames with timing jumps, modal mismatches, and feature distortions are accurately identified and targeted corrections are performed to ensure that single-frame data is time-aligned, modally synchronized, and feature reliable. As the most basic unit of quality control, frame-level verification focuses on frame-by-frame feature consistency. Priority is given to verifying temporal continuity and parameter matching; only frames that pass both criteria are considered valid. Frames that fail either criterion are deemed unqualified and trigger a targeted correction process.

[0116] Temporal continuity is used to measure the stationarity of feature changes between adjacent frames, and is calculated using the cosine similarity of fused features: ,in For the current frame fusion features, The fused feature of the previous frame has a value range of 0 to 1, and a qualified threshold S≥0.7 is set. When S<0.7, it is determined that there is a temporal jump, abrupt change in inter-frame features, or abnormal temporal connection.

[0117] Parameter matching degree is used to measure the consistency of multimodal data, and to statistically measure the percentage of overlap in parameters among the three modalities: vision, LiDAR, and millimeter-wave radar. ×100% A passing threshold M ≥ 0.8 is set. When M < 0.8, modal mismatch is considered to exist, indicating excessive deviation in multimodal data parameters and failure of spatiotemporal alignment. Feature distortion is determined by feature fidelity. A passing threshold F is set to ≥ 0.85. When F < 0.85, feature distortion, image blurring, point cloud incompleteness, or radar signal distortion are identified.

[0118] The verification process is executed frame by frame. Specifically, the temporal continuity (S) and parameter matching degree (M) are calculated first, and then the feature fidelity (F) is checked. If any indicator fails to meet the standard, the frame is marked as unqualified, and the corresponding correction strategy is triggered simultaneously. For temporal jump frames, timestamp reregistration is performed; for modal mismatch, parameter deviation correction is performed; and for feature distortion, feature enhancement repair is performed. For example, if a frame has a temporal continuity (S) of 0.48, a parameter matching degree (M) of 0.55, and a feature fidelity (F) of 0.72, all three indicators fail to meet the standard, and it is judged as a composite unqualified frame of temporal jump + modal mismatch + feature distortion. Targeted corrections are performed: temporal jumps are reregistered using 10ms-level timestamps; modal mismatch is corrected for the deviation between image and point cloud parameters; and feature distortion is addressed with edge enhancement and point cloud completion. After correction, S = 0.75, M = 0.83, and F = 0.91, all three indicators meet the standard, and the single frame data is compliant and usable.

[0119] For segment-level replacement, the authenticity of the scene and the validity of the information are verified. Continuous abnormal segments are located and replaced by matching them with a cached benchmark library, thus avoiding localized quality defects. Segment-level quality verification focuses on the overall quality of continuous time segments, emphasizing the verification of scene authenticity and information validity. It accurately locates continuous abnormal segments and completes segment replacement by matching them with a cached benchmark library of historical high-quality data. This prevents the spread of localized quality defects from the source, ensuring scene continuity, information integrity, and reliable features within the time segment. Segment-level verification uses 10-20 consecutive frames as the smallest verification unit, balancing scene continuity with computational efficiency and avoiding single-frame fluctuations from interfering with overall quality judgment.

[0120] Scene authenticity verification uses the percentage of semantic consistency in the scene as the formula: ×100%, setting a passing threshold P≥0.85; values ​​below this threshold are considered to indicate sudden scene changes, semantic jumps, or scene distortion. Information validity verification is quantified by the proportion of valid features, using the following formula: ×100%, set a qualified threshold I≥0.6, if it is lower than the threshold, it is judged that the information is incomplete, the features are missing or the noise interference is severe.

[0121] The semantic consistency ratio (P) and effective information ratio (I) of each segment are calculated. Any segment failing to meet the standard is marked as an abnormal segment. After locating the start and end frames of the abnormality, consecutive high-quality segments with the same working conditions, scene, and lighting conditions are retrieved from the cached benchmark library. Frame sequence replacement is then performed to ensure temporal continuity, modal alignment, and feature consistency after replacement. For example, in a consecutive 15-frame temporal segment, only 6 frames have dominant semantic meaning. The calculated P = 6 / 15 × 100% = 40%, which is lower than 0.85; the effective feature ratio I = 35%, which is lower than 0.6, classifying it as an abnormal segment with scene distortion and incomplete information. The replacement is completed by matching 15 high-quality segments from the cached benchmark library with the same urban road, daytime lighting, and unobstructed working conditions. After replacement, the scene semantic consistency P = 92%, the effective information ratio I = 78%, and the local distortion and incomplete information issues are completely eliminated, restoring the overall quality of the temporal segment to compliance.

[0122] Performing batch-level comprehensive quality assessment is a crucial step in ensuring the end-to-end reliability of multimodal data. It involves conducting a holistic quality evaluation of the entire batch of time-series data, preventing the accumulation of local anomalies that could lead to batch-wide failure, and ensuring the overall reliability and stability of the output data. The assessment process integrates four core verification indicators: time-series continuity, scenario authenticity, parameter matching degree, and the proportion of effective information. A weighted calculation yields a batch comprehensive quality score. Based on the score threshold, the batch is determined to meet the standards. If it fails to meet the standards, historical batches are backtracked for re-screening and processing, achieving closed-loop control from local verification to overall assurance.

[0123] The formula for calculating the overall quality score of a batch is: ,in: This represents the average batch temporal continuity, ranging from 0 to 1. This represents the average value of the batch scene authenticity, ranging from 0 to 1. This represents the average batch parameter matching degree, ranging from 0 to 1. The average percentage of valid information in a batch is 0 to 1. The weights are distributed as follows: temporal continuity 0.3, scene authenticity 0.3, parameter matching degree 0.2, and valid information percentage 0.2, prioritizing the overall stability of temporal sequence and scene.

[0124] A batch quality pass threshold Q ≥ 0.7 is set. When the overall score is ≥ 0.7, the batch is judged as meeting the quality standard and can be directly output; when the score is < 0.7, the batch is judged as failing the overall quality standard and triggers the batch backtracking mechanism. The backtracking rules are as follows: select high-quality batches with the same working conditions, scenarios, and environmental interference conditions in the historical database, and re-execute the entire process of hierarchical routing screening, anomaly repair, and quality verification to ensure that the final output batch meets the quality standards.

[0125] For example, after verifying the temporal sequence of 100 frames at a city intersection frame by frame, the average temporal continuity is statistically analyzed. =0.5; Mean of scene realism =0.4; mean parameter matching degree =0.5; Mean percentage of effective information =0.4. Substituting into the formula, Q = 0.3 × 0.5 + 0.3 × 0.4 + 0.2 × 0.5 + 0.2 × 0.4 = 0.15 + 0.12 + 0.10 + 0.08 = 0.45. The overall score is 0.45 < 0.7, indicating the batch is considered substandard overall. A backtracking process is triggered, retrieving high-quality batches from the same city intersection during daytime without rain or snow interference. The screening, repair, and verification processes are re-executed, ultimately outputting a batch with an overall score of 0.82, meeting the requirement of Q ≥ 0.7, indicating the batch is compliant and usable overall.

[0126] By integrating full-dimensional verification results, single-frame correction records, fragment replacement details, and batch backtracking information, a complete processing result archive is generated. This archive includes six core components: verification indicators, non-compliance types, handling methods, key parameters, before-and-after comparisons, and final quality rating. This ensures that the entire data processing process is traceable, verifiable, and reusable. The full-dimensional verification results cover frame-by-frame averages and extreme values ​​for four indicators: temporal continuity, scene authenticity, parameter matching degree, and effective information ratio. Single-frame correction records indicate the location, correction parameters, and effects of non-compliant frames such as temporal jumps, modal mismatches, and feature distortions. Fragment replacement details record the start and end frames of abnormal fragments, replacement baseline information, and changes in indicators before and after replacement. Batch backtracking information indicates backtracking trigger conditions, historical batch numbers, and re-screening results, ensuring that every step of the processing is clearly recorded.

[0127] Risk quantification uses a normalized risk scoring formula: ,in, , , , These are the corresponding risk values ​​for the indicators, ranging from 0 to 1. R ≥ 0.7 indicates high risk, 0.4 ≤ R < 0.7 indicates medium risk, and R < 0.4 indicates low risk. The risk level is directly included in the results file. Tiered screening performs differentiated processing based on quality level, boundary type, and anomaly type. Modal repair includes noise filtering, missing data completion, and spatiotemporal correction. Temporal optimization eliminates temporal fluctuations through interpolation smoothing and temporal reregistration. The entire closed-loop process ensures data quality.

[0128] The final quality rating uses a comprehensive score formula: Q≥0.85 is excellent, 0.7≤Q<0.85 is good, 0.6≤Q<0.7 is acceptable, and Q<0.6 is unacceptable. Excellent data is output directly, while the rest need to be checked twice.

[0129] For example, a time-series sample of a long-tailed rainstorm scene showed the following full-dimensional verification results: average temporal continuity 0.55, average scene realism 0.45, average parameter matching 0.5, and average effective information ratio 0.4. The non-compliant types were: temporal jumps (frames 20-25) and scene distortion (frames 30-45). The remediation method involved single-frame timestamp re-registration and feature fine-tuning for frames 20-25, and replacement of high-quality rainstorm scene segments from the cache for frames 30-45. Before repair, the risk score R=0.73 (high risk); after repair, R=0.32 (low risk), with a final comprehensive score Q=0.88 and a quality rating of excellent. The final output data is modally complete, spatiotemporally aligned, and feature-faithful, and can be directly used for intelligent driving model training, scene generalization, and algorithm iteration optimization.

[0130] In one implementation, such as Figure 2 As shown, this application also provides a system for filtering and processing multimodal data in long-tail scenarios of intelligent driving, including: The distributed ultrasound image preprocessing module 201 is used to perform size unification, type standardization and pixel normalization processing on multi-center distributed cardiac ultrasound images, complete single frame image block division, Gaussian filtering noise reduction, 3×3 Laplacian convolution texture scoring and adaptive mask generation, and output normalized temporal frame sequence, mask matrix and position index information. The privacy-compliant federated pre-training module 202 is used to build a federated learning architecture. It reconstructs masked image patches based on MAE encoding and decoding, updates and clips local gradients according to MSE loss, and removes offline clients from the central server. It iteratively updates the global model by aggregating local parameters according to the amount of data, so as to achieve privacy-compliant feature extraction that is available but not visible to the data. The sparse temporal feature encoding module 203 is used to perform sparse sampling on the temporal frames of cardiac ultrasound. It uses pre-trained ViT to extract single-frame spatial features and generates sine-cosine sparse temporal codes according to the acquisition index. After feature fusion, the temporal association is enhanced by Transformer multi-head self-attention and the frame-level temporal feature sequence is output. The dual-branch spatiotemporal feature fusion module 204 is used to construct a dual-branch spatiotemporal fusion network. It extracts global spatiotemporal features by combining 3D residual convolution with gated attention, and generates fine temporal features by fusing 2D residual convolution with temporal encoding. The two types of features are then concatenated to obtain high-dimensional spatiotemporal fusion feature information. The temporal aggregation and classification prediction module 205 is used to perform attention-weighted temporal aggregation on spatiotemporal fusion features and compress them into video-level features; the input fully connected classification head is activated by softmax and the output is multi-class probability prediction results including cardiac amyloidosis. The robust enhancement and auxiliary screening module 206 is used to adapt linear and convex array scanning modes through bidirectional polar coordinate transformation, simulate motion blur, Gaussian blur, and salt-and-pepper noise to carry out joint data enhancement; it integrates adaptive masking, sparse temporal coding and bi-branch spatiotemporal fusion to accurately capture the spatiotemporal features of cardiac ultrasound and generate intelligent auxiliary screening results for cardiac amyloidosis.

[0131] The various embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for evaluating the screening and processing method, electronic device, electronic device, and readable storage medium for multimodal data in long-tail scenarios of intelligent driving are basically similar to the above-described embodiments for the screening and processing method for multimodal data in long-tail scenarios of intelligent driving, and are therefore described simply. Relevant parts can be referred to in the descriptions of the above-described embodiments for the screening and processing method for multimodal data in long-tail scenarios of intelligent driving.

Claims

1. A method for filtering and processing multimodal data in long-tail scenarios of intelligent driving, characterized in that, include: Acquire the original features of multimodal data, construct a long-tailed scene multimodal data codebook that integrates scene semantic features, target contour parameters, spatiotemporal location labels, working condition information, and environmental interference parameters, and perform feature similarity matching retrieval to generate a candidate set of time-series data; Based on six core indicators—scene rarity, modal noise, data missing anomaly, spatiotemporal misalignment, target occlusion interference, and working condition matching error—the risk and quality normalization calculation of sample-level data is completed, and a four-level classification of invalid, inefficient, effective, and high-quality long-tail scene data is achieved. By combining temporal frame alignment, sensor frame rate calibration, motion posture change detection, and scene switching recognition, the perception of long-tail scene boundaries and working condition switching is realized, and the effective screening range of multimodal data and the stable constraints of non-boundary data with temporal continuity, modal matching, and scene fidelity are determined. Based on data quality level and scenario boundary information, a hierarchical routing and filtering decision rule is established, which includes direct retention, mild repair, modality completion, sample removal, and batch filtering. A data filtering benchmark library is built based on the historical high-quality data caching module. For modal anomaly data, noise filtering and removal, missing data interpolation and completion, and spatiotemporal offset correction are performed. For low-quality long-tailed samples, feature weight fine-tuning, redundant information removal, interference parameter suppression, and multimodal consistency forced alignment are performed. The multi-modal data output is verified in all dimensions, including temporal continuity, scene authenticity, parameter matching degree, and proportion of effective information. If the verification fails, multi-level conservative fallback processing, such as single-frame correction, fragment replacement, and batch backtracking, is performed. After risk quantification, hierarchical screening, modal repair, and temporal optimization, high-quality multi-modal data for long-tail scenarios of intelligent driving is generated.

2. The method as described in claim 1, characterized in that, By combining temporal frame alignment, sensor frame rate calibration, motion posture change detection, and scene transition recognition, long-tail scene boundary and condition transition perception are achieved. Effective filtering intervals for multimodal data and stable constraints for non-boundary data that ensure temporal continuity, modal matching, and scene fidelity are determined, including: Combining the boundary perception requirements of long-tail scenarios with the logic of working condition switching recognition, we introduce time alignment rules, frame rate calibration mechanisms, motion change detection and scene switching judgment strategies to determine a standardized perception scheme for time frame alignment, sensor frame rate calibration, motion posture detection and scene switching recognition. Perform timestamp synchronization and inter-frame registration on a frame-by-frame basis for multimodal time-series data to complete time-series frame alignment and eliminate modal time-series misalignment; Dynamic interpolation and resampling calibration are performed based on the actual frame rate differences of each sensor to unify the timing benchmark of multimodal data; The change in motion posture is calculated based on the feature differences of consecutive frames to identify abrupt changes in vehicle dynamic conditions. By integrating scene semantic features and inter-frame similarity determination, the switching boundaries of long-tail scenes can be accurately located; By integrating alignment results, calibration parameters, motion change information, and scene boundary signals, we determine the effective screening interval for multimodal data and the stability constraints for non-boundary data that are time-continuous, modally matched, and scene-fidelity preserved.

3. The method as described in claim 1, characterized in that, Based on data quality levels and scenario boundary information, a tiered routing and filtering decision rule is established, comprising direct retention, minor repair, modality completion, sample removal, and batch filtering. A data filtering benchmark library is constructed based on a historical high-quality data caching module, including: A hierarchical routing filtering decision framework is constructed based on data quality level and scenario boundary information, and a standardized filtering process is established for data classification judgment, boundary constraint verification, routing rule mapping and cache benchmark support. Multi-dimensional compliance verification is performed on sample quality level, scene boundary type, temporal continuity, and modal matching degree. Temporal features of modal integrity, feature credibility, and working condition adaptability are analyzed frame by frame. Abnormal data, including those with temporal misalignment, modal missingness, and scene distortion, are identified by using sparse weighting, noise interference level, and spatiotemporal deviation value. A benchmark library is built based on a historical high-quality data caching module. The hierarchical routing and screening results are integrated with quality level, boundary information, anomaly type, cache matching degree, direct retention, mild repair, modality completion, sample removal, and batch filtering.

4. The method as described in claim 1, characterized in that, For modal anomaly data, noise filtering, missing data interpolation completion, and spatiotemporal offset correction are performed. For low-quality long-tailed samples, feature weight fine-tuning, redundant information removal, interference parameter suppression, and multimodal consistency forced alignment are performed, including: A multimodal anomaly repair fusion mechanism is constructed based on modal anomaly identification, missing data localization, and spatiotemporal deviation calculation. Noise detection results are given priority processing, while missing data localization and spatiotemporal deviation are used for collaborative correction of anomaly composite scenarios. Multi-dimensional anomaly labels are generated for each sample. Modal noise frames, missing data segments, spatiotemporal offset frames, and feature distortion samples are marked as objects to be repaired, while the remaining samples are judged as low-quality optimization objects. Establish linkage rules between anomaly types and repair strategies, perform targeted corrections on abnormal data, implement feature optimization for low-quality samples, and complete the division of repair permissions and solidification of optimization constraints; The effectiveness of modal integrity, spatiotemporal consistency and feature fidelity after repair is verified. The correction results are mapped to the original time series frames. The number of extended target frames before and after the abnormal segment forms a repair transition zone to ensure that the repair process is accurately matched with the multimodal time series. By integrating anomaly identification results, repair operation sequences, and optimization constraint rules, modal anomaly correction and low-quality sample optimization result information is generated, which includes anomaly type, sample-level label, repair execution strategy, and quality optimization parameters.

5. The method as described in claim 4, characterized in that, The multimodal data output is validated across all dimensions, including temporal continuity, scene realism, parameter matching, and the proportion of effective information. If the validation fails, a multi-level conservative fallback process involving single-frame correction, segment replacement, and batch backtracking is implemented. Through risk quantification, hierarchical screening, modal repair, and temporal optimization, high-quality multimodal data for long-tail scenarios in intelligent driving is generated, including: In line with the requirements of full-dimensional output verification and quality assurance, a three-level control mechanism of frame-level verification, segment-level replacement and batch-level backtracking is introduced to determine standardized handling solutions for temporal continuity verification, scene authenticity verification, parameter matching degree comparison, effective information ratio evaluation and single frame deviation correction, abnormal segment replacement and historical batch backtracking. For frame-level verification, priority is given to checking the temporal continuity and parameter matching, identifying unqualified frames with temporal jumps, modal mismatches, and feature distortions, and performing targeted corrections. For fragment-level replacement, verify the authenticity of the scene and the validity of the information, locate continuous abnormal fragments and match them with the cached benchmark library to complete the replacement, thus avoiding local quality defects; For batch-level backtracking, the overall quality is determined by comprehensive verification results. If the quality is not up to standard, it is backtracked to historical high-quality batches for re-screening and processing to ensure the overall data reliability. By integrating verification results, correction operations, replacement strategies, and backtracking information, a full-dimensional verification and multi-level fallback processing result is generated, which includes verification indicators, non-compliance types, handling methods, and final quality rating. After risk quantification, hierarchical screening, modal repair, and time series optimization, high-reliability, high-consistency, and high-fidelity multimodal high-quality data for long-tail scenarios of intelligent driving is output.

6. A system for filtering and processing multimodal data in long-tail scenarios of intelligent driving, characterized in that, The system includes: The distributed ultrasound image preprocessing module is used to perform size unification, type standardization and pixel normalization processing on multi-center distributed cardiac ultrasound images. It completes single-frame image segmentation, Gaussian filtering noise reduction, 3×3 Laplacian convolution texture scoring and adaptive mask generation, and outputs normalized temporal frame sequence, mask matrix and position index information. The privacy-compliant federated pre-training module is used to build a federated learning architecture. It reconstructs masked image patches based on MAE encoding and decoding, updates and clips local gradients according to MSE loss, and removes offline clients from the central server. It iteratively updates the global model by aggregating local parameters according to the amount of data, thereby achieving privacy-compliant feature extraction where the data is available but not visible. The sparse temporal feature encoding module is used to perform sparse sampling on cardiac ultrasound time frames. It uses pre-trained ViT to extract single-frame spatial features and generates sine-cosine sparse temporal codes according to the acquisition index. After feature fusion, the temporal association is enhanced by Transformer multi-head self-attention and the frame-level temporal feature sequence is output. The dual-branch spatiotemporal feature fusion module is used to construct a dual-branch spatiotemporal fusion network. It extracts global spatiotemporal features by combining 3D residual convolution with gated attention, and generates fine temporal features by fusing 2D residual convolution with temporal encoding. The two types of features are then concatenated to obtain high-dimensional spatiotemporal fusion feature information. The temporal aggregation and classification prediction module is used to perform attention-weighted temporal aggregation on spatiotemporal fusion features and compress them into video-level features; the input fully connected classification head is activated by softmax and the output is multi-class probability prediction results including cardiac amyloidosis. The robust enhancement and auxiliary screening module is used to adapt linear and convex array scanning modes through bidirectional polar coordinate transformation, simulate motion blur, Gaussian blur, and salt-and-pepper noise to carry out joint data enhancement; it integrates adaptive masking, sparse temporal coding and bi-branch spatiotemporal fusion to accurately capture the spatiotemporal features of cardiac ultrasound and generate intelligent auxiliary screening results for cardiac amyloidosis.

7. An electronic device, characterized in that, include: First processor; The processor also includes a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the filtering and processing method for multimodal data in long-tail scenarios of intelligent driving as described in any one of claims 1 to 5 by executing the executable instructions.