Vehicle type recognition method and system based on multi-modal perception and dynamic feature matching
Patent Information
- Application Number
- CN202510949746.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-07-10
AI Technical Summary
[0003]然而,现行匹配框架仍把各通道视为静态可交换信息,只在单帧层面施加阈值一致性检验,缺乏基于时空连续性的约束与误差反馈
[0036] This invention constructs a multi-source co-current tensor by adaptively time-scaled synchronous fusion of visual, millimeter-wave, and wheel speed data, providing solid support for a unified feature foundation. It achieves real-time tracking of cross-domain features using a pseudo-topological tracking grid and elastic anchor points. The flicker hysteresis and velocity feedback amplitude parameters sensitively capture drift signals from visual distortion and dynamic abrupt changes, and a drift truncation index is generated in real-time via a time-error truncation operator to adaptively correct heterogeneous features and accurately identify complex dynamic disturbances. High-risk drift entries are retrieved from the matching memory and a correction matrix is constructed to perform targeted backfilling of the co-current tensor, achieving multi-dimensional alignment of visual, echo, and dynamic features. Finally, the corrected features are fused with the dynamic pattern mapping distance to efficiently output vehicle model tags and push them to billing and safety nodes, ensuring high accuracy and robustness of recognition under varying lighting conditions, occlusion, and bumpy environments, significantly reducing misclassification rates and improving reliability.
Smart Images

Figure CN120832530B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle model recognition technology, and more specifically, to a vehicle model recognition method and system based on multimodal perception and dynamic feature matching. Background Technology
[0002] In the rain-and-fog-filled nighttime traffic scenarios on elevated ramps, visual, millimeter-wave, and wheel speed signals simultaneously face diffuse reflection from light spots, multipath echoes from guardrails, and transient turbulence disturbances. Although multi-source observations can capture the outline, volume echo, and dynamic response of the same vehicle on a microsecond scale, these features frequently change with pitch, roll, and obstruction. The originally synchronized observation trajectories gradually branch off in the feature space, quietly forming "drift cracks."
[0003] However, the current matching framework still treats each channel as static, exchangeable information, applying threshold consistency checks only at the single-frame level, lacking constraints and error feedback based on spatiotemporal continuity. When the gap expands to a cross-frame scale, the algorithm merges irrelevant brightness peaks, radar ghost images, and torque drops into the same entity, causing the vehicle model label to jump to the wrong category during lane changes or sudden braking, triggering a chain reaction of billing disputes, accident retrospective misjudgments, and delayed triggering of active protection.
[0004] To address the aforementioned problems, a technical solution is provided. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a vehicle model recognition method and system based on multimodal perception and dynamic feature matching. This method constructs a multi-source co-current tensor by adaptively time-scaled synchronous fusion of visual, millimeter-wave, and wheel speed data. Real-time tracking of cross-domain features is achieved using a pseudo-topological tracking grid and elastic anchor points. The flicker hysteresis and velocity feedback amplitude parameters sensitively capture drift signals from visual distortion and dynamic abrupt changes. A drift truncation index is generated in real-time using a time-error truncation operator to adaptively correct heterogeneous features and identify complex dynamic disturbances. High-risk drift entries are retrieved from the matching memory and a correction matrix is constructed to perform directional backfilling of the co-current tensor, achieving multi-dimensional alignment of visual, echo, and dynamic features. Finally, the corrected features and dynamic mode mapping distance are fused to efficiently output vehicle model tags and push them to billing and security nodes, thus solving the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A vehicle model recognition method based on multimodal perception and dynamic feature matching includes the following steps:
[0008] S1: Preprocess multi-source signals and construct multi-source co-current tensors using adaptive time-scale alignment techniques;
[0009] S2: Extract optical flow peaks and echo peaks to construct a pseudo-topological tracking mesh, dynamically deploy elastic anchor points and refresh the index in real time;
[0010] S3: The flash hysteresis is obtained by detecting the peak displacement of the lamp group and cross-correlating it with the predicted trajectory of the elastic anchor point. The speed feedback return amplitude generated by the third derivative of the wheel speed curve and the integral of the braking interval is combined with the two for comprehensive analysis and written into the matching memory.
[0011] S4: Retrieve high-risk entries from the matching memory, generate a drift correction matrix, and backfill the multi-source co-current tensor;
[0012] S5: Perform feature extraction and matching analysis on the corrected multi-source co-current tensor, comprehensively evaluate the fit between the features and the vehicle model standard, determine the vehicle model category label, and push it to the billing and security nodes.
[0013] In a preferred embodiment, step S1 includes the following:
[0014] Visual frames, millimeter-wave echoes, and wheel speed data are acquired and preprocessed. Using the sampling frequency of the millimeter-wave echoes as the reference time axis, visual frames and wheel speed data are aligned by interpolation, features are extracted, and they are concatenated into a comprehensive feature vector to construct a multi-source co-current tensor.
[0015] In a preferred embodiment, step S2 includes the following:
[0016] Visual feature vectors and millimeter-wave echo feature vectors are extracted from multi-source co-current tensors. Motion vectors are generated by calculating the differences between visual feature vectors at adjacent time points. The point with the largest motion vector amplitude is identified as the optical flow peak. The point with the largest signal intensity is identified from the millimeter-wave echo feature vector as the echo peak. An initial mesh is constructed using the optical flow peak and the echo peak as nodes. The topological relationship is determined by calculating the spatial distance between nodes and the consistency of motion direction. The node positions are updated over time to adapt to vehicle motion, forming a pseudo-topological tracking mesh.
[0017] In a preferred embodiment, step S2 further includes the following:
[0018] In the pseudo-topological tracking grid, an elastic anchor point is selected, with the initial position being the weighted center of the optical flow peak and the echo peak. The weights are dynamically adjusted according to the signal reliability. The anchor point position is dynamically updated and the index is refreshed by comparing the deviation between the current position and the predicted position of the elastic anchor point, thus realizing the spatiotemporal continuous tracking of vehicle features.
[0019] In a preferred embodiment, step S3 includes the following:
[0020] Visual feature vectors are extracted from the multi-source synchronous tensor, and the displacement of the flash peak of the light group is detected. The flash hysteresis is calculated by cross-correlation analysis with the predicted trajectory of the anchor point. Wheel speed feature vectors are extracted from the multi-source synchronous tensor, the third derivative of the wheel speed curve is calculated, and the third derivative is integrated in the braking interval to generate the speed feedback return amplitude. The flash hysteresis and speed feedback return amplitude are input into the time error truncation operator to calculate the drift truncation index, and the drift truncation index is written into the matching memory.
[0021] In a preferred embodiment, step S4 includes the following:
[0022] High-risk entries are retrieved from the matching memory, and the time points when the drift cutoff index exceeds a preset threshold are identified. High-risk entries are sorted by drift severity according to the drift cutoff index and its changing trend. A drift correction matrix is generated using the drift cutoff index and flicker hysteresis. The drift correction matrix contains correction coefficients that reflect the intensity of interference.
[0023] In a preferred embodiment, step S4 further includes the following:
[0024] The drift correction matrix is applied to the corresponding time point feature vectors of the multi-source co-current tensor to correct the feature vectors and realign the multi-source features on the time axis.
[0025] In a preferred embodiment, step S5 includes the following:
[0026] The system receives the corrected multi-source synchronous tensor, extracts features from it using a convolutional neural network to generate a feature map, calculates the Euclidean distance between the feature map and a predefined vehicle model template to generate a residual value, extracts wheel speed-related features from the corrected multi-source synchronous tensor to generate a dynamic feature vector, and uses a dynamic time warping algorithm to calculate the distance between the dynamic feature vector and the standard power mode of each vehicle model.
[0027] In a preferred embodiment, step S5 further includes the following:
[0028] The reciprocal of the residual value and the reciprocal of the distance value are combined and weighted to generate a comprehensive discrimination score. The vehicle category with the highest comprehensive discrimination score is selected as the vehicle label and pushed to the billing and safety node.
[0029] A vehicle model recognition system based on multimodal perception and dynamic feature matching includes:
[0030] Signal fusion module: preprocesses multi-source signals and constructs multi-source co-current tensors using adaptive time-scale alignment technology;
[0031] Feature tracking module: Extracts optical flow peaks and echo peaks to construct a pseudo-topological tracking mesh, dynamically deploys elastic anchor points and refreshes the index in real time;
[0032] Drift detection module: The flicker hysteresis is obtained by detecting the displacement of the flash peak of the lamp group and cross-correlating it with the predicted trajectory of the elastic anchor point. The speed feedback return amplitude generated by the third derivative of the wheel speed curve and the braking interval integral is combined with the two to perform a comprehensive analysis and write it into the matching memory.
[0033] Drift correction module: retrieves high-risk entries from the matching memory, generates a drift correction matrix, and backfills the multi-source co-current tensor;
[0034] Vehicle model recognition module: Performs feature extraction and matching analysis on the corrected multi-source simultaneous tensor, comprehensively evaluates the fit between the features and the vehicle model standard, determines the vehicle model category label, and pushes it to the billing and security nodes.
[0035] The technical effects and advantages of the vehicle model recognition method and system based on multimodal perception and dynamic feature matching in this invention are as follows:
[0036] This invention constructs a multi-source co-current tensor by adaptively time-scaled synchronous fusion of visual, millimeter-wave, and wheel speed data, providing solid support for a unified feature foundation. It achieves real-time tracking of cross-domain features using a pseudo-topological tracking grid and elastic anchor points. The flicker hysteresis and velocity feedback amplitude parameters sensitively capture drift signals from visual distortion and dynamic abrupt changes, and a drift truncation index is generated in real-time via a time-error truncation operator to adaptively correct heterogeneous features and accurately identify complex dynamic disturbances. High-risk drift entries are retrieved from the matching memory and a correction matrix is constructed to perform targeted backfilling of the co-current tensor, achieving multi-dimensional alignment of visual, echo, and dynamic features. Finally, the corrected features are fused with the dynamic pattern mapping distance to efficiently output vehicle model tags and push them to billing and safety nodes, ensuring high accuracy and robustness of recognition under varying lighting conditions, occlusion, and bumpy environments, significantly reducing misclassification rates and improving reliability. Attached Figure Description
[0037] Figure 1 This is a flowchart illustrating the vehicle model recognition method based on multimodal perception and dynamic feature matching of the present invention.
[0038] Figure 2 This is a schematic diagram of the vehicle model recognition system based on multimodal perception and dynamic feature matching of the present invention. Detailed Implementation
[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0040] Example 1: Figure 1This invention presents a vehicle model recognition method based on multimodal perception and dynamic feature matching, comprising:
[0041] S1: Preprocess multi-source signals and construct multi-source co-current tensors using adaptive time-scale alignment techniques;
[0042] S2: Extract optical flow peaks and echo peaks to construct a pseudo-topological tracking mesh, dynamically deploy elastic anchor points and refresh the index in real time;
[0043] S3: The flash hysteresis is obtained by detecting the peak displacement of the lamp group and cross-correlating it with the predicted trajectory of the elastic anchor point. The speed feedback return amplitude generated by the third derivative of the wheel speed curve and the integral of the braking interval is combined with the two for comprehensive analysis and written into the matching memory.
[0044] S4: Retrieve high-risk entries from the matching memory, generate a drift correction matrix, and backfill the multi-source co-current tensor;
[0045] S5: Perform feature extraction and matching analysis on the corrected multi-source simultaneous tensor, comprehensively evaluate the fit between the features and the vehicle model standard, determine the vehicle model category label, and push it to the billing and security nodes.
[0046] In the complex nighttime traffic scenarios of elevated ramps shrouded in rain and fog, vehicle model recognition faces multiple challenges. Visual signals are affected by diffuse reflection from light spots, millimeter-wave signals are interfered with by multipath echoes from guardrails, and wheel speed signals fluctuate due to transient bumps and disturbances. These factors make it difficult for single-sensor observation data to accurately characterize vehicle features, especially when the vehicle is pitching, tilting, or obstructed, as the feature trajectories of multi-source observations gradually bifurcate in the feature space, forming "drift cracks." Traditional vehicle model recognition methods rely on single-frame threshold consistency checks, neglecting spatiotemporal continuity constraints and error feedback, which easily leads to incorrect matching across frame scales, affecting toll accuracy, safety policy triggering, and the reliability of accident backtracking. To address this issue, this invention proposes a vehicle model recognition method based on multimodal perception and dynamic feature matching. By fusing visual frames, millimeter-wave echoes, and wheel speed curves, a multi-source co-current tensor is constructed, and combined with dynamic feature matching technology, temporal alignment and accurate recognition of vehicle features are achieved. Step S1, as the entry fusion step, is the foundation of the entire solution. It aims to integrate multi-source data into a unified data structure, providing reliable input for subsequent pseudo-topology tracking, drift correction, and vehicle type identification.
[0047] The purpose of step S1 is to fuse the visual frame, millimeter-wave echo, and wheel speed curve into a multi-source co-current tensor. Adaptive timescales ensure the alignment of the multi-source data on the time axis, providing a consistent data foundation for subsequent steps. The specific processing logic is as follows, divided into three sub-steps: data acquisition and preprocessing, adaptive timescale generation, and multi-source co-current tensor construction.
[0048] Step S1 includes the following:
[0049] S1.1, Data Acquisition and Preprocessing
[0050] Data acquisition and preprocessing are the starting steps for constructing multi-source synchronous tensors. For nighttime scenes with rain and fog on elevated ramps, it is necessary to process the characteristics of visual frames, millimeter-wave echoes, and wheel speed curves separately to reduce environmental interference and extract effective information.
[0051] Visual Frame Acquisition and Preprocessing: High-definition cameras deployed on elevated ramps capture visual images of the vehicle, generating a raw visual frame sequence. Due to the effects of rain, fog, and diffuse reflection from nighttime light spots, the raw visual frame sequence contains noise and blur, weakening the clarity of the vehicle outline. The preprocessing process first utilizes wavelet transform technology to decompose the image into different frequency components, removing high-frequency noise while preserving low-frequency features of the vehicle outline, generating denoised visual frames. Next, contrast stretching technology is applied to enhance the salience of the vehicle's headlights and shape by adjusting the brightness range of image pixels, generating enhanced visual frames. These enhanced visual frames more clearly present the vehicle's visual features, facilitating subsequent extraction of outline and headlight information.
[0052] Millimeter-wave echo acquisition and preprocessing: Echo signals from vehicles are acquired using millimeter-wave radar to generate raw echo sequences. Due to the multipath effect of guardrails, virtual image interference is mixed in the raw echo sequences, affecting the accuracy of vehicle volume characteristics. The preprocessing process first uses a bandpass filter, setting a specific frequency range to filter out low-frequency background noise and high-frequency random interference, generating filtered echoes. Next, envelope detection is performed on the filtered echoes. By identifying the trend of signal amplitude changes, the echo peak corresponding to the vehicle volume characteristics is extracted to generate preprocessed echoes. Preprocessed echoes can more accurately reflect the vehicle's volume and distance information, reducing the interference of virtual images on subsequent analysis.
[0053] Wheel speed curve acquisition and preprocessing: Wheel speed data is acquired through vehicle wheel speed sensors to generate a raw wheel speed sequence. Due to the influence of transient bumps, the raw wheel speed sequence contains spikes, affecting the accuracy of speed and acceleration calculations. The preprocessing process first uses a moving average filtering technique to smooth out spikes by calculating the average value of wheel speed data within a continuous time window, generating smoothed wheel speeds. Next, the smoothed wheel speeds are normalized to map the wheel speed data of different vehicles to a uniform numerical range, eliminating differences in speed dimensions, and generating preprocessed wheel speeds. Preprocessed wheel speeds can stably reflect the dynamic characteristics of the vehicle, providing reliable input for subsequent feature extraction.
[0054] After the above processing, enhanced visual frames, preprocessed echoes, and preprocessed wheel speeds are obtained, which serve as intermediate data for subsequent steps.
[0055] S1.2, Adaptive timescale generation:
[0056] Adaptive time stamp generation is used to solve the problem of inconsistent sampling frequencies and timestamps for visual frames, millimeter-wave echoes, and wheel speed curves. It aligns the three types of signals to a unified reference time axis to ensure time synchronization.
[0057] First, a reference time axis is determined, using the sampling frequency of the millimeter-wave echo as the reference frequency, as it is typically high. The reciprocal of the reference frequency is defined as the time interval, generating a series of equally spaced time points to form the reference time axis. Next, time alignment processing is performed on three types of signals: For enhanced visual frames, since their sampling frequency is lower than the reference frequency of the reference time axis, cubic spline interpolation is used to estimate the value at each reference time point based on existing data points to generate aligned visual frames; for preprocessed echoes, since their sampling frequency is consistent with the reference time axis, the original values are directly used to generate aligned echoes; for preprocessed wheel speeds, since their sampling frequency may differ from the reference time axis, linear interpolation is used to estimate the value at each reference time point based on adjacent data points to generate aligned wheel speeds. After alignment processing, aligned visual frames, aligned echoes, and aligned wheel speeds are obtained on a unified reference time axis, providing time-consistent data for subsequent fusion.
[0058] S1.3, Construction of multi-source synchronous tensors:
[0059] Multi-source synchronous tensor construction is the process of fusing aligned visual frames, aligned echoes, and aligned wheel speeds into a high-dimensional tensor to comprehensively represent vehicle features.
[0060] First, feature extraction is performed: vehicle contour and light cluster features are extracted from aligned visual frames, and edge and brightness distribution are identified using image processing techniques to form visual feature vectors; echo peak and distance features are extracted from aligned echoes, and echo feature vectors are formed by analyzing signal amplitude and time delay; speed and acceleration features are extracted from aligned wheel speeds, and wheel speed feature vectors are formed by calculating the rate of change within a time window. These feature vectors represent the key information of each signal at each time point. Next, tensor concatenation is performed. At each reference time point, the visual feature vector, echo feature vector, and wheel speed feature vector are sequentially combined into a comprehensive feature vector. For the comprehensive feature vector sequence across the entire reference time axis, a multi-source co-current tensor is constructed, with a matrix structure representing the number of time points and the dimension of the comprehensive features. The final output multi-source co-current tensor contains fused features from visual, millimeter-wave, and wheel speed signals, providing rich temporal information for subsequent analysis.
[0061] Through the above technical logic, step S1 provides a stable and accurate multi-source concurrent tensor for the entire vehicle model recognition system, ensuring efficient vehicle feature extraction and temporal alignment in nighttime scenes with rain and fog on elevated ramps.
[0062] Step S1 has generated a multi-source simultaneous tensor by fusing visual frames, millimeter-wave echoes, and wheel speed curves, laying the data foundation for subsequent processing. However, single-frame threshold verification cannot effectively constrain cross-frame feature drift, necessitating the introduction of a spatiotemporal continuity mechanism to ensure the stability and accuracy of feature tracking. Step S2 addresses this context by constructing a feature tracking framework to accommodate vehicle dynamics and environmental interference, providing reliable support for subsequent drift correction and vehicle model identification.
[0063] The goal of step S2 is to construct a pseudo-topological tracking mesh based on optical flow peaks and echo peaks, dynamically deploy elastic anchor points, and refresh the index in real time to achieve continuous spatiotemporal tracking of vehicle features. The following describes in detail the specific processing logic of step S2, which is divided into three sub-steps: extraction of optical flow peaks and echo peaks, construction of the pseudo-topological tracking mesh, and deployment of elastic anchor points and index refresh.
[0064] Step S2 includes the following:
[0065] S2.1, Extraction of optical flow peak and echo peak:
[0066] Optical flow peak and echo peak extraction are key feature points of visual and millimeter-wave signals obtained from multi-source co-current tensors, which are used to construct pseudo-topological tracking meshes and track vehicle features.
[0067] Visual feature vectors are extracted from the multi-source co-current tensor, representing the vehicle's outline and headlight information in the visual frame. By comparing the visual feature vectors at adjacent time points, the difference between them is calculated to generate a motion vector, reflecting the vehicle's displacement in the visual frame. The point with the largest motion vector amplitude is defined as the optical flow peak, indicating the location of the most significant vehicle motion in the visual frame. Next, millimeter-wave echo feature vectors are extracted from the multi-source co-current tensor, containing vehicle volume and distance information. By analyzing the intensity of the echo signal, the point with the highest intensity is identified, defined as the echo peak, representing the significant reflection location of the vehicle in the millimeter-wave radar. The computational logic for extracting optical flow peaks and echo peaks lies in using visual signals to capture vehicle motion features and combining them with distance and volume features provided by millimeter-wave signals to form a multi-source complementary keypoint set.
[0068] The rationale for extracting optical flow peaks and echo peaks is that visual signals and millimeter-wave signals respectively provide vehicle motion information and physical characteristics. These two signals are complementary in rain, fog, and nighttime scenes, comprehensively characterizing the vehicle's dynamic changes and spatial position. Using the point of maximum motion vector amplitude as the optical flow peak and the point of maximum echo signal intensity as the echo peak allows for precise localization of key vehicle features. This method reduces redundant data in the multi-source simultaneous tensor by extracting key feature points, improving computational efficiency. Simultaneously, it provides accurate input for constructing the pseudo-topological tracking mesh, ensuring the reliability and accuracy of subsequent tracking.
[0069] S2.2, Pseudo-topological tracking mesh construction:
[0070] The pseudo-topology tracking mesh is constructed based on optical flow peaks and echo peaks to generate a dynamically adjusted mesh structure to adapt to changes in vehicle motion and attitude.
[0071] An initial mesh is constructed in the feature space using optical flow peaks and echo peaks as nodes. The topological relationships between nodes are determined by calculating the spatial distance and motion direction consistency between the optical flow peaks and echo peaks. Specifically, the spatial distance is obtained by measuring the straight-line distance between two points in the feature space, used to assess the spatial proximity between nodes; the motion direction consistency is obtained by comparing the differences in the motion vector directions of the optical flow peaks and echo peaks, used to confirm the correlation of nodes in motion. The initial mesh is thus formed, reflecting the spatial distribution and motion relationship of the optical flow peaks and echo peaks. Next, as time progresses, the vehicle's motion causes changes in the positions of the optical flow peaks and echo peaks. By calculating the displacements of the optical flow peaks and echo peaks at adjacent time points, the positions of the nodes in the mesh are updated, allowing the mesh structure to be dynamically adjusted, forming a pseudo-topological tracking mesh. The adjustment process ensures that the mesh always remains consistent with the actual distribution of vehicle features.
[0072] Because the positions of feature points constantly change during vehicle movement, fixed meshes cannot adapt to this dynamic nature. However, by constructing an initial mesh with consistent spatial distance and direction of motion, and dynamically adjusting the mesh structure based on displacement, it is possible to effectively track the spatiotemporal changes of vehicle features. The pseudo-topological tracking mesh maintains its match with the actual vehicle trajectory through dynamic adjustment, providing a stable feature tracking framework. This enhances the system's adaptability to changes in vehicle attitude in complex scenes, thereby improving the robustness and accuracy of tracking.
[0073] S2.3, Elastic Anchor Deployment and Index Refresh:
[0074] Elastic anchor point deployment and index refresh select and dynamically update key nodes in the pseudo-topology tracking grid to achieve stable tracking of vehicle features.
[0075] In the pseudo-topological tracking mesh, key nodes representing the center of vehicle features are selected and defined as elastic anchor points. The initial position of the elastic anchor point is determined by calculating the weighted center of the optical flow peak and echo peak. The calculation of the weighted center comprehensively considers the reliability weights of visual and millimeter-wave signals, and the weights are dynamically adjusted according to signal strength and noise level to ensure accurate positioning. After deployment, the elastic anchor point adaptively adjusts its position as the vehicle features change. Then, over time, the position of the elastic anchor point is updated by comparing the deviation between the current position and the expected position predicted based on motion, and its index in the mesh is refreshed. The update process adjusts the anchor point coordinates based on the magnitude and direction of the deviation to ensure that it is always anchored to the center of the vehicle features. The index refresh remaps the relationship between the anchor point and the mesh nodes, maintaining the continuity of tracking.
[0076] Because the vehicle feature center changes with movement, fixed anchor points cannot continuously track the target. However, by calculating the initial position using weighted centers and dynamically adjusting the anchor points, the system can flexibly adapt to feature changes. Index refresh maintains the alignment of the anchor points with the actual feature centers through deviation correction. Elastic anchor points provide stable reference points for vehicle features, and combined with the index refresh mechanism, this ensures the continuity of tracking in time and space, significantly improving the accuracy and stability of feature matching in dynamic scenes.
[0077] Step S2 begins with the extraction of optical flow peaks and echo peaks. It generates key feature points from multi-source signals, providing input for the construction of a pseudo-topological tracking mesh. This mesh utilizes these feature points to form a dynamic grid that adapts to changes in vehicle motion. Elastic anchor point deployment and index refresh further optimize key point tracking within the mesh, ensuring spatiotemporal continuity. Together, these techniques achieve stable tracking of vehicle features, making it suitable for nighttime scenarios involving rain and fog on elevated ramps, providing efficient feature extraction and temporal alignment capabilities for vehicle model recognition systems.
[0078] Step S2 constructs a pseudo-topological tracking mesh based on optical flow peaks and echo peaks, and deploys elastic anchor points to achieve spatiotemporal continuity tracking of features. However, vehicle dynamics (such as acceleration, braking, or lane changes) and environmental disturbances (such as rain, fog, light spots, or guardrail echoes) cause deviations between the predicted trajectory of the elastic anchor points and the actual feature trajectory, resulting in feature drift. To address this issue, step S3 performs quantitative analysis on feature drift within the anchor point slippage monitoring period. By detecting changes in the displacement of the light group flash peaks and wheel speed curves, the drift cutoff index is calculated, providing a precise basis for subsequent drift correction.
[0079] Step S3 involves detecting the flash ripple peak displacement of the light group and cross-correlating it with the predicted trajectory of the elastic anchor point to calculate the flash ripple hysteresis. Simultaneously, the third derivative of the wheel speed curve is calculated, and the negative impact is integrated in the braking interval to generate the speed feedback return amplitude. Finally, the two parameters are input into the time-displacement truncation operator to calculate the drift truncation index and write it into the matching memory. The detailed technical logic is as follows:
[0080] Step S3 includes the following:
[0081] S3.1, Detection of flash peak displacement and calculation of flash hysteresis in lamp assembly:
[0082] The purpose of detecting the flicker peak displacement and calculating the flicker hysteresis of the lamp group is to quantify the degree of visual feature drift and provide a timing basis for subsequent drift correction.
[0083] Visual feature vectors, containing lighting brightness information, are extracted from a multi-source co-current tensor. By analyzing the changes in lighting brightness over time, peak detection technology is used to identify the peak brightness points, denoted as the flicker peak positions. The flicker peak displacement is defined as the difference between the flicker peak positions at adjacent time points, specifically calculated by subtracting the previous time point from the current time point's flicker peak position, characterizing the dynamic changes of the flicker in the feature space. Subsequently, cross-correlation analysis is performed between the flicker peak displacement and the elastic anchor point predicted trajectory to calculate the flicker lag. Cross-correlation analysis compares the similarity between the flicker peak displacement and the elastic anchor point predicted trajectory at different time offsets, determining the time offset that maximizes the similarity; this time offset is the flicker lag. The flicker lag represents the time lag of the lighting flicker relative to the elastic anchor point predicted trajectory, measuring the degree of visual feature drift.
[0084] The detection of lamp flash peak displacement and the calculation of flash hysteresis are based on the fact that lamp flash is an important component of vehicle visual characteristics, and its displacement can reflect the actual situation of vehicle motion and attitude changes. Cross-correlation analysis is used to calculate the flash hysteresis, which can accurately quantify the temporal misalignment between the lamp flash and the predicted trajectory of the elastic anchor point. Temporal domain analysis captures the dynamic characteristics of visual feature drift, providing accurate temporal information, thereby improving the accuracy of drift detection.
[0085] S3.2, Calculation of the third derivative of the wheel speed curve and generation of the speed feedback amplitude:
[0086] The purpose of calculating the third derivative of the wheel speed curve and generating the speed feedback return amplitude is to quantify the degree of drift of the vehicle's dynamic characteristics, especially under dynamic conditions such as braking, to provide a dynamic basis for drift correction.
[0087] Wheel speed feature vectors are extracted from the multi-source synchronous tensor, and these feature vectors contain wheel speed data. The third derivative of the wheel speed curve, i.e., jerk, is calculated by differentially dividing the wheel speed data by a third factor. Jerk characterizes the transient properties of wheel speed changes and reflects subtle fluctuations in vehicle dynamics. Next, within the braking range, the third derivative of the wheel speed curve is integrated to calculate the cumulative effect of negative impact, generating the speed feedback reflection amplitude. The braking range is determined by analyzing the second derivative of the wheel speed curve, i.e., acceleration; when the acceleration is negative, the start and end times of braking are identified. The speed feedback reflection amplitude is the integral value of the third derivative of the wheel speed within the braking range, reflecting the cumulative amount of wheel speed dynamic response during braking and used to measure the degree of drift in dynamic characteristics.
[0088] The calculation of the third derivative of the wheel speed curve and the generation of the speed feedback reflection amplitude are based on the characteristic that wheel speed data directly reflects the vehicle's dynamic state. The third derivative can capture the rate of acceleration change, which is especially suitable for analysis during critical moments such as braking. The speed feedback reflection amplitude quantifies the dynamic characteristic drift during braking by integrating the third derivative. High-order differential and integral analysis are used to accurately capture the transient changes in vehicle dynamic characteristics, providing quantitative indicators at the dynamic level.
[0089] S3.3, Calculation and writing of the drift truncation index to the matching memory:
[0090] The purpose of calculating and writing the drift cutoff index into the matching memory is to combine the drift quantization results of visual and dynamic features, generate the drift cutoff index, and store it for subsequent correction.
[0091] The flicker hysteresis and fast feed back amplitude are input into the time-error truncation operator to calculate the drift truncation index. The time-error truncation operator is calculated by multiplying the flicker hysteresis by its corresponding preset scaling factor and then dividing by the fast feed back amplitude by its corresponding preset scaling factor. To avoid a zero denominator, a very small positive number is added as a correction to generate the drift truncation index. The drift truncation index provides a unified drift quantification index by integrating the temporal misalignment of visual and dynamic features. Subsequently, the drift truncation index is written to a matching memory, which is the storage structure for drift information. Each write operation records the drift truncation index corresponding to the current time point, ensuring that the matching memory contains complete historical drift data.
[0092] The calculation and writing of the drift cutoff index into the matching memory is based on the need for comprehensive analysis of drift characteristics from both visual and dynamic features. A time-error truncation operator correlates flicker hysteresis and fast-feedback amplitude to generate a comprehensive drift index. Writing to the matching memory ensures the integrity of the historical drift information. The drift cutoff index provides a simple and effective method for drift quantification, and the matching memory's storage mechanism guarantees the traceability of drift data, thereby improving the accuracy of drift correction and the system's temporal continuity.
[0093] Step S3 begins with the detection of flash peak displacement and the calculation of flash hysteresis to quantify the degree of visual feature drift. Next, it quantifies the degree of dynamic feature drift by calculating the third derivative of the wheel speed curve and generating the speed feedback return amplitude. Finally, it integrates the flash hysteresis and speed feedback return amplitude to calculate the drift cutoff index and writes it into the matching memory. This collectively achieves comprehensive quantitative analysis of feature drift, providing the vehicle model recognition system with efficient drift detection and correction capabilities in nighttime scenarios involving rain and fog on elevated ramps.
[0094] Step S3 calculates the drift truncation index and stores it in the matching memory by detecting the flicker peak displacement and wheel speed curve changes of the light group, providing a quantitative basis for drift correction. However, the drift truncation index only identifies the degree and location of feature misalignment and needs to be transformed into an operable correction tool to align the feature bifurcations of the multi-source co-current tensor on the temporal axis. Step S4, against this backdrop, generates a drift correction matrix and replenishes the multi-source co-current tensor by retrieving high-risk entries from the matching memory, thereby achieving temporal realignment of multi-source features.
[0095] Step S4, based on the matching memory and drift truncation index generated in step S3, generates a drift correction matrix by retrieving high-risk entries and replenishes the multi-source co-current tensor, thereby realigning the multi-source features on the temporal axis. The detailed technical logic is as follows:
[0096] Step S4 includes the following:
[0097] S4.1, Retrieve high-risk entries:
[0098] The process of retrieving high-risk entries aims to identify time points with severe feature drift from the matching memory, providing clear targets for subsequent correction. First, the drift cutoff index is extracted from the matching memory and compared with a preset threshold. Time points where the drift cutoff index exceeds the preset threshold are defined as high-risk entries. The preset threshold is dynamically adjusted based on the interference level of rain and fog scenarios to ensure that the identification of high-risk entries can adapt to changes in different environmental conditions. Next, the difference in drift cutoff indices between adjacent time points is calculated, specifically the difference between the current time point's drift cutoff index and the previous time point's drift cutoff index, to assess the drift trend. If the calculated difference is positive, it indicates that the drift has intensified at the current time point, and this time point is prioritized for inclusion in the high-risk entry set. Subsequently, combining the magnitude of the drift cutoff index and the drift trend, the high-risk entry set is sorted according to the severity of drift, ensuring that time points with the most significant drift impact are addressed first.
[0099] The retrieval of high-risk entries relies on a comprehensive analysis of the drift cutoff index and drift trend, enabling precise identification of time points where feature drift is severe. By introducing a preset threshold that dynamically adjusts based on the degree of scene interference, and prioritizing cases where drift intensifies in conjunction with drift trend analysis, this method ensures the targeted nature and effectiveness of the correction process. By quantifying the severity and trend of drift, the system can optimize the allocation of correction resources, thereby improving adaptability and correction accuracy in complex environments.
[0100] S4.2, Generate the drift correction matrix:
[0101] The process of generating the drift correction matrix aims to construct a transformation tool for high-risk entries to correct feature drift in multi-source co-current tensors. First, the drift correction matrix is defined as a transformation matrix with the same feature dimension as the multi-source co-current tensor. For each high-risk entry, the drift correction matrix is generated based on information about the drift cutoff index and flicker hysteresis. Specifically, the calculation method is as follows: first, take an identity matrix, add the result of multiplying the drift cutoff index by the drift adjustment matrix, and then multiply the result by the correction coefficient to obtain the final drift correction matrix. The correction coefficient is determined based on the interference intensity of diffuse reflection and multipath echoes in the scene to ensure that the correction strength matches the degree of environmental interference. The selection of the drift adjustment matrix depends on the sign of the flicker hysteresis: if the flicker hysteresis is positive, a positive hysteresis adjustment matrix is used; if the flicker hysteresis is negative, a negative hysteresis adjustment matrix is used; if the flicker hysteresis is zero, a zero matrix is used, indicating no adjustment is needed. The design of the drift adjustment matrix fully considers the directionality of feature drift to ensure the accuracy of the correction.
[0102] The drift correction matrix generation method combines information from the drift cutoff index and flicker hysteresis to create a targeted correction tool for each high-risk entry. Through dynamic adjustment of the correction coefficients and directional selection of the drift correction matrix, this method ensures the accuracy and environmental adaptability of the correction process. Employing matrix transformations to correct eigenvectors maintains the structural consistency of the feature space, thereby improving the stability and reliability of the correction effect.
[0103] S4.3, Completing the multi-source co-current tensor:
[0104] The process of restoring the multi-source co-current tensor aims to correct the feature vectors of high-risk entries using a drift correction matrix, thereby achieving realignment of multi-source features along the temporal axis. For each high-risk entry, the feature vector of the multi-source co-current tensor at that time point is transformed using the drift correction matrix. Specifically, the feature vector is multiplied by the drift correction matrix to generate the corrected feature vector. For time points not identified as high-risk entries, their feature vectors remain unchanged and are not adjusted. By applying the above correction operation to all high-risk entries, the feature vector sequence of the multi-source co-current tensor is updated to a sequence containing the corrected feature vectors, thus generating the corrected multi-source co-current tensor and ensuring the alignment consistency of multi-source features along the temporal axis.
[0105] The multi-source co-current tensor restoration is a targeted application based on the drift correction matrix. It corrects only the eigenvectors of high-risk entries while preserving the original state of the eigenvectors of non-high-risk entries. By using local correction rather than global adjustment, it reduces computational resource consumption, improves system processing efficiency, and ensures the temporal continuity and consistency of the eigenvector sequence, providing accurate multi-source feature data for subsequent processing.
[0106] The technical logic of step S4 begins with retrieving high-risk entries and clearly identifying the time points where feature drift is severe. Then, by generating a drift correction matrix, a precise correction tool is constructed for these high-risk entries. Finally, by replenishing the multi-source co-temporal tensor, the correction matrix is applied to complete the correction of the feature vectors, thereby achieving temporal alignment of multi-source features.
[0107] After processing in steps S1 to S4, the data has been fused into a corrected multi-source synchronous tensor, and temporal alignment of features has been achieved through drift correction. The corrected multi-source synchronous tensor includes the contour features of the visual signal, the volumetric echo of millimeter waves, and the dynamic response of wheel speed. The features are realigned on the temporal axis, eliminating drift cracks caused by vehicle pitch, roll, or environmental interference (such as rain, fog, light spots, and guardrail multipath echoes). However, in the face of severe conditions such as vehicle acceleration and lane changes, as well as complex interferences such as diffuse reflection of light spots and transient bumps, temporal alignment alone is insufficient to ensure the accuracy and robustness of vehicle model recognition. Step S5, as the final step of the recognition scheme, needs to receive the corrected multi-source synchronous tensor, perform deep feature extraction and comprehensive analysis on it, generate accurate vehicle model labels, and push them to the billing and security nodes to meet the requirements of accurate toll classification and precise triggering of security policies.
[0108] Step S5 includes the following:
[0109] S5.1, Feature Extraction and Residual Analysis:
[0110] The feature extraction and residual analysis process aims to extract spatiotemporal features from the corrected multi-source co-current tensor and evaluate the matching degree between the features and each vehicle model by comparing them with predefined vehicle model templates. First, a convolutional neural network (CNN) is applied to extract features from the corrected multi-source co-current tensor. The CNN generates feature maps through a sliding window operation of multiple convolutional kernels. These feature maps capture the fusion characteristics of the spatial contours of visual frames, the volume distribution of millimeter-wave echoes, and the dynamic changes in wheel speed curves in the spatiotemporal dimension. The feature map generation process ensures that the comprehensive features of the multi-source data are preserved and enhanced. Next, the generated feature maps are compared with the predefined vehicle model templates. The vehicle model templates represent the spatiotemporal feature distributions of different vehicle models under ideal conditions, containing typical feature patterns of each model. The comparison process generates residual values by calculating the Euclidean distance between the feature maps and the vehicle model templates. The smaller the residual value, the higher the matching degree between the feature map and the corresponding vehicle model template.
[0111] Feature extraction and residual analysis rely on the powerful feature extraction capabilities of convolutional neural networks (CNNs) and the intuitive matching and evaluation capabilities of Euclidean distance. CNNs, through the operation of multiple convolutional kernels, can automatically learn and extract high-level features from corrected multi-source simultaneous tensors, adapting to feature variations in complex scenarios. Residual analysis, by comparing with predefined vehicle model templates, can directly quantify the similarity between feature maps and various vehicle models.
[0112] S5.2, Distance Calculation in Power Mode:
[0113] The process of calculating the distance between dynamic modes aims to extract dynamic features from the corrected multi-source synchronous tensor and evaluate the degree of matching of these dynamic features by comparing them with the standard dynamic modes of each vehicle model. First, wheel speed-related features are extracted from the corrected multi-source synchronous tensor to generate dynamic feature vectors. These vectors contain dynamic parameters such as velocity and acceleration, reflecting the vehicle's motion state at a given time point. The extraction process is based on the temporal components of the wheel speed curve, ensuring that the dynamic features are aligned with the features of the visual frame and millimeter-wave echo on the time axis. Next, a dynamic time warping algorithm is used to calculate the distance between the dynamic feature vectors and the standard dynamic modes of each vehicle model. The standard dynamic mode is a sequence of dynamic features for different vehicle models under standard conditions. The dynamic time warping algorithm works by comparing each pair of points in the time series between the dynamic feature vectors and the standard dynamic modes, finding the optimal matching path between the two sequences, and accumulating the minimum matching cost between points on the path to obtain the total distance value. The smaller the distance value, the higher the degree of matching between the dynamic feature vectors and the corresponding standard dynamic mode.
[0114] The calculation of dynamic pattern distance utilizes the Dynamic Time Warping (VTW) algorithm to process time-series data, effectively addressing the nonlinear variations and time delays of dynamic features along the time axis. The VTW algorithm provides a robust method for measuring temporal similarity by finding the optimal matching path and calculating the minimum matching cost. The extraction and comparison process of dynamic feature vectors enhances the recognition system's ability to perceive vehicle dynamic behavior, significantly improving recognition accuracy, especially under severe conditions such as acceleration and braking. Using a standard dynamic pattern as the comparison benchmark ensures the controllability and consistency of the matching process.
[0115] S5.3, Vehicle Model Label Recognition and Output:
[0116] The vehicle model label identification and output process aims to integrate the results of feature extraction, residual analysis, and power mode distance calculation to generate vehicle model labels and push them to the billing and security nodes. First, a comprehensive discrimination score is calculated by combining the residual values generated from residual analysis and the distance values generated from power mode distance calculation. The comprehensive discrimination score is calculated by multiplying the reciprocal of the residual value and the reciprocal of the distance value by preset weighting coefficients and then summing them. The reciprocal operation transforms the residual values and distance values into matching metrics, and the weighting coefficients are used to adjust the contribution of the residual analysis results and power mode distance results in the discrimination; a higher score indicates a higher matching degree. Next, at each time point, the comprehensive discrimination scores of all vehicle model categories are compared, and the vehicle model category with the highest score is selected as the identification result, generating a vehicle model label. Finally, the generated vehicle model label is pushed to the billing and security nodes for subsequent charging classification and security policy triggering.
[0117] Vehicle model label discrimination and output integrates matching information from feature maps and vehicle model templates, as well as matching information from dynamic feature vectors and standard power patterns, through comprehensive discrimination scores, ensuring the comprehensiveness and accuracy of the recognition results. The introduction of weighting coefficients allows the system to adjust the contribution ratio of residual analysis and power pattern distance according to different scenarios and operating conditions, enhancing the system's flexibility and adaptability. The output of vehicle model labels directly serves the practical application needs of billing and safety nodes, improving the system's practicality and reliability.
[0118] The technical logic of step S5 starts with feature extraction and residual analysis. It uses a convolutional neural network to extract the spatiotemporal features of the corrected multi-source synchronous tensor and generates residual values by comparing them with a predefined vehicle model template. Then, it extracts dynamic feature vectors by calculating the distance of the power mode and generates distance values by comparing them with the standard power mode. Finally, it integrates the residual values and distance values by comprehensively judging the score, generates vehicle model tags, and pushes them to the billing and safety nodes.
[0119] Example 2: Figure 2 The present invention provides a vehicle model recognition system based on multimodal perception and dynamic feature matching, comprising:
[0120] Signal fusion module: preprocesses multi-source signals and constructs multi-source co-current tensors using adaptive time-scale alignment technology;
[0121] Feature tracking module: Extracts optical flow peaks and echo peaks to construct a pseudo-topological tracking mesh, dynamically deploys elastic anchor points and refreshes the index in real time;
[0122] Drift detection module: The flicker hysteresis is obtained by detecting the displacement of the flash peak of the lamp group and cross-correlating it with the predicted trajectory of the elastic anchor point. The speed feedback return amplitude generated by the third derivative of the wheel speed curve and the braking interval integral is combined with the two to perform a comprehensive analysis and write it into the matching memory.
[0123] Drift correction module: retrieves high-risk entries from the matching memory, generates a drift correction matrix, and backfills the multi-source co-current tensor;
[0124] Vehicle model recognition module: Performs feature extraction and matching analysis on the corrected multi-source simultaneous tensor, comprehensively evaluates the fit between the features and the vehicle model standard, determines the vehicle model category label, and pushes it to the billing and security nodes.
[0125] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0126] It should be noted that the system of the present invention can be deployed on the device itself to realize embedded applications, or it can run on a PC or other terminal with a user interface, thereby meeting a variety of hardware environments and usage requirements.
[0127] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
[0128] It should be noted that, in this document, the use of relational terms such as "first" and "second" is merely to distinguish one entity or operation from another, and does not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0129] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A vehicle model recognition method based on multimodal perception and dynamic feature matching, characterized in that, Including the following steps: S1: Preprocess multi-source signals and construct multi-source co-current tensors using adaptive time-scale alignment techniques; S2: Extract optical flow peaks and echo peaks to construct a pseudo-topological tracking mesh, dynamically deploy elastic anchor points and refresh the index in real time; Step S2 includes the following: Visual feature vectors and millimeter-wave echo feature vectors are extracted from multi-source synchronous tensors. Motion vectors are generated by calculating the differences between visual feature vectors at adjacent time points. The point with the largest motion vector amplitude is identified as the optical flow peak. The point with the largest signal intensity is identified from the millimeter-wave echo feature vector as the echo peak. An initial grid is constructed using the optical flow peak and the echo peak as nodes. The topological relationship is determined by calculating the spatial distance between nodes and the consistency of motion direction. The node positions are updated over time to adapt to vehicle motion, forming a pseudo-topological tracking grid. In the pseudo-topological tracking grid, an elastic anchor point is selected. The initial position is the weighted center of the optical flow peak and the echo peak. The weight is dynamically adjusted according to the signal reliability. The anchor point position is dynamically updated and the index is refreshed by comparing the deviation between the current position and the predicted position of the elastic anchor point, so as to realize the spatiotemporal continuous tracking of vehicle features. S3: The flash hysteresis is obtained by detecting the peak displacement of the lamp group and cross-correlating it with the predicted trajectory of the elastic anchor point. The speed feedback return amplitude generated by the third derivative of the wheel speed curve and the integral of the braking interval is combined with the two for comprehensive analysis and written into the matching memory. Step S3 includes the following: Visual feature vectors are extracted from the multi-source synchronous tensor and the displacement of the flash peak of the lamp group is detected. The flash hysteresis is calculated by cross-correlation analysis with the predicted trajectory of the anchor point. Wheel speed feature vectors are extracted from the multi-source synchronous tensor, the third derivative of the wheel speed curve is calculated, and the third derivative is integrated in the braking interval to generate the speed feedback return amplitude. Input the flicker hysteresis and fast feedback return amplitude into the time error truncation operator, calculate the drift truncation index, and write the drift truncation index into the matching memory; The logic for obtaining the flicker peak displacement of the lamp group is as follows: analyze the change of lamp group brightness over time, identify the peak brightness point of the lamp group through peak detection, and record the position of the peak point as the flicker peak position of the lamp group; calculate the difference between the flicker peak positions of the lamp group corresponding to adjacent time points to obtain the flicker peak displacement of the lamp group. The calculation method of the time-fault cutoff operator is to divide the product of the flicker hysteresis and its corresponding preset proportional coefficient by the product of the fast feed back amplitude and its corresponding preset proportional coefficient; to avoid the case where the denominator is zero, a very small positive number is added as a correction to generate the drift cutoff index. S4: Retrieve high-risk entries from the matching memory, generate a drift correction matrix, and backfill the multi-source co-current tensor; S5: Perform feature extraction and matching analysis on the corrected multi-source co-current tensor, comprehensively evaluate the fit between the features and the vehicle model standard, determine the vehicle model category label, and push it to the billing and security nodes.
2. The vehicle model recognition method based on multimodal perception and dynamic feature matching according to claim 1, characterized in that, Step S1 includes the following: Visual frames, millimeter-wave echoes, and wheel speed data are acquired and preprocessed. Using the sampling frequency of the millimeter-wave echoes as the reference time axis, visual frames and wheel speed data are aligned by interpolation, features are extracted, and they are concatenated into a comprehensive feature vector to construct a multi-source co-current tensor.
3. The vehicle model recognition method based on multimodal perception and dynamic feature matching according to claim 2, characterized in that, Step S4 includes the following: High-risk entries are retrieved from the matching memory, and the time points when the drift cutoff index exceeds a preset threshold are identified. High-risk entries are sorted by drift severity according to the drift cutoff index and its changing trend. A drift correction matrix is generated using the drift cutoff index and flicker hysteresis. The drift correction matrix contains correction coefficients that reflect the intensity of interference.
4. The vehicle model recognition method based on multimodal perception and dynamic feature matching according to claim 3, characterized in that, Step S4 also includes the following: The drift correction matrix is applied to the corresponding time point feature vectors of the multi-source co-current tensor to correct the feature vectors and realign the multi-source features on the time axis.
5. The vehicle model recognition method based on multimodal perception and dynamic feature matching according to claim 4, characterized in that, Step S5 includes the following: The system receives the corrected multi-source synchronous tensor, extracts features from it using a convolutional neural network to generate a feature map, calculates the Euclidean distance between the feature map and a predefined vehicle model template to generate a residual value, extracts wheel speed-related features from the corrected multi-source synchronous tensor to generate a dynamic feature vector, and uses a dynamic time warping algorithm to calculate the distance between the dynamic feature vector and the standard power mode of each vehicle model.
6. The vehicle model recognition method based on multimodal perception and dynamic feature matching according to claim 5, characterized in that, Step S5 also Includes the following: The reciprocal of the residual value and the reciprocal of the distance value are combined and weighted to generate a comprehensive discrimination score. The vehicle category with the highest comprehensive discrimination score is selected as the vehicle label and pushed to the billing and safety node.
7. A vehicle model recognition system based on multimodal perception and dynamic feature matching, used to implement the vehicle model recognition method based on multimodal perception and dynamic feature matching as described in any one of claims 1-6, characterized in that, include: Signal fusion module: preprocesses multi-source signals and constructs multi-source co-current tensors using adaptive time-scale alignment technology; Feature tracking module: Extracts optical flow peaks and echo peaks to construct a pseudo-topological tracking mesh, dynamically deploys elastic anchor points and refreshes the index in real time; Drift detection module: The flicker hysteresis is obtained by detecting the displacement of the flash peak of the lamp group and cross-correlating it with the predicted trajectory of the elastic anchor point. The speed feedback return amplitude generated by the third derivative of the wheel speed curve and the braking interval integral is combined with the two to perform a comprehensive analysis and write it into the matching memory. Drift correction module: retrieves high-risk entries from the matching memory, generates a drift correction matrix, and backfills the multi-source co-current tensor; Vehicle model recognition module: Performs feature extraction and matching analysis on the corrected multi-source simultaneous tensor, comprehensively evaluates the fit between the features and the vehicle model standard, determines the vehicle model category label, and pushes it to the billing and security nodes.
Citation Information
Patent Citations
Intelligent target distribution method and system for collaborative interception
CN119485687A
Systems and Methods for Localizing an Autonomous Vehicle
US20250091604A1