A multi-modal fusion vehicle outer contour measurement system and a measurement method
The multimodal fusion vehicle profile measurement system utilizes time synchronization and lightweight neural networks to achieve spatiotemporal synchronization and feature fusion of multimodal data, solving the problem of limited measurement accuracy and robustness in existing technologies and improving the accuracy and adaptability of vehicle profile measurement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI JIANCHUANG SPACE INFORMATION TECH CO LTD
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-29
AI Technical Summary
Existing vehicle profile measurement technologies lack precise spatiotemporal synchronization, depth complementary perception, and continuous iteration capabilities, resulting in limited measurement accuracy and robustness. In particular, it is difficult to achieve high-precision vehicle profile measurement in complex environments.
A multimodal fusion vehicle profile measurement system is adopted, including a terminal perception layer, an edge processing layer, and a collaborative management layer. The spatiotemporal synchronization of multimodal sensors is achieved through a timing synchronization component. Data processing and learning are combined with a lightweight multimodal neural network to ensure that multimodal data can complement and fuse features under the same spatiotemporal reference.
It achieves deep fusion of multimodal data in the physical and semantic dimensions, improves the robustness and accuracy of the measurement system, can accurately identify vehicle attributes and precisely extract physical dimensions in complex environments, and has the ability to continuously iterate and optimize.
Smart Images

Figure CN122116309A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle profile measurement, and more particularly to a multimodal fusion vehicle profile measurement system and method. Background Technology
[0002] Vehicle external dimension measurement is a core component of intelligent transportation and off-site enforcement systems, playing a crucial role in improving the efficiency of overloading control and ensuring road safety. Achieving high-precision, automated dimension estimation is not only key to reducing costs and increasing efficiency in overloading control, but also the technological foundation for building an intelligent regulatory system.
[0003] However, existing sensing solutions still have significant limitations in practical applications. First, due to the lack of effective spatiotemporal synchronization and triggering constraints, sampling delays and observation differences are common among heterogeneous sensors, making it difficult to guarantee the consistency of multimodal data. Second, traditional technologies often rely on single modalities or simple feature superposition, making it difficult to achieve deep correlation between semantic attributes and geometric accuracy under conditions of drastic changes in illumination or complex occlusion. This hinders the effective construction of multimodal fusion objects for the vehicle under test, resulting in limited measurement accuracy and robustness. Furthermore, existing systems mostly employ static preset algorithms, lacking adaptive iteration and model evolution capabilities based on real data samples. Performance cannot be continuously optimized with the accumulation of sample size, limiting the long-term stability of the system in complex and variable environments.
[0004] In summary, existing technologies lack a vehicle outline measurement solution that features precise spatiotemporal synchronization, deep complementary perception, and continuous iterative evolution. Summary of the Invention
[0005] In order to overcome the above-mentioned technical defects, the purpose of this invention is to provide a multimodal fusion vehicle outline measurement system and measurement method.
[0006] This invention discloses a multimodal fusion vehicle profile measurement system, comprising: Terminal perception layer; the terminal perception layer includes multimodal sensor components and timing synchronization components; the multimodal sensor components include camera modules and lidar modules, which measure the outline of the vehicle under test through various measurement methods; the timing synchronization components include a time synchronization module, which is used to synchronize the system time of each module in the multimodal sensor components; Edge processing layer; The edge processing layer includes computing units deployed with lightweight multimodal neural networks and is communicatively connected to the terminal perception layer; It is used to process the information collected by the terminal perception layer and obtain preliminary measurement results; The collaborative management layer communicates with both the terminal perception layer and the edge processing layer to verify the preliminary measurement results; and learns based on the processing results of the edge processing layer, and maintains a lightweight multimodal neural network based on the learning results.
[0007] Preferably, the time synchronization component includes at least one of the following modules: Attitude reference module; The attitude reference module includes an inertial measurement unit, which is used to acquire the attitude information of the terminal perception layer in real time, and provide attitude compensation parameters for the edge processing layer based on the attitude information; Spatiotemporal reference module; used to provide a unified time and position reference for each module in the multimodal sensor assembly; The time synchronization module is used to distribute the high-precision time signal or local time signal output by the spatiotemporal reference module to the multimodal sensor components and the edge processing layer to achieve time alignment between the multimodal sensor components and the edge processing layer.
[0008] A second aspect of the present invention discloses a multimodal fusion vehicle outline measurement method, based on any one of the foregoing multimodal fusion vehicle outline measurement systems, the method comprising: Based on the timing synchronization component, the multimodal sensor component is synchronously triggered to sample the vehicle under test and acquire multimodal acquisition data at the same timestamp; the multimodal acquisition data includes point cloud data acquired by the lidar module and image data acquired by the camera module; Spatiotemporal alignment and delay compensation are performed on multimodal acquisition data to unify the multimodal acquisition data under the same spatiotemporal reference. Multimodal target detection is performed by the edge processing layer, and the multimodal acquired data is matched to construct a multimodal fusion object of the vehicle under test, thereby achieving complementary and unified data features between different modalities; Based on the multimodal fusion object, the three-dimensional outline dimensions of the vehicle under test are estimated; The multimodal fusion object and its 3D outline dimensions are uploaded to the collaborative management layer, where they are archived and studied.
[0009] Preferably, based on the timing synchronization component, the multimodal sensor component is synchronously triggered to sample the vehicle under test, and the multimodal acquisition data at the same timestamp includes: Hardware timing trigger: Utilize GNSS timing, PPS or PTP hardware timing units to provide a unified sampling trigger signal and high-precision timestamp for each sensor; Standardized encapsulation: Multimodal acquisition data is written into a unified buffer queue according to a standard frame structure, and the status, including frame loss and over-temperature, is recorded.
[0010] Preferably, spatiotemporal alignment and delay compensation are performed on the multimodal acquisition data to unify the multimodal acquisition data under the same spatiotemporal reference, including: Time interpolation: Based on a unified time standard, interpolation and reconstruction are performed on multimodal acquisition data to map the multimodal acquisition data to the same reference time. Delay compensation: Estimate the total communication delay of each module in the multimodal sensor assembly, and suppress noise by smoothing filter on the total delay to obtain the compensated multimodal acquisition data; Motion correction: Geometric correction is performed on each sampling point using the attitude change information of the carrier to eliminate data distortion caused by the motion of the vehicle under test; Quality detection: Based on the multimodal acquisition data of the previous moment, predict the multimodal acquisition data of the current moment; compare the predicted multimodal acquisition data with the multimodal acquisition data acquired at the current moment. If the difference is greater than a preset threshold, mark the moment as low-quality alignment and use the predicted multimodal acquisition data for subsequent operations.
[0011] Preferably, performing multimodal target detection includes: Perform forward inference based on convolutional neural networks on image data to extract visual feature vectors of the target, including semantic category and encoded appearance, texture information; Perform point cloud neural network processing on point cloud data to extract geometric feature vectors of the target, including its size, shape, and spatial distribution; Semantic discrimination is performed using visual feature vectors, and three-dimensional geometric precision and scale are provided using geometric feature vectors to obtain a multimodal feature set of the vehicle under test.
[0012] Preferably, matching the multimodal acquired data includes: Geometric gated filtering: Projecting point cloud data onto the image plane in the image data, and performing depth statistical consistency verification on the point cloud data falling within the detection range using the median absolute deviation to remove background outliers; Association confidence assessment: Using a pre-defined neural network model, the association confidence of geometric feature vectors and visual feature vectors is scored; Global optimal matching: A cost matrix is constructed based on the association confidence score, and a global optimization algorithm is used to solve for the best matching pair between the image target and the point cloud instance in order to construct the multimodal fusion object of the vehicle under test.
[0013] Preferably, matching the detected multimodal data further includes: Adhesion risk assessment: A Gaussian mixture model is used to fit the distribution pattern of the point cloud data on a plane perpendicular to the forward direction of the vehicle under test, and the separation index is calculated to identify whether there is a risk of multi-vehicle adhesion. Secondary segmentation of adhesion: If there is a risk of multiple vehicles being stuck together, the stuck point cloud data is split into multiple independent vehicle point cloud sets by combining the abrupt change locations of the point cloud data, the outline boundary of the vehicle under test identified by the image data, and the speed difference of each point in the point cloud data. Multimodal fusion object reconstruction: The split vehicle point cloud set is used as point cloud data, and multimodal data matching is performed again to construct a multimodal fusion object.
[0014] Preferably, estimating the three-dimensional outline dimensions of the vehicle under test based on the multimodal fusion object includes: Establish a local coordinate system: assign weights to the point cloud data in the multimodal fusion object, determine the yaw angle of the vehicle under test through the weighted principal component analysis algorithm, and transform the point cloud data into a three-dimensional coordinate system constructed with the length, width and height of the vehicle under test according to the yaw angle; Robust dimensional calculation: The point cloud distribution along each axis of the three-dimensional coordinate system is statistically analyzed, the difference between the high quantile and the low quantile is calculated, and the difference is used as the estimated three-dimensional dimensions of the vehicle under test.
[0015] Preferably, estimating the three-dimensional outline dimensions of the vehicle under test based on the multimodal fusion object further includes: Constructing the bounding box: Based on the estimated 3D dimensions, construct the bounding box of the vehicle under test in the 3D coordinate system; Calculate the fitting: Calculate the distance vector from each point in the point cloud data to the bounding box surface, and determine the fitting quality index based on the statistical distribution of the distance vector; Determine measurement reliability: Calculate the reliability of this size estimate based on the fitting quality index, and use it as the basis for judging the over-limit alarm.
[0016] Compared with existing technologies, the above technical solution has the following advantages: 1. This application constructs a multimodal fusion measurement system for exterior contour measurement, achieving deep fusion of multimodal data in both physical and semantic dimensions. The scheme ensures temporal consistency at the sensing source through a timing synchronization component, eliminates observational differences between heterogeneous sensors through spatiotemporal alignment, and combines multimodal target detection to achieve complementary advantages between visual semantics and point cloud geometry. Compared to traditional measurement methods, this scheme effectively overcomes the perception bottleneck of a single sensor under complex conditions such as adverse weather, drastic changes in lighting, or target occlusion, ensuring that the system can accurately identify vehicle attributes and precisely extract physical dimensions, significantly improving the overall robustness of the sensing system and the objectivity and accuracy of the measurement results. Furthermore, through archiving by the collaborative management layer, the measurement data and results can be learned to continuously iterate the measurement system and further improve its measurement accuracy. 2. A three-layer architecture of terminal perception, edge processing, and collaborative management, combined with a lightweight edge-side neural network and a cloud-based closed-loop learning mechanism, ensures real-time perception and continuous algorithm evolution. At the hardware level, by integrating an attitude reference module and a timing synchronization component, and utilizing hardware-level trigger signals (such as PPS / PTP) to achieve microsecond-level sampling alignment, the uncertainty and jitter caused by software polling are completely eliminated. Simultaneously, real-time pose information provided by the inertial measurement unit provides dynamic compensation for the edge processing layer, effectively correcting perception biases caused by gantry vibration or environmental interference. This in-depth design, from the physical layer to the computing architecture, ensures high fidelity of the original data and the integrity of the evidence chain. 3. In the perception processing flow, the system eliminates point cloud distortion caused by sensor communication delays and vehicle motion through a four-level alignment logic of time interpolation, delay compensation, and motion correction, ensuring seamless integration of multi-source observations under the same spatiotemporal reference. Based on this, the system utilizes a parallel neural network to extract the target's apparent texture and spatial geometric features, establishing a depth consistency gating system using the median absolute deviation to remove background noise that mistakenly enters the detection box due to overlapping viewpoints. Subsequently, a pre-defined neural network model is used to comprehensively evaluate spatial geometric compatibility and cross-modal semantic similarity, achieving deep coupling between visual features and point cloud geometric features, and intelligently scoring the generated multimodal fusion object. This feature-level complementary mechanism solves the problem of mismatch in complex backgrounds, ensuring the uniqueness and continuity of each physical instance. 4. To address the issue of point cloud adhesion, a Gaussian mixture model is used to fit the lateral distribution pattern. The separation index is used to scientifically diagnose the adhesion risk of vehicles side-by-side. Furthermore, a secondary segmentation process is performed, integrating point cloud density abrupt changes, visual contour constraints, and motion vector differences, to suppress the misconstruction of multimodal fusion objects caused by multi-vehicle adhesion and the resulting severe overwidth false alarms. In the final measurement stage, weighted principal component analysis is used to correct the vehicle's posture, and quantile statistics are used to remove noise flypoints. Finally, a directed bounding box is constructed, and the point cloud fitting residual is calculated to provide a quantitative reliability assessment for each measurement conclusion. This decision-making mechanism based on robust statistics and a quality closed loop ensures that the over-limit alarm conclusions have clear statistical significance and legal credibility. Attached Figure Description
[0017] Figure 1 Functional architecture diagram of the multimodal fusion vehicle outline measurement system provided in this application; Figure 2 Hardware architecture diagram of the multimodal fusion vehicle outline measurement system provided in this application; Figure 3 A schematic diagram of the integrated intelligent measurement terminal provided in this application; Figure 4A flowchart illustrating the multimodal fusion vehicle outline measurement method provided in this application; Figure 5 An example diagram illustrating the construction of the bounding box in the multimodal fusion vehicle outline measurement method provided in this application; Figure 6 The maintenance architecture diagram of the neural network model in the multimodal fusion vehicle outline measurement method provided in this application; Figure 7 This is a schematic diagram of the visual monitoring and interactive interface provided in this application. Detailed Implementation
[0018] The advantages of the present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments.
[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0020] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0021] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, can be interpreted as "when," "in response to determination," or "when," or "in the event of a determination." In the description of this invention, it should be understood that the terms "longitudinal", "lateral", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0022] In the description of this invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection" and "linking" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two components. They can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.
[0023] In the following description, suffixes such as "module," "part," or "unit" used to denote elements are used only for the convenience of the description of the invention and have no specific meaning in themselves. Therefore, "module" and "part" can be used interchangeably.
[0024] Please see Figures 1-3 , Figure 1 Functional architecture diagram of the multimodal fusion vehicle outline measurement system provided in this application; Figure 2 Hardware architecture diagram of the multimodal fusion vehicle outline measurement system provided in this application; Figure 3 A schematic diagram of the integrated intelligent measurement terminal provided in this application.
[0025] like Figures 1-3 As shown, the first aspect of the present invention discloses a multimodal fusion vehicle profile measurement system, comprising: Terminal perception layer; the terminal perception layer includes multimodal sensor components and timing synchronization components; the multimodal sensor components include camera modules and lidar modules, which measure the outline of the vehicle under test through various measurement methods; the timing synchronization components include a time synchronization module, which is used to synchronize the system time of each module in the multimodal sensor components; Edge processing layer; The edge processing layer includes computing units deployed with lightweight multimodal neural networks and is communicatively connected to the terminal perception layer; It is used to process the information collected by the terminal perception layer and obtain preliminary measurement results; The collaborative management layer communicates with both the terminal perception layer and the edge processing layer to verify the preliminary measurement results; and learns based on the processing results of the edge processing layer, and maintains a lightweight multimodal neural network based on the learning results.
[0026] This application constructs a multimodal fusion measurement system for exterior contour measurement, achieving deep fusion of multimodal data in both physical and semantic dimensions. The scheme ensures temporal consistency of the sensing sources through a timing synchronization component, eliminates observational discrepancies between heterogeneous sensors through spatiotemporal alignment, and combines multimodal target detection to achieve complementary advantages between visual semantics and point cloud geometry. Compared to traditional measurement methods, this scheme effectively overcomes the perception bottleneck of single sensors under complex conditions such as adverse weather, drastic changes in lighting, or target occlusion, ensuring that the system can accurately identify vehicle attributes and precisely extract physical dimensions, significantly improving the overall robustness of the perception system and the objectivity and accuracy of the measurement results.
[0027] Furthermore, based on this layered, decoupled, and collaborative system, this invention ensures real-time perception and instant control performance on the roadside while endowing the system with global optimization and continuous evolution capabilities. It proposes an engineering system that is elastic, scalable, manageable, and algorithmically evolvable, significantly enhancing the reliability, adaptability, and long-term viability of the entire technology in real-world complex environments.
[0028] The above is an explanation of the basic concept of this application. The specific implementation of the solution will be described below.
[0029] like Figures 2-3 As shown, in one possible implementation, the terminal sensing layer and edge processing layer of the multimodal fusion vehicle profile measurement system proposed in this invention can be integrated into an integrated intelligent measurement terminal. This integrated intelligent measurement terminal adopts a modular design, including a multi-source sensing module (which can be understood as a multimodal sensor assembly), an edge intelligent computing module (which can be understood as an edge processing layer), a communication management module (used for communication with the collaborative management layer), and a power management module, all highly integrated within a protective, heat-dissipating, and environmentally adaptable chassis, forming a fully functional, independently operable intelligent sensing and computing node.
[0030] Those skilled in the art will understand that there are no restrictions on the specific implementation methods of the aforementioned modules.
[0031] In one possible implementation, the camera module in the multimodal sensor assembly is configured to: employ a high-resolution industrial camera, preferably with wide dynamic range, low-light imaging capability and anti-glare design, support the configuration of a supplementary lighting device, and acquire high-quality image information including vehicle body texture, color, license plate and lane environment, providing basic data for vehicle target detection, semantic segmentation and outer contour boundary extraction.
[0032] The lidar is configured as a solid-state or hybrid solid-state lidar, featuring a high beam count and ranging accuracy, capable of stably outputting 3D point cloud data in complex environments such as rain, fog, strong light, and backlight. By clustering and contour fitting the spatial reflection points on the vehicle surface, the 3D geometry and spatial position of the vehicle's outline are obtained, providing dominant information for accurate measurement of the outline dimensions.
[0033] Furthermore, the time synchronization component includes at least one of the following modules: Attitude Reference Module: The attitude reference module includes an inertial measurement unit (IMU) for real-time acquisition of attitude information from the terminal's perception layer and for providing attitude compensation parameters to the edge processing layer based on this information. Specifically, it can be an IMU or similar unit used to acquire real-time attitude change information (including pitch, roll, and yaw angles) of the terminal mounting bracket according to usage requirements. This information is used to provide compensation parameters to reduce the impact of factors such as gantry vibration, foundation settlement, and thermal deformation on calibration results, ensuring the stability and accuracy of external parameters during long-term operation.
[0034] The spatiotemporal reference module provides a unified time and position reference for each module in the multimodal sensor assembly. Specifically, it can be a BeiDou / GPS dual-mode timing receiver, a GNSS timing module, or a ground-based timing device. This provides the terminal with a unified UTC (Coordinated Universal Time) time and position reference, enabling clock calibration and coordinate systems between the terminal and multiple front-end devices, and providing spatiotemporal reference for cross-lane and cross-site data association and joint analysis.
[0035] The time synchronization module distributes the high-precision time signal output by the spatiotemporal reference module or the local time signal to the multimodal sensor components and the edge processing layer to achieve time alignment between the multimodal sensor components and the edge processing layer. Specifically, it can select a PPS / IRIG-B hardware timing unit, an IEEE1588PTP timing unit, or a combination thereof. This allows the high-precision time signal output by the spatiotemporal reference module to be distributed to the camera module, the lidar module, and the edge processing layer. By outputting PPS pulses, synchronization trigger signals, and a unified timestamp, it achieves microsecond-level alignment of the acquisition times of multiple sensors. When external timing is unavailable, it can switch to a local high-stability oscillator for timekeeping and monitor and alarm on the synchronization quality.
[0036] Those skilled in the art will understand that the above sub-modules can be set individually or combined arbitrarily as needed to meet the spatiotemporal synchronization accuracy requirements of different engineering scenarios. Through the collaboration of attitude reference, spatiotemporal reference, and time synchronization modules, a highly reliable multidimensional spatiotemporal reference is established. The attitude compensation parameters provided by the inertial measurement unit (IMU) effectively correct the pose deviations of the sensors caused by dynamic vibrations; a unified UTC time and position reference ensures the consistency between global positioning and local sensing; and microsecond-level time alignment eliminates the sampling time difference between heterogeneous sensors, preventing fusion failures caused by spatiotemporal misalignment from a physical level, laying a solid foundation for high-precision sensing with high data fidelity.
[0037] In one possible implementation, the edge processing layer can be integrated into the chassis as an edge intelligent computing module. This module can employ an embedded computing platform or an industrial-grade edge server, integrating a multi-core CPU, a graphics processing unit (GPU), and / or a neural network acceleration unit (NPU), and configured with high-speed memory and solid-state storage to meet the real-time computing requirements of high-frame-rate point clouds and videos. The edge intelligent computing module can complete most of the high-bandwidth, real-time intelligent computing tasks close to the data source, significantly reducing reliance on cloud computing power and link bandwidth. This enables millisecond-level measurement of vehicle dimensions and rapid over-limit assessment, while providing a structured, high-quality data foundation for subsequent secondary verification, model iteration, and cross-site collaborative analysis in the cloud.
[0038] Optionally, the integrated intelligent measurement terminal of the present invention supports a split deployment scheme. By reserving high-speed interfaces such as Ethernet and PCIe, it standardizes the connection between external edge computing units such as industrial computers and edge servers and the terminal chassis to adapt to different scale scenarios and existing system transformation needs.
[0039] The communication management module is used to realize data interaction and control command transmission between the integrated intelligent measurement terminal and external systems. It is a key interface connecting end-side devices, the network transmission layer, and the cloud platform. This module can integrate hardware and software components such as industrial Ethernet interfaces, wireless communication units, and secure access units, supporting multi-link access and redundancy switching to ensure reliable reporting and transmission of vehicle outline measurement data and over-limit alarm information in complex field environments. The control and communication module integrates 4G / 5G communication modules, satellite communication modules, Ethernet / WIFI interfaces, and time synchronization units to realize data uploading, remote updates, device management, and multi-terminal collaboration. Optionally, the communication management module includes, but is not limited to, the following sub-modules: Wired access module: Provides a gigabit Ethernet interface for accessing fiber optic leased lines, campus networks or government private networks to achieve high-speed and stable communication with the upper-level platform or roadside control system; optional integrated industrial switch functions support port isolation, VLAN segmentation and link redundancy.
[0040] Wireless communication module: Configurable with 4G / 5G cellular communication module, Wi-Fi module and optional satellite communication module, used to realize wireless transmission of measurement results, alarm information and electronic evidence packets in wired conditions or emergency scenarios, and supports multi-link optimization, link health monitoring and automatic switching to ensure the basic communication capabilities of critical services in weak network or network outage conditions.
[0041] Edge control module: Reserves industrial interfaces such as RS485, CAN, LVDS, RS422, and digital I / O for linkage with external devices such as information boards, audible and visual alarms, barrier gates, and road gate controllers to realize functions such as local alarm display, lane control, and on-site prompts; supports the forwarding and execution of control commands issued from the cloud or triggered by local rules.
[0042] The communication management module supports multiple upper-layer protocols such as TCP / IP, MQTT, HTTP / HTTPS, and WebSocket, enabling the uploading of structured data, alarm events, and log information encapsulated according to business needs. Optional deployment of VPN, dedicated tunnel, or APN private network access ensures encrypted transmission and secure isolation with the cloud platform. The communication management module incorporates a link status monitoring and fault alarm mechanism, reporting communication quality, packet loss rate, latency, and other operational indicators to the terminal edge intelligent computing module in real time. When the system detects an anomaly in the main link, it automatically switches to a backup link or enters a "local caching + delayed upload" mode, ensuring that measurement data and evidence information are not lost and alarm linkage is not interrupted even under fluctuating network conditions or short-term interruptions. This improves the reliability and engineering availability of the entire vehicle exterior intelligent measurement and over-limit automatic alarm system.
[0043] In one possible implementation, a power management module is integrated as the core power supply and control component of the integrated intelligent measurement terminal. It is responsible for providing centralized, stable, and monitorable power to all core modules, including multi-source sensing, edge intelligent computing, and communication management. This is a crucial foundation for ensuring the continuous and robust operation of the entire system and achieving long-term unattended operation. The power management module supports one or more combinations of power inputs from various power sources, such as AC380V, 220V, DC48V, and backup batteries. It outputs multiple isolated DC12V power supplies to power sensors, the computing platform, and communication equipment separately. This unit integrates overvoltage, overcurrent, short-circuit, and reverse connection protection functions to ensure a stable and controlled operating power supply to each functional module even under conditions of power grid fluctuations and power surges.
[0044] The power management module has built-in circuitry for collecting voltage, current, and temperature data, monitoring the output status of each power source, battery remaining power, and health status in real time. It then reports this information to the edge intelligent computing module and cloud platform in the form of status indicators or alarm flags. When the front end detects abnormal voltage, current, or temperature exceeding limits, it can trigger power degradation strategies (such as reducing computational load or sensor operating frequency) or execute a safety shutdown procedure. The cloud platform can also perform operation and maintenance analysis and formulate maintenance plans based on historical power status data.
[0045] Optionally, the power management module includes AC / DC surge protectors, common-mode / differential-mode filters, and grounding surge protection units to mitigate the impact of lightning strikes, power grid surges, and the start-up and shutdown of high-power equipment on the terminal power supply lines. This reduces the influence of electromagnetic interference on sensors and computing devices, improving the system's immunity and reliability in harsh outdoor electromagnetic environments. The power management module provides a highly reliable, multi-source redundant, monitorable, and controllable power supply foundation for the vehicle outline intelligent measurement and over-limit automatic alarm system. This not only improves the terminal's survivability under power grid fluctuations, lightning interference, and power outages, but also provides crucial data support for subsequent remote operation and maintenance and health management, thereby significantly improving the overall system's engineering availability and lifecycle reliability.
[0046] Through the above-mentioned integrated hardware and software design, the integrated intelligent measurement terminal of the present invention is not a simple stacking of functions, but rather a collaborative optimization at multiple levels such as structure, power supply, synchronization, communication and computing. This significantly reduces the system's full life cycle operation and maintenance costs and operational complexity, transforming vehicle outline measurement from an engineering task that relies on meticulous maintenance into a standardized service capability that can be deployed on a large scale and is easy to maintain. This provides a stable, accurate and cost-effective source of basic data and decision-making basis for application scenarios such as highway overload control and intelligent transportation.
[0047] The above is an explanation of the terminal perception layer and edge processing layer configuration provided in this application.
[0048] Furthermore, the collaborative management layer can be implemented based on a cloud-based intelligent management platform. This platform serves as the intelligent and business center of the entire system. It can be flexibly deployed in private clouds, government clouds, or local data centers, depending on the customer's data security requirements and existing IT infrastructure. It preferably adopts a microservices and containerized architecture, integrating functions such as model training and evaluation, federated learning coordination, model repository management, canary releases and rollbacks, device and policy management, and auditing and access control. This platform is responsible for aggregating, processing, and analyzing massive amounts of data from the integrated intelligent measurement terminals across the entire domain, achieving a closed loop from edge perception to cloud intelligence, providing unified decision support and centralized management for vehicle weight enforcement operations.
[0049] The cloud service platform can be functionally divided into several collaborative business modules to jointly achieve centralized management and intelligent empowerment of the front-end integrated intelligent measurement terminal. The platform first configures a data aggregation and governance module to uniformly access structured measurement data, event alarm information, and equipment operating status parameters reported by various terminals. It performs standardized parsing, cleaning, deduplication, and correlation on data from different sources and in different formats, and establishes a standardized data model and index system according to station, lane, vehicle identification, and time label, forming a vehicle outline archive and spatiotemporal trajectory database that can be uniformly called by upper-level business.
[0050] The intelligent analysis and comprehensive judgment module performs secondary verification and correlation analysis on the edge-side preprocessing and preliminary judgment results. On the one hand, it performs trajectory connection and consistency checks on the detection records of the same vehicle at different times and different sections to identify suspicious over-limit behaviors, repeated violation characteristics and abnormal measurement results. On the other hand, it performs batch statistics and mining on multi-site and multi-time period data to realize traffic flow analysis, vehicle type structure analysis, distribution and trend judgment of over-limit behaviors, etc., providing data support for macro governance and policy optimization.
[0051] The model training and intelligent service module provides centralized training, validation, version management, and release capabilities for algorithm models such as visual detection, point cloud segmentation, outline size regression, quality assessment, and confidence fusion. This module can utilize anonymized samples and structured summary data aggregated in the cloud to periodically or on-demand retrain and performance evaluation of the model. It also combines a federated learning coordination mechanism to aggregate local updates from various edge nodes, forming a global model more adapted to real-world traffic enforcement scenarios. The updated model is then distributed to the front-end terminals via a remote upgrade mechanism, enabling continuous evolution of front-end recognition and measurement capabilities.
[0052] The visualization and business interaction module presents underlying data and analysis results intuitively in the form of charts, maps, timelines, etc. through a data dashboard, real-time monitoring view, alarm management interface, and statistical report components. It supports operation and law enforcement personnel to view the operating status of equipment at each site, traffic flow and overload distribution in real time, retrieve the complete evidence chain of a specified vehicle or event, and review, confirm, annotate and process alarm records, realizing the digitalization and full traceability of business processes.
[0053] The unified equipment operation and maintenance and security management module is used to centrally monitor the status of all front-end terminals and edge nodes, manage parameter configuration, manage software and firmware versions, remotely diagnose faults and process maintenance work orders. It also controls access and audits key operations such as user access, parameter changes, model release and evidence retrieval. Combined with security policies and abnormal behavior detection mechanisms, it improves the overall security and compliance of the system, reduces the frequency of manual inspections and operation and maintenance costs, and supports the large-scale deployment and full lifecycle management of intelligent vehicle outline measurement and automatic over-limit alarm systems.
[0054] Based on the collaborative design of the above-mentioned integrated intelligent measurement terminal and cloud-based intelligent management platform, this invention forms a closed-loop system from edge perception, edge intelligent processing, and cloud-based comprehensive analysis and operation and maintenance, which can provide highly reliable, scalable, and sustainably optimized system support for road overload control operations in complex road environments.
[0055] Please see Figure 4 , Figure 4 This is a flowchart illustrating the multimodal fusion vehicle profile measurement method provided in this application.
[0056] A second aspect of the present invention discloses a multimodal fusion vehicle outline measurement method, based on any one of the foregoing multimodal fusion vehicle outline measurement systems, the method comprising: Based on the timing synchronization component, the multimodal sensor component is synchronously triggered to sample the vehicle under test and acquire multimodal acquisition data at the same timestamp; the multimodal acquisition data includes point cloud data acquired by the lidar module and image data acquired by the camera module; Spatiotemporal alignment and delay compensation are performed on multimodal acquisition data to unify the multimodal acquisition data under the same spatiotemporal reference. Multimodal target detection is performed by the edge processing layer, and the multimodal acquired data is matched to construct a multimodal fusion object of the vehicle under test, thereby achieving complementary and unified data features between different modalities; Based on the multimodal fusion object, the three-dimensional outline dimensions of the vehicle under test are estimated; The multimodal fusion object and its 3D outline dimensions are uploaded to the collaborative management layer, where they are archived and studied.
[0057] Based on the aforementioned system, the multimodal fusion vehicle outline measurement method provided in this application establishes a closed-loop process of "synchronous triggering - spatiotemporal alignment - feature association - robust estimation," achieving deep complementarity of multimodal data. By organically combining the geometric accuracy of LiDAR with the semantic discrimination capability of cameras, it overcomes the performance bottleneck of single sensors under extreme conditions such as severe weather, drastic changes in lighting, or target occlusion. This globally collaborative strategy ensures that the system can accurately identify vehicle categories and accurately measure outline dimensions, significantly reducing the risk of false alarms caused by environmental interference and improving the automation and accuracy of overloading enforcement.
[0058] The above is a basic description of the multimodal fusion vehicle outline measurement method provided in this application. The specific implementation methods of each step will be explained below: In one possible implementation, based on a timing synchronization component, the multimodal sensor components are synchronously triggered to sample the vehicle under test, acquiring multimodal data at the same timestamp, including: Hardware timing trigger: Utilize GNSS timing, PPS or PTP hardware timing units to provide a unified sampling trigger signal and high-precision timestamp for each sensor.
[0059] Specifically, on the system side, the integrated intelligent measurement terminal drives sampling with a fixed period Δt. Each sensor module generates a timestamped data frame at the sampling time and writes it to a unified buffer queue for subsequent synchronization and fusion. Hardware timing provides a unified sampling trigger signal and high-precision timestamps for multiple sensors, and each sensor data frame is sampled with a local timestamp for the nth frame. in This is the sensor's own sampling period.
[0060] Furthermore, multi-source sensor data, including images, point clouds, and inertial navigation data, are written into the buffer queue terminal according to a standard frame structure: each sensor frame is encapsulated in a unified structure as follows: in: , }, And record the status word (Frame drops, overheating, synchronization lock status, etc.) are used as quality priors.
[0061] Perform status and integrity checks on sampled frames and generate raw quality labels to support degradation control. Calculate a base quality score for each frame. when Frames that are marked as low-confidence are preferred in the initial stage, and their weight is reduced or a degradation mode is triggered during subsequent fusion.
[0062] By introducing hardware-level timing synchronization using GNSS, PPS, or PTP, the system achieves source-level sampling locking. The hardware triggering mechanism completely eliminates the uncertainty and jitter caused by software polling, ensuring extremely high temporal consistency of data from the moment of generation. Simultaneously, through standardized encapsulation and status recording (such as frame loss and over-temperature), the system can monitor the health of the sensing link in real time, providing transparent physical prior information for subsequent degradation control and data verification, enhancing the integrity of the sensing evidence chain and judicial credibility. Furthermore, through archiving by the collaborative management layer, measurement data and results can be learned to continuously iterate the measurement system and further improve its measurement accuracy.
[0063] In one possible implementation, spatiotemporal alignment and delay compensation are performed on the multimodal acquisition data to unify the multimodal acquisition data under the same spatiotemporal reference, including: Time interpolation: Based on a unified time standard, interpolation reconstruction is performed on multimodal acquisition data to map the multimodal acquisition data to the same reference time.
[0064] Specifically, using the master clock inference trigger time t as a unified alignment reference, all sensor observation data are mapped to the same time axis to generate aligned observations. For arbitrary asynchronous measurement data (Such as image data, point cloud data), the target time-time data is reconstructed using linear interpolation: in For high-frequency moving targets, a higher-order motion model can be selected for interpolation (such as a second-order uniform acceleration model).
[0065] Optionally, the system supports clock model calibration; if sensor clock drift exists, a clock synchronization model is established. in Clock drift coefficient, Clock skew is estimated and compensated online using periodic time synchronization protocols such as PTP / NTP.
[0066] Delay compensation: Estimate the total communication delay of each module in the multimodal sensor assembly, and suppress noise by smoothing filtering to obtain the compensated multimodal acquisition data.
[0067] Furthermore, the system estimates the total delay of each sensor channel online: in The inherent measurement delay of the sensor includes camera exposure time, lidar scanning cycle, etc. Data transmission latency includes system bus transmission and network latency, etc. The preprocessing computation delay mainly includes the delay caused by feature extraction and point cloud filtering.
[0068] Compensate the interpolated data to the alignment time: To suppress delay estimation noise, exponential smoothing is introduced for delay smoothing filtering: in It is a smoothing factor that is adaptively adjusted based on the characteristics of delayed fluctuations.
[0069] Motion correction: Geometric correction is performed on each sampling point using the attitude change information of the vehicle to eliminate data distortion caused by the motion of the vehicle under test.
[0070] For rotating lidar, each point in the point cloud carries its sampling time. The system eliminates point cloud distortion caused by carrier motion through motion compensation. It utilizes changes in carrier pose. (Estimated via IMU or odometry) Correct for each point: in Indicates from Time to Align Rigid body transformation of the carrier coordinate system at any given time.
[0071] To compensate for the impact of older sampling points by assigning time decay weights to the point cloud: in This is the time decay coefficient, typically taken as a fraction of the scan period. This weight is used for subsequent fusion weighting, so that points closer to the alignment time contribute more.
[0072] Define motion compensation residual metric: when At this time, the motion model optimization or recompensation process is triggered.
[0073] Quality detection: Based on the multimodal acquisition data of the previous moment, predict the multimodal acquisition data of the current moment; compare the predicted multimodal acquisition data with the multimodal acquisition data acquired at the current moment. If the difference is greater than a preset threshold, mark the moment as low-quality alignment and use the predicted multimodal acquisition data for subsequent operations.
[0074] Specifically, the alignment effect is verified through multi-source data consistency checks: in These are predictions based on motion models. To measure the covariance of the sensor. If If the threshold is exceeded, the system will revert to the predicted value and mark the frame as low-quality alignment.
[0075] Through the four-level processing flow of time interpolation, delay compensation, motion correction, and quality verification, the system achieves precise alignment of multi-sensor data in time and space, providing a highly consistent input foundation for subsequent fusion sensing.
[0076] In one possible implementation, performing multimodal object detection includes: Perform forward inference based on convolutional neural networks on image data to extract visual feature vectors of the target, including semantic category and encoded appearance, texture information; In the vision branch, the system inputs an image at time t. Perform forward inference based on a convolutional neural network (CNN), the network model is denoted as... The target set output during a single forward propagation: in: For the goal 2D bounding box . For the target category (e.g., cars, trucks, pedestrians, motorcycles). To detect confidence level, which reflects the probability of the target's existence. It is extracted from the target region (RoI) or within the mask. The 3D deep feature vector encodes the appearance, texture, and local structural information of the target, and is the key basis for subsequent data association and target re-identification (Re-ID). These are the optimized parameters for the visual neural network.
[0077] This module preferably adopts a single-stage or two-stage detection architecture that balances accuracy and speed, including but not limited to the YOLO series, EfficientDet variants, or lightweight versions of Mask R-CNN, and undergoes model pruning and quantization to adapt to edge computing resource constraints.
[0078] Perform point cloud neural network processing on point cloud data to extract geometric feature vectors of the target, including its size, shape, and spatial distribution; In the point cloud branch, to fully utilize 3D geometric information and enhance detection capabilities in visually degraded scenarios such as occlusion and low light, the system can optionally deploy a point cloud neural network. Directly on the original or voxelized point cloud The process is performed to output the target proposal in three-dimensional space and its corresponding geometric features: in: A 3D target proposal is typically represented by a 3D bounding box. Or it may be given in the form of a prototype (such as the target center or the foreground cluster). It is a D'-dimensional geometric feature vector extracted from point cloud clusters, which encodes invariant features such as the target's size, shape, spatial distribution, and local point patterns. These are the parameters of the point cloud neural network.
[0079] Preferred network PointNet++, VoxelNet, or advanced architectures based on sparse convolution can be used to maintain stable feature extraction capabilities even under conditions of sparse and unevenly distributed point clouds.
[0080] Semantic discrimination is performed using visual feature vectors, and three-dimensional geometric precision and scale are provided using geometric feature vectors, thereby generating a cross-modal feature set for constructing multimodal fusion objects.
[0081] Specifically, the edge intelligent processing platform of the present invention supports visual networks. With cloud neural networks Parallel processing, simultaneously obtaining results on the edge: Precise 2D localization, semantic category, and high-dimensional visual feature set ; Precise 3D location, scale information, and complementary geometric feature set The two types of outputs mentioned above together form the input basis for subsequent multimodal data association and fusion calculations, realizing multimodal perception with "feature-level complementarity": the visual branch uses convolutional neural networks to deeply mine the texture, category, and appearance information of the target, solving the semantic problem of "distinguishing similar car models"; the point cloud branch uses a dedicated model to extract three-dimensional spatial structural features, ensuring the absolute accuracy of scale measurement. This parallel and clearly defined processing approach enables the system to maintain stable perceptual output through feature complementarity even in visual degradation (such as at night) or sparse point cloud scenarios, significantly improving robustness under various conditions.
[0082] After obtaining the above multimodal feature set, the multimodal acquisition data can be matched.
[0083] In one possible implementation, matching multimodal acquisition data includes: Geometric gated filtering: Point cloud data is projected onto the image plane in the image data, and depth statistical consistency is checked on the point cloud data falling within the detection range by the median absolute deviation to remove background outliers.
[0084] Specifically, firstly, using the external parameters obtained from the aforementioned precise calibration, the current lidar point cloud is... Project onto the image plane. For each visual detection box... Collect all laser points whose projections fall within the frame to form a preliminary candidate point set: in, Point cloud in radar coordinate system, This is its projection position on the image plane.
[0085] Considering that relying solely on projection-based bounding boxes introduces two typical types of interference: first, background points such as the ground, guardrails, and green belts may mistakenly fall into the detection box due to overlapping viewpoints; second, when vehicles in front or behind obstruct the view, the point cloud of the rear vehicle may be incorrectly assigned to the front vehicle. Therefore, this system introduces a consistency gating based on depth statistics on the candidate point set to remove depth outliers. Let the depth component of the candidate point set in the camera coordinate system be vector Z. Robust statistics using the median and median absolute deviation (MAD) are used to identify and filter out depth outliers. in, It is a candidate point set The vector formed by the depth values of all points in the camera coordinate system. This indicates that the median is taken, and the constant 1.4826 makes MAD a consistent estimator of the standard deviation of the normal distribution. The refined point set is formed by retaining points that satisfy the following conditions: threshold Typically, a value of 2.5 to 3.0 is used, corresponding to a confidence interval of approximately 98%-99% under a normal distribution. This step effectively removes background points and misassigned points caused by occlusion, resulting in a high-purity target candidate point cloud, laying the foundation for subsequent accurate geometric attribute calculations and reliable association.
[0086] Association confidence assessment: Using a pre-defined neural network model, the association confidence of geometric feature vectors and visual feature vectors is scored; After completing the initial projection screening and depth geometry gating, the visual instances need to be... Point cloud examples A one-to-one matching relationship is established between them. Traditional matching strategies based on manual metrics such as IoU and center distance are prone to failure when the target is partially occluded, the point cloud is extremely sparse, or the viewpoint difference is large. To address this, this invention introduces a lightweight multilayer perceptron (MLP) to construct a learning-based cross-modal association confidence score. For each pair of candidate associations... Construct feature vectors that include spatial geometric and semantic features: Spatial geometric compatibility characteristics: Visual inspection box Point cloud detection The cross-union ratio (CUI) between two-dimensional bounding boxes projected onto the image plane measures the degree of overlap in two-dimensional space.
[0087] The Euclidean distance between the visually approximate 3D position (obtained by back projection or ground plane assumption) and the 3D center of the point cloud.
[0088] The difference between velocity estimated based on visual optical flow and velocity measured based on point cloud registration or radar.
[0089] Cross-modal semantic similarity features: Apparent feature vectors extracted from the visual branch Geometric feature vectors extracted from point cloud branches The cosine similarity between the two modalities reflects their consistency in the semantic embedding space.
[0090] The above features are concatenated and input into the MLP, and then processed by the Sigmoid function. Mapped to The interval is used to obtain the learning-based association confidence score: Global optimal matching: A cost matrix is constructed based on the association confidence score, and a global optimization algorithm is used to solve for the best matching pair between the image target and the point cloud instance in order to construct the multimodal fusion object of the vehicle under test.
[0091] The aforementioned scores A comprehensive assessment of visual targets was conducted. With lidar target The probability of belonging to the same multimodal fusion object. A cost matrix is constructed based on this. And the Hungarian algorithm is used to solve for the global optimal matching. : This learning-based association strategy can adaptively balance geometric constraints and feature similarity, significantly improving the robustness of cross-modal pairing in occluded, sparse point cloud, and complex background scenarios.
[0092] The above describes the method for matching multimodal acquired data. However, in some scenarios, point cloud overlap may occur.
[0093] Therefore, in one possible implementation, matching the detected multimodal data also includes: Adhesion risk assessment: A Gaussian mixture model is used to fit the distribution pattern of the point cloud data on a plane perpendicular to the direction of travel of the vehicle under test, and the separation index is calculated to identify whether there is a risk of multi-vehicle adhesion.
[0094] Specifically, the internal structure of the point cloud clusters is evaluated from the perspective of lateral distribution. Using the lateral axis perpendicular to the vehicle's main driving direction (or the direction of the principal components of the point cloud) as the analysis benchmark, the distribution of the point cloud along this axis is statistically analyzed, and a Gaussian mixture model (GMM) is used for fitting to obtain the mean and variance of the two principal components. Define the separation index: when Greater than the preset threshold When the current point cloud cluster is determined to have a clear bimodal or multimodal structure in its horizontal distribution, it is a high-probability cluster that is stuck together, and the secondary segmentation process is triggered as follows.
[0095] Secondary segmentation of adhesion: If there is a risk of multiple vehicles being stuck together, the stuck point cloud data is split into multiple independent vehicle point cloud sets by combining the abrupt change locations of the point cloud data, the outline boundary of the vehicle under test identified by the image data, and the speed difference of each point in the point cloud data. Multimodal fusion object reconstruction: The split vehicle point cloud set is used as point cloud data, and multimodal data matching is performed again to construct a multimodal fusion object.
[0096] In the secondary segmentation stage, the system comprehensively utilizes three types of complementary clues for joint decision-making: Geometric segmentation cues: Based on geometric features such as abrupt changes in point cloud density and curvature, three-dimensional clustering algorithms such as Euclidean clustering and region growing are used to achieve natural segmentation in spatial structure; Visual guidance cues: Using the fine pixel-level masks obtained by image instance segmentation, the masks of different vehicles are back-projected into three-dimensional space as strong constraints for dividing the point cloud boundaries of each vehicle. Motion consistency cue: Combining short-time trajectory tracking results, the motion vector field of point clouds within a cluster is analyzed, and regions with significantly inconsistent motion directions or velocities are divided into different sub-clusters.
[0097] The new point cloud subclusters generated by the above secondary segmentation are regarded as candidate objects to be generated, and the association process is re-executed in combination with the unassigned visual features to achieve correct splitting and tracking of multimodal fusion objects.
[0098] Once the multimodal fusion object is established, the three-dimensional outline of the vehicle can be measured.
[0099] In one possible implementation, estimating the three-dimensional outline dimensions of the vehicle under test based on a multimodal fusion object includes: Establish a local coordinate system: assign weights to the point cloud data in the multimodal fusion object, determine the yaw angle of the vehicle under test through the weighted principal component analysis algorithm, and transform the point cloud data into a three-dimensional coordinate system constructed with the length, width and height of the vehicle under test according to the yaw angle; Robust dimensional calculation: The point cloud distribution along each axis of the three-dimensional coordinate system is statistically analyzed, the difference between the high quantile and the low quantile is calculated, and the difference is used as the estimated three-dimensional dimensions of the vehicle under test.
[0100] After obtaining a clean 3D point cloud associated with a specific target, the core task of this module is to robustly and accurately estimate the 3D dimensions (length, width, and height) of the vehicle's outline and its yaw angle on the horizontal plane on a discrete point set. Addressing the common problems in practical engineering such as point cloud noise, outliers, and incomplete observations caused by occlusion, this invention abandons the simple extremum method, which is extremely sensitive to outliers, and constructs a size estimation process that integrates robust statistics and quality assessment.
[0101] First, the vehicle's principal orientation (yaw angle) on the horizontal plane is fundamental for constructing the target's own coordinate system and calculating its length, width, and height. To reduce the interference of outliers on the principal orientation estimation, this invention employs a weighted covariance matrix combined with PCA to obtain a robust principal orientation. (Target point cloud) Points in Assign weights The weights can be set based on the point's reflection intensity, neighborhood density, or the correlation confidence output by the aforementioned modules. Points with low reflection intensity or located at the cluster edge can be assigned smaller weights. The weighted centroid is calculated based on these weights. The corresponding weighted covariance matrix is: For the weighted covariance matrix Perform eigenvalue decomposition and extract the eigenvector corresponding to the largest eigenvalue. As the principal axis of the target in the horizontal plane (XY plane). The target's yaw angle. : Furthermore, rotate all points around the vertical axis (Z-axis). Then, translate it to a local coordinate system with the centroid as the origin to align the point cloud: in, This is the corresponding rotation matrix. After this step, the vehicle's longitudinal direction is aligned with the local coordinate system's X′ axis, and its lateral direction is aligned with the Y′ axis, providing a unified and stable reference system for subsequent size estimation.
[0102] After orientation alignment, directly estimating L / W / H by taking the difference between the maximum and minimum projections on each axis is easily "stretched" or "compressed" by a small amount of noise or residual sticky points. Therefore, this invention uses quantile differences as a size estimator, which is naturally robust to outliers. Specifically, the 99th percentiles of the aligned point cloud in the local X′ (major axis), Y′ (width axis), and Z′ (height axis) coordinates are calculated respectively. With 1% quantile The difference is the length, width, and height of the target: By pruning approximately 1% of the extreme points (quantile pairs in) (Adjust as needed within the interval) effectively suppresses the influence of occasional noise points, residual adhesion points and local abnormal reflections on the size results, making the estimated value closer to the actual physical outline of the vehicle, and significantly reducing inter-frame jitter and occasional jumps.
[0103] Please see Figure 5 , Figure 5 An example diagram illustrating the construction of the bounding box in the multimodal fusion vehicle outline measurement method provided in this application; Furthermore, based on the multimodal fusion object, estimating the three-dimensional outline dimensions of the vehicle under test also includes: Constructing the bounding box: Based on the estimated 3D dimensions, construct the bounding box of the vehicle under test in the 3D coordinate system; Calculate the fitting: Calculate the distance vector from each point in the point cloud data to the bounding box surface, and determine the fitting quality index based on the statistical distribution of the distance vector; Determine measurement reliability: Calculate the reliability of this size estimate based on the fitting quality index, and use it as the basis for judging the over-limit alarm.
[0104] Providing only the size value is insufficient to support subsequent filtering and alarms. This invention further quantifies the fitting quality of the size estimation for this frame, providing downstream modules with directly usable quality metrics. Specifically, in the local coordinate system, using... Construct a 3D directed bounding box (OBB) for the vehicle based on the parameters, and define its extent in the local coordinate system as follows: For each point Calculate the signed distance vector from it to the surface of the bounding box: in, , , To retrieve the maximum value element-wise. When the point is inside the bounding box, When the point is outside the box, Pointing to the nearest box surface, it reflects the degree to which a point deviates from the ideal outline. The residual is defined as: To avoid a single extreme residual dominating the overall evaluation, the system employs residual ensembles. The 95th percentile is used as the quality index for the size fit of this frame: The smaller this metric, the better the match between the point cloud and the fitted bounding box, and the more reliable the size estimation. Combining metrics such as point cloud quantity and spatial coverage, a point cloud quality score can be constructed. This data is then used as the core input for subsequent temporal filtering and alarm engines, adaptively adjusting measurement noise, weights, and thresholds. This automatically corrects the observation bias of individual multimodal fusion objects when point cloud quality is poor, ensuring the statistical stability and measurement reliability of the vehicle outline dimensions output.
[0105] The above is a description of a complete measurement process for this application. Some optional solutions will be provided below as supplementary information.
[0106] In dynamic traffic scenarios, accurate and stable speed estimation is fundamental to advanced functions such as speeding warnings and trajectory prediction. Speed measurements from single sensors each have limitations: visual optical flow depends on texture and is susceptible to illumination interference, while point cloud registration is unstable under sparse conditions. Therefore, this system designs a speed estimation framework based on adaptive Kalman filtering for each multimodal fusion object. By intelligently weighting and fusing observation data with different characteristics, robust three-dimensional speed estimation is achieved.
[0107] In a preferred embodiment, the system computes complementary velocity observations in parallel and evaluates their quality in real time. For visual optical flow observations, sparse feature points are tracked within the image region associated with the target. Pixel displacements are estimated using the Lucas-Kanade optical flow algorithm, and combined with the average depth given by the point cloud, the two-dimensional pixel flow is converted into a three-dimensional velocity vector in the camera coordinate system. Its observational uncertainty The system estimates the mean square error of feature point tracking, depth dispersion, and image signal-to-noise ratio. For point cloud registration observation, an improved ICP algorithm is used to solve for the optimal rigid body transformation of point cloud clusters belonging to the same target in two adjacent frames. The displacement increment is extracted from the transformation matrix and divided by the inter-frame time to obtain the velocity. Its observational uncertainty This is related to registration residuals, point cloud overlap rate, and geometric integrity (whether there are sufficient planar / edge features). Optionally, millimeter-wave radar radial velocity can also be introduced as a third type of observation to enhance robustness in visually degraded scenarios such as rain, fog, and nighttime.
[0108] The above multi-source velocity observations are stacked to form the observation vector at the current moment. The uncertainty of each observation source is modeled as an observation noise covariance matrix: Each of them It is a 3×3 covariance matrix, representing the uncertainty of the corresponding observations in each direction of three-dimensional space. To simplify the calculation, isotropic assumptions can be made, degenerating into scalar variance form: .
[0109] The system employs either an Extended Kalman Filter (EKF) or an Unscented Kalman Filter (UKF) to perform time-series estimation of the target state. The state vector is defined as: .
[0110] in These represent position, velocity, and acceleration, respectively. Based on the target's historical motion characteristics, a uniform velocity (CV), uniform acceleration (CA), or coordinated turn (CT) model is adaptively selected as the state transition function. Process noise covariance Online adjustments to random dynamics, such as amplifying lateral and acceleration noise components when detecting sudden braking or sharp turns, to reflect the increased uncertainty in the motion model.
[0111] Filtering prediction steps: in, It is a state transition function The Jacobian matrix can be approximated by an unscented transformation.
[0112] When multi-source observations are obtained Then, the system first calculates the quality score for each observation source. The score integrates internal sensor metrics (such as the number of optical flow tracking points and ICP residuals) with spatiotemporal consistency metrics (such as deviation from predicted values). Subsequently, the quality score is mapped to a scaling factor of the observation noise covariance matrix to achieve quality-driven adaptive updates. in, As the reference noise level, To prevent small values from being divided by zero, this design ensures that higher quality corresponds to lower observation noise, leading to greater confidence in the Kalman filter during updates; conversely, lower quality results in a automatically reduced weight.
[0113] In the update step, the Kalman gain is first calculated. : It integrates prediction and observation: in, It is the observation function. Its Jacobian matrix. For velocity observations, The velocity component is usually extracted directly from the state.
[0114] Filtered state velocity components The system outputs the current optimal velocity estimate for the target. To balance smoothness and real-time performance, the system can overlay a short-time-window median or slight low-pass filter on the filtered result, while strictly controlling the phase delay to avoid introducing significant lag.
[0115] To enhance robustness under abnormal operating conditions, this module also includes an observation source anomaly management mechanism: when the quality score of an observation source remains below the threshold for an extended period or exhibits systematic deviations from other observation sources or predicted values, it can be temporarily removed from the fusion process. Updates will then rely solely on the remaining high-quality observations and the motion model, while appropriately increasing process noise to reflect additional uncertainties. When the quality of all observations falls below the threshold, the system automatically switches to pure prediction mode, relying solely on the motion model to extrapolate velocity and explicitly amplifying the output covariance. This allows downstream alarm modules to adopt more conservative risk strategies (such as raising the trigger threshold or downgrading to a "suspected overspeeding" flag).
[0116] Ultimately, this module not only outputs the optimal velocity estimate It also provides its estimated covariance. (Right now (The part corresponding to the speed). This uncertainty information is directly passed to the subsequent confidence management and alarm decision-making modules to construct speed safety margins and probabilistic judgment conditions, ensuring that overspeed alarms for multimodal fusion objects have clear judgment criteria and statistical credibility under complex operating conditions.
[0117] In complex traffic scenarios, single-frame observations inevitably contain noise and uncertainty. To avoid false alarms and missed alarms caused by instantaneous anomalies, this invention introduces a systematic quality assessment, confidence fusion, and uncertainty management mechanism at the target level, providing "credible evidence" for each measurement result. This ensures that alarms are no longer based on a single threshold comparison, but on intelligent judgment based on statistical significance and reliability constraints.
[0118] In one implementation, the system evaluates target observations from three complementary dimensions, obtaining point cloud quality scores, visual quality scores, and cross-modal alignment quality scores. First, the geometric quality of the point cloud is scored, as the quality of the point cloud directly determines the reliability of the 3D size estimation. Point cloud quality scoring. It combines the number of point clouds, coverage, and goodness of fit: in: It is the number of valid point clouds associated with the target. This reflects the concept of "diminishing marginal utility". It refers to the spatial coverage of the point cloud over the target's 3D bounding box. For example, it can be described by the ratio of the point cloud's span in each axis to the estimated size, or by the point projection ratio of each face. Low coverage usually means that the target is partially occluded. This is the residual statistic from the point cloud to the fitted bounding box, measuring the degree of fit between the point cloud and the regular cube model. Excessive residuals may be caused by point cloud noise, irregular target shapes (such as protruding cargo), or adhesion. It's the Sigmoid function, which maps linear combinations to the interval (0,1). Coefficients Offline data calibration was used to balance the contributions of each indicator. The larger the value, the more complete the point cloud, the better the fit, and the more reliable the size estimate.
[0119] Visual inspection quality scores can be directly output by the detection network as category confidence scores. Provided, and if necessary, additional factors such as mask quality, edge sharpness, and occlusion level can be added for weighting. Higher... This means that the target has a clear outline and semantic definition in the visual field, making the determination of the target's existence and category more reliable.
[0120] Cross-modal alignment quality score: Cross-modal alignment quality is used to evaluate the consistency between visual and lidar observations and is a key indicator for verifying data correlation and calibration stability. It can be defined as: in, It is the average distance (in pixels) between the associated point cloud projected onto the image plane and the boundary of the visual detection box or instance segmentation mask. Parameter Sensitivity to errors has been controlled. Low alignment quality may indicate miscorrelation, extrinsic drift, or severe occlusion.
[0121] To obtain a global reliability metric that can directly drive subsequent decisions, this system integrates the above three types of quality scores into a target-level comprehensive confidence score. In the preferred embodiment, a geometrically weighted form (product structure) is used to reflect the weakest link effect; if the quality of any dimension is too low, the overall credibility will be significantly reduced. Among them, the index weight Used to adjust the importance of different dimensions, it is usually set to 1, indicating equal weight. In practice, the system supports automatic adjustment based on the acquisition scenario; for example, in nighttime or rainy / foggy weather, the visual weight can be appropriately reduced. . The function ensures that the output value is within a reasonable probability range.
[0122] The overall confidence level serves as a unified gating signal for subsequent modules. On one hand, in target tracking filtering, it acts as a state update gating signal; only... Above the threshold Only observations with a value of 0.3 are used to update the target's state (position, size, velocity); otherwise, the filter only performs predictions to prevent low-quality observations from contaminating the stable trajectory. On the other hand, as an alarm triggering gate, the triggering of any service alarm (overwidth, overheight, etc.) requires, in addition to satisfying the physical quantity exceeding the limit, also to satisfy... Higher than the preset confidence threshold (e.g., 0.5) to suppress false alarms caused by single-frame false detections or noise peaks from the source.
[0123] For key decision-making quantities such as external dimensions and speed, this system further explicitly models measurement uncertainties and introduces a safety margin into alarm decisions, ensuring that out-of-limit conclusions have clear statistical confidence significance and achieving robust alarm decisions. Taking width measurement as an example, in one implementation, width uncertainty... It can be obtained by any or a combination of the following methods: Analytical method: Calculation based on the error propagation formula of the point cloud distribution and the fitted model; Statistical method: Variance estimation of multi-frame samples based on local time windows; Learning methods: based on quality features such as point cloud quality Coverage Predictions The learning-based error regression model.
[0124] Based on this, a safety margin for width determination is defined: in, The quantiles of the standard normal distribution are used to cover random measurement errors. This is the systematic error margin, used to cover long-term deviations such as calibration residuals and mechanical deformation, and is calibrated by long-term statistics.
[0125] In determining the excessive width, this invention employs a conservative strategy: an alarm is triggered only if the conservative value, after deducting uncertainties, still exceeds the limit. Width alarm triggering conditions: That is: if the measured value, after subtracting its uncertainty margin, still exceeds the regulatory threshold... Furthermore, the system only recognizes a credible overwidth event as having occurred if the overall confidence level is sufficiently high. Similarly, separate parameters can be constructed for height, length, and speed. This achieves a unified determination of uncertainty constraints. Similarly, for velocity estimation, the velocity covariance matrix output by the Kalman filter is used. Obtain the standard deviation of each velocity component Build a safety margin for speed to ensure robust judgment in overspeed warnings.
[0126] In a preferred embodiment, this module outputs a structured quality and uncertainty description vector for each multimodal fusion object, including but not limited to the overall confidence level. Quality decomposition indicators Used for system debugging, health monitoring, and online calibration and diagnosis; standard deviation of key physical quantities Safety margin corresponding to the dimension The information includes the confidence level parameters used. This information is directly used by the subsequent time-series filtering and alarm rule engine, enabling the system to automatically reduce the impact of single-frame observations or downgrade them to suspected / pending alarms when faced with sensor degradation, partial obstruction, or environmental deterioration, rather than outputting hard alarms lacking evidence. Therefore, this invention achieves evidence-based, explainable, and quantifiable alarm decisions in engineering deployments, providing crucial support for building a highly reliable and verifiable off-site enforcement system for road traffic enforcement.
[0127] Based on the aforementioned high-quality measurement results, quantified uncertainty, and comprehensive confidence level, a joint probabilistic judgment, hysteresis state machine, and multi-frame confirmation alarm decision mechanism is constructed, so that the system output is not an immediate response to single-frame noise, but a stable and auditable violation event.
[0128] Taking width as an example, the system treats it as a random variable with mean and variance, rather than a single scalar, including the optimal estimate: (Usually the filtered estimate), the uncertainty estimate output from the previous step. At a given regulatory threshold When the actual width of the current vehicle exceeds the threshold, the probability is calculated as follows: in This is the cumulative distribution function of the standard normal distribution. This probability considers both the magnitude of the excess and the uncertainty of the estimate: the greater the excess and the smaller the variance, the higher the probability of exceeding the limit; when uncertainty is high, even if the measured value is slightly higher than the threshold, the probability of exceeding the limit will be suppressed, making the decision more cautious. Similarly, separate functions can be constructed for indicators such as vehicle height, length, and speed. The equal probability form provides a unified probability basis for multiple types of alarms.
[0129] To avoid alarm jitter caused by single-frame noise or brief changes in target attitude, the system maintains a hysteresis state machine for each target alarm type. The input is the probability of consecutive frames exceeding limits. The output is a stable alarm status. The state transition rule is as follows: Among them, the trigger threshold This ensures that the alarm state is only activated when the confidence level is high. Release threshold. In satisfying Formation of probability hysteresis interval To avoid repeatedly switching around the threshold, trigger confirmation frame count. and release confirmation frame number These represent the number of consecutive frames required for triggering and releasing the alarm, used to suppress single-frame spikes, ensure alarm continuity, avoid flickering, and guarantee sufficient "acknowledgment" and "hold" time. This state machine can be applied to all alarm types, including over-width, over-height, over-length, and speeding restricted vehicles, using the same framework. It supports configuring corresponding probability parameters and frame number thresholds for specific services.
[0130] Before proceeding with probability and hysteresis determination, the system first employs a comprehensive confidence level derived from the preceding steps. Quality gating is performed only if the following conditions are met: Current frame observations only participate The calculation and state machine update process is as follows: When the confidence level is below a threshold, the observation in this frame is only used for internal diagnostics and does not drive alarm state changes, thereby automatically suppressing false alarms in cases of severe occlusion, sensor malfunction, or calibration mismatch. This mechanism can automatically suppress erroneous alarms in cases of severe occlusion, sensor malfunction, or calibration mismatch.
[0131] When the state machine detects a transition from 0 to 1 in an alarm (i.e., the alarm event is officially triggered), the system generates a standardized event evidence package at the edge, along with corresponding structured event records and evidence identifiers. The specific evidence content and chain management are uniformly described by the subsequent evidence chain mechanism. This design elevates alarm output from "frame-level transient" to "event-level result," facilitating event-based handling and auditing by upper-level law enforcement systems.
[0132] Therefore, this invention upgrades the vehicle outline over-limit alarm from a traditional hard threshold trigger to an intelligent decision-making module with probabilistic reasoning and self-evaluation capabilities, providing more reliable and verifiable technical support for non-site enforcement of road over-limit control and intelligent traffic supervision.
[0133] To achieve continuous system optimization and compliance auditing capabilities, this invention further designs a complete evidence chain generation mechanism and a cloud-based collaborative learning framework. The system performs real-time sensing and alerting at the edge and continuously evolves from operational data through a closed-loop learning mechanism. This ensures that alert results are traceable, verifiable, replayable, and rollbackable without compromising real-time performance and stability, and supports the continuous evolution of models and policies. Specifically, this includes: End-side generation and solidification of the chain of evidence To ensure the traceability of alarms and the relevance of model optimization, when preset trigger conditions are met, the terminal generates a structured evidence package and locally stores it. Trigger conditions include, but are not limited to: When an alarm event is triggered, the status of any alarm type, such as overwidth, overheight, or overspeed, changes from 0 to 1. Difficult samples trigger the following issues: low overall confidence, high output entropy, hovering around the threshold, and inconsistency across models. System anomalies are triggered, such as suspected drift in external parameters, sensor malfunctions, low long-term consistency, and clock synchronization anomalies.
[0134] The evidence package uses a hierarchical structure and includes at least: Raw / Summary Data Layer: Key raw data or its fingerprints within the window before and after the alarm (such as images / point clouds / radar frames, key frame sequences, summary features, ROI cropping, etc.). Feature and intermediate result layer: spatiotemporal alignment results, filtering / fusion results, target tracking trajectory, reprojection error, residual statistics, quality labels, etc. Decision output layer: alarm type, threshold, confidence level, risk level, and explanatory information (such as contribution / rule hit / key evidence fragments); Context and configuration layer: Device ID, timestamp, geographic / attitude information, sensor health status, calibration version, model version, parameter version, software version, operating load, etc. Integrity and tamper-proof layer: Generate hash digests for each layer of data in the evidence package and construct a Merkle root or hash chain; at the same time, add a trusted timestamp and a device-side signature / certificate chain to form a verifiable integrity proof.
[0135] The evidence solidification method in this invention includes: writing the evidence package into a local evidence library on the terminal side that can only be added and not modified; forming a chain structure in the manner of "event index-evidence package hash-precedence hash" to ensure that the evidence is traceable and difficult to forge; and performing hierarchical compression, desensitization and encryption on the evidence package according to the network and storage status, supporting offline caching and retransmission after recovery.
[0136] Data upload, verification, and archiving management.
[0137] Edge terminals prioritize and adaptively upload evidence packets based on bandwidth: difficult samples and high-risk alarms are uploaded first; other evidence can be uploaded in batches with delays or only fingerprint digests are uploaded.
[0138] Upon receiving the evidence package, the cloud platform first performs rigorous integrity and authenticity verification, including: verifying certificates and signatures, recalculating hashes / Merkle roots, and comparing chained preceding hashes to determine if there has been any tampering or omissions. Only evidence packages that pass verification are then archived and indexed according to metadata such as site, time, vehicle model, and event type, and written to both a hot storage area for rapid retrieval and a cold storage area for long-term retention, while simultaneously generating non-repudiable audit logs. Through this mechanism, it can be ensured that any alarm event can be accurately reproduced and interpreted in the cloud under the same model version and parameter configuration environment, providing reliable technical support for subsequent review, arbitration, and judicial evidence collection.
[0139] Building upon evidence archiving, the cloud platform further unifies the governance of evidence packages, constructing a high-quality sample library for training and collaborative learning. This sample library covers at least four main types of samples: alarm samples from real-world business scenarios, low-confidence, difficult samples reflecting the limits of algorithm capabilities, false alarm / missed alarm samples for targeted correction, and environmental drift samples characterizing distribution drift due to changes in lighting, weather, vehicle type, etc.
[0140] Building upon this foundation, the cloud sequentially performs sample deduplication, quality scoring and grading management, automatic generation of weak labels based on rules / models, and manual sampling and verification of key samples to ensure that the data entering the training process meets engineering requirements in terms of accuracy, consistency, and compliance. For typical weak scenarios (such as nighttime, rain and fog, point cloud adhesion, and special vehicle models), specialized reinforcement subsets can be constructed according to labels, increasing their sampling frequency or loss weights during training to achieve targeted enhancement.
[0141] Cloud-based collaborative learning and strategy generation.
[0142] Based on the processed sample data, the cloud can flexibly select or combine between federated learning mode and centralized training mode according to the deployment environment and data compliance requirements: In federated learning mode, training tasks and hyperparameter configurations are distributed from the cloud to each edge client. The edge clients use local evidence data to fine-tune the model and generate parameter updates at a relatively small learning rate. The cloud side uses a secure aggregation protocol to collect updates from all clients, synchronously injecting differential privacy noise, and simultaneously filtering clients and removing abnormal updates based on update norm, directional similarity, and performance contribution. The client selection strategy comprehensively considers data quality, device and link stability (quality awareness), geographical location, road network, time, and weather diversity (diversity constraints), thereby obtaining a robust global model while ensuring privacy and security.
[0143] In the centralized training mode, the cloud only conducts centralized training based on the original data that has passed the de-identification and compliance verification, or the synthetic data reconstructed from the evidence fingerprint. Optionally, based on the reinforcement sample set built on typical weak working conditions and key business scenarios, the sampling frequency or loss weight of relevant samples is increased during the training process, and higher weights are given to the reinforcement sample set corresponding to long-tail scenarios such as night, rain and fog, point cloud adhesion, and special vehicle models, so as to systematically improve the robustness and generalization ability of the model under key working conditions.
[0144] End-to-end closed-loop distribution and gray-scale rollback.
[0145] Please see Figure 6 , Figure 6 The maintenance architecture diagram of the neural network model in the multimodal fusion vehicle outline measurement method provided in this application.
[0146] like Figure 6 As shown, based on the aforementioned learning results, the lightweight multimodal neural network can be maintained, and its recognition accuracy can be maintained and improved through version updates.
[0147] To avoid introducing systemic risks when deploying new models, the cloud performs a rigorous evaluation and distribution process after each round of global model updates: First, an offline evaluation is conducted. The new model is then tested on a reserved validation set in the cloud, covering core indicators such as detection accuracy, size estimation error, and alarm accuracy. Versions that do not meet the standards are directly eliminated.
[0148] After approval, a small-scale, canary release will be conducted, prioritizing the deployment of the new version to a limited number of representative sites or lane edge nodes for A / B comparison. This allows the new model to run in the same scenario as the current production model and compare actual performance. During the canary release, continuous monitoring of metrics will be performed, with a focus on false alarm rate, network latency, and edge computing power usage. Any abnormal degradation will trigger an automatic rollback.
[0149] All distribution, activation, and rollback actions are written into the audit chain, forming a closed loop with the corresponding evidence package, model version, and parameter version, ensuring that the entire process is auditable and traceable.
[0150] The model distribution process is accompanied by a digital signature and version management mechanism, specifically including: All distributed model packages carry digital signatures, and the edge devices are only allowed to update and load after the signature validity verification is completed; The update adopts a gradual strategy, first deploying it in a canary manner on a few sites, and then gradually expanding the coverage after verifying stability; To address performance fluctuations after deployment, the edge will retain the most recent N historical versions. If a new model's metrics deteriorate, it can be quickly rolled back to a stable version, thereby ensuring the long-term security and controllability of the system.
[0151] Please see Figure 7 , Figure 7 This is a schematic diagram of the visual monitoring and interactive interface provided in this application.
[0152] like Figure 7 As shown, by way of example, the present invention also provides a visual monitoring and interactive interface for road overload control operations, which presents measurement results, alarm events, evidence chains and model version status in a visual manner in the form of charts, maps, timelines, etc. It supports querying and searching by station, lane, time, vehicle type and other dimensions, viewing the complete evidence chain and decision path of any event, and completing alarm review, confirmation, annotation and circulation, so as to realize the integrated closed-loop management of perception, decision-making, evidence collection and optimization.
[0153] Preferably, the visual monitoring and interactive interface has the following functional modules: The on-site video area is used to display the roadside monitoring scene and the currently inspected vehicles in real time, and overlay information such as vehicle detection frames, lane lines, and over-limit status signs, so that operators can intuitively perceive the on-site traffic conditions and measurement results. The structured event list area records each passing vehicle and its corresponding detection event, including at least the following fields: passage time, license plate number, vehicle type, measured length, width and height, type of over-limit, amount of over-limit, confidence level and processing status. It supports searching and sorting by license plate number, time, station and other conditions. The event-level evidence preview area displays key images of the selected vehicle (such as close-up of the vehicle and close-up of the rear), close-up of the license plate, and the corresponding license plate recognition results. It can also play video clips of the vehicle passing through, realizing the integrated display of panoramic images, structured data, and evidence screenshots. The statistics and operational status dashboard area displays key indicators such as vehicle traffic flow, number of overloaded vehicles, distribution of various types of overloads, pass rate, and equipment online rate in the current time period through a combination of charts and numbers. It can also be filtered and compared by dimensions such as station, lane, and time interval to support macro-level situation assessment and operation and maintenance decisions.
[0154] It should be noted that the embodiments of the present invention have better implementability and are not intended to limit the present invention in any way. Any person skilled in the art may use the above-disclosed technical content to change or modify it into equivalent effective embodiments. However, any modifications or equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.
Claims
1. A multimodal fusion vehicle outline measurement system, characterized in that, include: Terminal perception layer; The terminal perception layer includes a multimodal sensor assembly and a timing synchronization assembly; the multimodal sensor assembly includes a camera module and a lidar module, which measure the outline of the vehicle under test through various measurement methods; the timing synchronization assembly includes a time synchronization module, which is used to synchronize the system time of each module in the multimodal sensor assembly; An edge processing layer; the edge processing layer includes computing units deployed with lightweight multimodal neural networks and is communicatively connected to the terminal perception layer; it is used to process the information collected by the terminal perception layer and obtain preliminary measurement results; Collaborative management; The collaborative processing layer is communicatively connected to both the terminal perception layer and the edge processing layer to verify the preliminary measurement results; and learns based on the processing results of the edge processing layer, and maintains the lightweight multimodal neural network according to the learning results.
2. The multimodal fusion vehicle outline measurement system as described in claim 1, characterized in that, The time synchronization component includes at least one of the following modules: Attitude reference module; the attitude reference module includes an inertial measurement unit, used to acquire the attitude information of the terminal perception layer in real time, and provide attitude compensation parameters for the edge processing layer based on the attitude information; Spatiotemporal reference module; Used to provide a unified time and position reference for each module in the multimodal sensor assembly; A time synchronization module is used to distribute the high-precision time signal or local time signal output by the spatiotemporal reference module to the multimodal sensor assembly and the edge processing layer to achieve time alignment between the multimodal sensor assembly and the edge processing layer.
3. A multimodal fusion method for measuring vehicle outline, characterized in that, The multimodal fusion vehicle profile measurement system based on any one of claims 1-2 includes: Based on the timing synchronization component, the multimodal sensor component is synchronously triggered to sample the vehicle under test and acquire multimodal acquisition data at the same timestamp; the multimodal acquisition data includes point cloud data acquired by the lidar module and image data acquired by the camera module; Spatiotemporal alignment and delay compensation are performed on the multimodal acquisition data to unify the multimodal acquisition data under the same spatiotemporal reference. Multimodal target detection is performed by the edge processing layer, and the multimodal acquired data is matched to construct a multimodal fusion object of the vehicle under test, thereby achieving complementary and unified data features between different modalities; Based on the multimodal fusion object, the three-dimensional outline dimensions of the vehicle under test are estimated; The multimodal fusion object and the three-dimensional outline dimensions are uploaded to the collaborative management layer for archiving and learning.
4. The multimodal fusion vehicle outline measurement method as described in claim 3, characterized in that, The timing synchronization component, which synchronously triggers the multimodal sensor component to sample the vehicle under test and acquires multimodal data at the same timestamp, includes: Hardware timing trigger: Utilize GNSS timing, PPS or PTP hardware timing units to provide a unified sampling trigger signal and high-precision timestamp for each sensor; Standardized encapsulation: The multimodal acquisition data is written into a unified buffer queue according to a standard frame structure, and the status, including frame loss and over-temperature, is recorded.
5. The multimodal fusion vehicle outline measurement method as described in claim 3, characterized in that, The step of performing spatiotemporal alignment and delay compensation on the multimodal acquisition data to unify the multimodal acquisition data under the same spatiotemporal reference includes: Time interpolation: Based on a unified time standard, the multimodal acquisition data is interpolated and reconstructed to map the multimodal acquisition data to the same reference time. Delay compensation: Estimate the total communication delay of each module in the multimodal sensor assembly, and suppress noise by smoothing filtering on the total delay to obtain compensated multimodal acquisition data; Motion correction: Geometric correction is performed on each sampling point using the attitude change information of the carrier to eliminate data distortion caused by the motion of the vehicle under test; Quality detection: Based on the multimodal acquisition data of the previous moment, predict the multimodal acquisition data of the current moment; compare the predicted multimodal acquisition data with the multimodal acquisition data acquired at the current moment. If the difference is greater than a preset threshold, mark the moment as low-quality alignment and use the predicted multimodal acquisition data for subsequent operations.
6. The multimodal fusion vehicle outline measurement method as described in claim 4, characterized in that, The multimodal target detection includes: Perform forward inference based on a convolutional neural network on the image data to extract visual feature vectors of the target, including semantic category, encoded appearance, and texture information; Point cloud neural network processing is performed on the point cloud data to extract geometric feature vectors of the target, including the target size, shape, and spatial distribution; The visual feature vectors are used for semantic discrimination, and the geometric feature vectors are used to provide three-dimensional geometric accuracy and scale, so as to obtain the multimodal feature set of the vehicle under test.
7. The multimodal fusion vehicle outline measurement method as described in claim 6, characterized in that, The matching of the multimodal acquired data includes: Geometric gated filtering: The point cloud data is projected onto the image plane in the image data, and the depth statistical consistency check is performed on the point cloud data falling within the detection range by the median absolute deviation to remove background outliers; Association confidence assessment: Using a pre-defined neural network model, the association confidence of the geometric feature vector and the visual feature vector is scored; Global optimal matching: Based on the score of the association confidence, a cost matrix is constructed, and a global optimization algorithm is used to solve for the best matching pair between the image target and the point cloud instance to construct the multimodal fusion object of the vehicle under test.
8. The multimodal fusion vehicle outline measurement method as described in claim 7, characterized in that, The matching of the detected multimodal data also includes: Adhesion risk assessment: A Gaussian mixture model is used to fit the distribution pattern of the point cloud data on a plane perpendicular to the forward direction of the vehicle under test, and the separation index is calculated to identify whether there is a risk of multi-vehicle adhesion. Secondary segmentation of adhesion: If there is a risk of multiple vehicles being stuck together, the stuck point cloud data is split into multiple independent vehicle point cloud sets by combining the abrupt change position of the point cloud data, the outline boundary of the vehicle under test identified by the image data, and the speed difference of each point in the point cloud data. Multimodal fusion object reconstruction: The split vehicle point cloud set is used as point cloud data, and multimodal data matching is performed again to construct a multimodal fusion object.
9. The multimodal fusion vehicle outline measurement method as described in claim 3, characterized in that, The estimation of the three-dimensional outline dimensions of the vehicle under test based on the multimodal fusion object includes: Establish a local coordinate system: assign weights to the point cloud data in the multimodal fusion object, determine the yaw angle of the vehicle under test through a weighted principal component analysis algorithm, and transform the point cloud data into a three-dimensional coordinate system constructed with the length, width and height directions of the vehicle under test according to the yaw angle; Robust size calculation: The point cloud distribution along each axis of the three-dimensional coordinate system is statistically analyzed, the difference between the high quantile and the low quantile is calculated, and the difference is used as the estimated three-dimensional size of the vehicle under test.
10. The multimodal fusion vehicle outline measurement method as described in claim 9, characterized in that, The estimation of the three-dimensional outline dimensions of the vehicle under test based on the multimodal fusion object further includes: Constructing the bounding box: Based on the estimated three-dimensional dimensions, construct the bounding box of the vehicle under test in the three-dimensional coordinate system; Calculate the fitting: Calculate the distance vector from each point in the point cloud data to the bounding box surface, and determine the fitting quality index based on the statistical distribution of the distance vector; Determine measurement reliability: Calculate the reliability of this size estimate based on the fitted quality index, and use it as the basis for judging the over-limit alarm.