A traffic data intelligent collection method and system based on Beidou multimodal fusion
By monitoring the changes in Beidou positioning status and high-precision map layer height, dynamically selecting radar or video data as the dominant benchmark, and generating space-time compensation vectors, the problem of multi-source space-time benchmark inaccuracy caused by Beidou signal obstruction is solved, and the reliability and accuracy of traffic data collection are improved.
Patent Information
- Application Number
- CN202511055733.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-30
AI Technical Summary
In dynamic and complex traffic scenarios, existing technologies experience Beidou satellite signal obstruction, leading to degradation of the positioning mode and inaccurate multi-source space-time references, resulting in reduced reliability of traffic status perception, which is particularly prominent in areas such as interchange hubs and urban canyons.
By monitoring the positioning status changes of Beidou positioning data, triggering the benchmark calibration command, using high-precision maps to obtain layer heights, activating the high-confidence compensation mode, selecting radar point cloud data or video detection frame data as the dominant benchmark, generating spatiotemporal compensation vectors, calibrating spatiotemporal coordinates, and achieving accurate association of multi-source data.
It improves the spatiotemporal alignment accuracy of multi-source perception data, suppresses the benchmark drift caused by satellite signal shielding, ensures the high reliability of traffic event flow in complex urban environments, and solves the trajectory splitting problem caused by multi-source spatiotemporal benchmark inaccuracy.
Smart Images

Figure CN120559693B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic system data fusion, and more specifically, to a method and system for intelligently collecting traffic data based on Beidou multimodal fusion. Background Art
[0002] Current intelligent transportation systems widely utilize multi-source sensing technologies for data collection. The Beidou satellite navigation system, with its high-precision positioning and 24 / 7 service capabilities, has become a core tool for vehicle trajectory monitoring and traffic flow analysis. To enhance data coverage, existing technologies integrate heterogeneous sensors such as video surveillance and millimeter-wave radar to build collaborative roadside sensing systems. At the data processing level, the industry is widely promoting standardized data interface specifications and enabling real-time aggregation and cleaning of multi-source traffic data based on distributed platforms.
[0003] Existing technologies suffer from the defect of multi-source spatiotemporal reference inaccuracy in dynamic and complex traffic scenarios. Beidou signals are easily blocked by the environment, which triggers positioning mode degradation. Sensors such as vision and radar produce dynamic drift of spatiotemporal references during the multimodal data fusion process due to their own time delays and coordinate system differences. As a result, the Beidou trajectory, video detection frame and radar point cloud data of the same target cannot be accurately associated in the spatiotemporal dimension, causing target position deviation and trajectory splitting, significantly reducing the reliability of traffic status perception. This problem is particularly prominent in areas with unstable satellite signals such as interchange hubs and urban canyons. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method and system for intelligent collection of traffic data based on Beidou multimodal fusion to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for intelligently collecting traffic data based on Beidou multimodal fusion, comprising:
[0007] S1. Obtain Beidou positioning data, radar point cloud data, and video detection frame data of the target object;
[0008] S2. Monitor the positioning status change information of Beidou positioning data, and trigger the benchmark calibration instruction when the Beidou signal is degraded from a fixed solution to a floating solution;
[0009] S3. When the benchmark calibration command is triggered, the layer height of the target object is extracted from the high-precision map;
[0010] S4. When the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold, the high confidence compensation mode is activated;
[0011] S5: When the high confidence compensation mode is activated, the radar point cloud data is selected as the dominant benchmark; when it is not activated, the video detection frame data is selected as the dominant benchmark; based on the preset mapping relationship between the layer height and the compensation boundary, the dynamic compensation boundary coefficient is generated to constrain the spatial fusion range of the dominant benchmark and generate the spatiotemporal compensation vector;
[0012] S6. Superimpose the spatiotemporal compensation vector onto the positioning data in the reduced-order stage to generate calibrated spatiotemporal coordinates, associate the radar point cloud data with the video detection frame data, and convert it into a standardized traffic event stream output.
[0013] Furthermore, the Beidou positioning data, radar point cloud data, and video detection frame data of the target object are obtained, including:
[0014] Collect Beidou positioning data of target objects in real time through Beidou positioning terminals;
[0015] Scan the target object with millimeter-wave radar to generate radar point cloud data and record the timestamp of the radar point cloud data;
[0016] Capture the video stream of the target object through the video surveillance device, extract the video detection frame data based on the target detection model, and record the timestamp of the video detection frame data;
[0017] Align the timestamps of Beidou positioning data, radar point cloud data, and video detection frame data to the same time base.
[0018] Furthermore, the positioning state change information of the Beidou positioning data is monitored, and when the Beidou signal is degraded from a fixed solution to a floating solution, a reference calibration instruction is triggered, including:
[0019] Real-time analysis of positioning status words from Beidou positioning data;
[0020] When the positioning status word indicates that the positioning mode changes from fixed solution to floating solution, the state change time point is recorded as the starting time of the reduction order;
[0021] Continuously monitor the changes in the positioning status word of Beidou positioning data based on the start time of the reduction order;
[0022] When the floating solution state lasts until the current moment and the length of time reaches the preset trigger threshold, a reference calibration instruction is generated;
[0023] The generated reference calibration instruction is associated with the start time of the order reduction and stored.
[0024] Furthermore, when the benchmark calibration instruction is triggered, the layer height of the target object is extracted from the high-precision map, including:
[0025] Obtaining the reduced-order start time associated with the reference calibration instruction;
[0026] Extract the target object position coordinates corresponding to the start time of order reduction from Beidou positioning data;
[0027] Convert the target object's position coordinates into matching coordinates in the high-precision map coordinate system;
[0028] Query the layer height attribute field of the HD map based on the matching coordinates;
[0029] Extract the value of the layer height attribute field as the layer height where the target object is located.
[0030] Furthermore, when the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold, the high confidence compensation mode is activated, including:
[0031] Get the target object's layer height and the start time of order reduction;
[0032] Calculate the duration of the scale reduction based on the current time and the start time of the scale reduction;
[0033] Compare the floor height with the preset floor height threshold; compare the continuous downgrade time with the preset time limit threshold;
[0034] When the layer height is greater than the layer height threshold and the continuous order reduction time exceeds the time limit threshold, a high confidence compensation mode activation instruction is generated;
[0035] When the floor height is less than or equal to the floor height threshold or the continuous order reduction time does not exceed the time limit threshold, the high confidence compensation mode is maintained in the inactive state.
[0036] Furthermore, when the high-confidence compensation mode is activated, the radar point cloud data is selected as the dominant benchmark, and when it is not activated, the video detection frame data is selected as the dominant benchmark. Based on the preset mapping relationship between the layer height and the compensation boundary, a dynamic compensation boundary coefficient is generated to constrain the spatial fusion range of the dominant benchmark, and a spatiotemporal compensation vector is generated, including:
[0037] According to the high confidence compensation mode activation state:
[0038] When the high confidence compensation mode is activated, the radar point cloud data is selected as the dominant benchmark based on the low failure probability characteristics of the radar sensor in a high-rise environment;
[0039] When the high confidence compensation mode is not activated, the video detection frame data is selected as the dominant benchmark based on the motion continuity characteristics of the video sensor in the low-level environment;
[0040] Obtain the floor height of the target object as a characterization parameter of the environmental interference entropy value;
[0041] Calculate the dynamic compensation boundary coefficient based on the preset mapping relationship between the floor height and the compensation boundary;
[0042] Use dynamic compensation boundary coefficients to limit the position correction amplitude of the dominant reference during the spatial fusion process;
[0043] A spatiotemporal compensation vector is generated based on the deviation between the dominant reference and the Beidou positioning data after limiting the position correction amplitude.
[0044] Furthermore, as the floor height increases, the corresponding environmental interference entropy value increases, and the dynamic compensation boundary coefficient increases synchronously.
[0045] Furthermore, the spatiotemporal compensation vector is superimposed on the positioning data in the order reduction stage to generate calibrated spatiotemporal coordinates. The radar point cloud data and the video detection frame data are then correlated and converted into a standardized traffic event stream output, including:
[0046] Obtain Beidou positioning data and time-space compensation vectors in the reduced-order stage;
[0047] The time-space compensation vector is superimposed on the BeiDou positioning data in the reduced-order stage to generate the calibrated time-space coordinates;
[0048] Associating the calibrated spatiotemporal coordinates with the radar point cloud data;
[0049] Align the calibrated spatiotemporal coordinates with the video detection frame data by timestamp;
[0050] Extract target size attributes from the associated radar point cloud data, and extract target motion trajectory attributes from the video detection frame data;
[0051] The target size attribute and target motion trajectory attribute are fused to generate a standardized traffic event stream containing event type, timestamp, and spatial coordinates.
[0052] On the other hand, the present invention provides a traffic data intelligent collection system based on Beidou multimodal fusion, comprising:
[0053] Multi-source perception module, used to obtain Beidou positioning data, radar point cloud data and video detection frame data of the target object;
[0054] The downgrade monitoring module is used to monitor the positioning status change information of Beidou positioning data, and triggers the benchmark calibration instruction when the Beidou signal is downgraded from a fixed solution to a floating solution;
[0055] The floor height extraction module is used to extract the floor height of the target object from the high-precision map when the benchmark calibration instruction is triggered;
[0056] A mode activation module is used to activate the high confidence compensation mode when the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold;
[0057] The benchmark compensation module selects radar point cloud data as the dominant benchmark when the high-confidence compensation mode is activated, and selects video detection frame data as the dominant benchmark when it is not activated. It generates dynamic compensation boundary coefficients based on the preset mapping relationship between layer height and compensation boundary to constrain the spatial fusion range of the dominant benchmark and generate a spatiotemporal compensation vector.
[0058] The event generation module is used to superimpose the spatiotemporal compensation vector onto the positioning data in the reduced-order stage to generate calibrated spatiotemporal coordinates, associate the radar point cloud data with the video detection frame data, and convert it into a standardized traffic event stream output.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] 1. Through a dual threshold judgment mechanism based on floor height parameters and positioning status, the high-confidence compensation mode is dynamically activated when Beidou positioning is downgraded, significantly improving the spatiotemporal alignment accuracy of multi-source perception data. Specifically, building floor height is used as a quantitative indicator of environmental interference intensity, and a dynamic mapping relationship between floor height and compensation boundary is established. This allows the spatial fusion range to adaptively adjust with building density, breaking through the limitations of traditional fixed compensation thresholds. In high-rise dense areas, the compensation boundary of the radar-dominated benchmark is automatically expanded, effectively suppressing benchmark drift caused by satellite signal obstruction. In low-rise open areas, a video-dominated benchmark is used to maintain high-resolution spatial constraints.
[0061] 2. The generation mechanism of spatiotemporal compensation vectors integrates dual corrections for spatial offset and time delay. Dynamic compensation boundary coefficients constrain the data fusion range of the dominant benchmark, ensuring that the compensation amount always matches the current environmental interference intensity. The process of calibrating spatiotemporal coordinates and associating them with multi-source data adopts a collaborative strategy of spatial position matching and timestamp alignment. Radar point clouds achieve sub-meter spatial association through geometric center nearest neighbor search, and video data is dynamically aligned through sliding windows to eliminate inter-frame delays. The resulting standardized traffic event stream maintains high reliability in event type determination even in complex urban environments through a cross-validation mechanism for target size attributes and motion trajectory attributes, solving the trajectory splitting problem caused by misalignment of multi-source spatiotemporal benchmarks. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 This is a flow chart of a method for intelligently collecting traffic data based on Beidou multimodal fusion according to the present invention;
[0063] Figure 2 This is a structural schematic diagram of a traffic data intelligent collection system based on Beidou multimodal fusion in the present invention. DETAILED DESCRIPTION
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0065] Example 1: Figure 1 The present invention provides a method for intelligently collecting traffic data based on Beidou multimodal fusion, comprising:
[0066] S1. Obtain Beidou positioning data, radar point cloud data, and video detection frame data of the target object;
[0067] S2. Monitor the positioning status change information of Beidou positioning data, and trigger the benchmark calibration instruction when the Beidou signal is degraded from a fixed solution to a floating solution;
[0068] S3. When the benchmark calibration command is triggered, the layer height of the target object is extracted from the high-precision map;
[0069] S4. When the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold, the high confidence compensation mode is activated;
[0070] S5: When the high confidence compensation mode is activated, the radar point cloud data is selected as the dominant benchmark; when it is not activated, the video detection frame data is selected as the dominant benchmark; based on the preset mapping relationship between the layer height and the compensation boundary, the dynamic compensation boundary coefficient is generated to constrain the spatial fusion range of the dominant benchmark and generate the spatiotemporal compensation vector;
[0071] S6. Superimpose the spatiotemporal compensation vector onto the positioning data in the reduced-order stage to generate calibrated spatiotemporal coordinates, associate the radar point cloud data with the video detection frame data, and convert it into a standardized traffic event stream output.
[0072] S1. Obtain Beidou positioning data, radar point cloud data, and video detection frame data of the target object. The specific implementation is as follows:
[0073] Beidou positioning data of the target object is collected in real time through a Beidou positioning terminal. In specific implementation, the Beidou positioning terminal is installed on the top of the target vehicle. It has a built-in dual-frequency receiving chip that receives navigation signals transmitted by Beidou satellites in real time. During the signal reception phase, the Beidou positioning terminal first performs carrier phase and pseudorange measurements, and then obtains the three-dimensional geographic coordinates of the target object by calculating satellite ephemeris data. These three-dimensional geographic coordinates include longitude, latitude, and elevation information, and the corresponding acquisition timestamp is recorded with millisecond-level accuracy. Multipath mitigation technology is used during the acquisition process to eliminate interference from building-reflected signals. Specifically, the antenna array's beamforming technology is used to identify the direction of the signal source, and the weight coefficient for suppressing non-line-of-sight path signals is set within a reasonable range, such as 0.3 to 0.5. Beidou positioning data is output at a fixed frequency, such as a 10 Hz sampling frequency. Each output includes the location coordinates, the positioning precision dilution value, and the corresponding system timestamp.
[0074] Millimeter-wave radar scans the target object to generate radar point cloud data and record timestamps for the radar point cloud data. The millimeter-wave radar is deployed on a roadside pole, for example, at a height of 6 meters above the ground, with the radar's central axis tilted at a 15-degree angle to the horizontal. The millimeter-wave radar transmits a frequency-modulated continuous wave (FMCW) in a specific frequency band, such as 77 GHz. When the electromagnetic wave encounters a target vehicle, a reflection echo is generated. The radar receiver extracts the difference frequency signal through mixing and converts it into a range-Doppler spectrum using a fast Fourier transform. After detecting the peaks in the spectrum, a constant false alarm rate (CFAR) algorithm is used to filter out noise interference points, retaining valid target reflection points to form the radar point cloud data. Each radar point cloud data point contains the target's radial range, azimuth, and velocity relative to the radar, and a microsecond-accurate timestamp is recorded by the radar's internal clock. The radar scanning period is set to a fixed value, for example, generating a complete frame of radar point cloud data every 50 milliseconds.
[0075] Video surveillance equipment captures a video stream of the target object. The object detection model extracts video detection frame data and records the timestamp of the video detection frame data. The video surveillance equipment uses an industrial camera with a specific pixel count, such as 2 megapixels, to capture traffic scene video streams at a fixed rate, such as 25 frames per second. The video stream is then processed by a pre-trained object detection model, which utilizes a deep convolutional neural network architecture and consists of three main components: a feature extraction layer, a region proposal network, and a classification and regression layer. The feature extraction layer first generates a multi-scale feature map. The region proposal network then generates candidate object regions within the feature map. Finally, the classification and regression layer outputs the object category confidence and bounding box coordinates. The bounding box coordinates are expressed in pixel coordinates and include four parameters: the x-coordinate of the object's upper left corner in the image, the y-coordinate of the upper left corner, and the width and height. When each video detection frame data is generated, the corresponding exposure moment is recorded as a timestamp using the hardware trigger signal from the video capture device. The timestamp accuracy is controlled within a reasonable error range, for example, within 10 milliseconds.
[0076] The timestamps of Beidou positioning data, radar point cloud data, and video detection frame data are aligned to the same time base. This time base uses the Coordinated Universal Time (UTC) standard. Time synchronization is achieved through the following steps: First, during system initialization, the Beidou positioning terminal's timing function obtains accurate Coordinated Universal Time (UTC) as the reference time source. The millimeter-wave radar synchronizes its clock with the reference time source via the Network Time Protocol (NTP). This synchronization process employs clock offset compensation algorithms, such as the Markov clock offset compensation algorithm, to keep the radar's internal clock error within a reasonable range, for example, within 1 millisecond. Video surveillance equipment synchronizes with the reference time source via hardware trigger signals, recording the reference time corresponding to the rising edge of the synchronization pulse during each frame capture. The timestamp alignment process specifically includes: establishing a time series with the reference time as the horizontal axis; sorting the Beidou positioning data by acquisition timestamp; performing interpolation resampling of the radar point cloud data based on the frame generation time, such as linear interpolation or cubic spline interpolation, to ensure that the sampling interval is consistent with the Beidou positioning data; and matching the video detection frame data to the nearest time node based on the exposure timestamp. Finally, a time-aligned three-source dataset is generated. The data packet at the same time node contains Beidou positioning data, radar point cloud data, and video detection frame data at the corresponding moment.
[0077] S2. Monitor the positioning status change information of Beidou positioning data. When the Beidou signal is degraded from a fixed solution to a floating solution, trigger the benchmark calibration instruction. The specific implementation is as follows:
[0078] The positioning status word is parsed from Beidou positioning data in real time. Specifically, the positioning status word is a specific field embedded in the Beidou positioning data message. This field occupies a specific byte length, for example, 4 bytes, and contains the positioning mode identifier, satellite number information, and positioning accuracy indicators. The parsing process is implemented through the following steps: first, the status field is extracted from the received Beidou positioning data message; the extracted byte data is decomposed bit by bit, where specific bits represent the positioning mode code. A specific positioning mode code value indicates a fixed solution, while another specific value indicates a floating solution; the middle byte indicates the number of satellites involved in the solution; and the highest byte indicates the horizontal positioning precision value. The parsing program runs at a fixed frequency, for example, every 100 milliseconds, to ensure real-time monitoring of positioning status changes.
[0079] When the positioning status word indicates that the positioning mode changes from a fixed solution to a floating solution, the state change time point is recorded as the starting moment of the reduction. The specific judgment logic is: in two consecutive analysis operations, if the previous positioning mode code value is the corresponding value of the fixed solution and the current positioning mode code value is the corresponding value of the floating solution, then it is determined that a state change has occurred in which the fixed solution is reduced to a floating solution. The time point is recorded in the following method: when a state change is detected, the current system time is immediately obtained. The time is provided by the Beidou timing module and is accurate to the millisecond level, and it is marked as the starting moment of the reduction. The system time is expressed in Coordinated Universal Time, for example, it is recorded in the format of year-month-day hour: minute: second. millisecond. The starting moment of the reduction is stored in the temporary buffer as the key time reference, and the status continuous monitoring flag is triggered at the same time.
[0080] Based on the start time of downgrade, the system continuously monitors changes in the positioning status word of Beidou positioning data. This continuous monitoring process utilizes a sliding time window mechanism, with the time window length set as a multiple of a preset trigger threshold. For example, if the preset trigger threshold is 5 seconds, the time window is set to 10 seconds. Within this time window, the system performs the following operations at a fixed frequency, such as 200 milliseconds: parses the positioning status word in the newly received Beidou positioning data; checks whether the positioning mode code value corresponds to a floating solution; if so, updates the floating solution persistence counter; and if the positioning mode code value changes to a fixed solution, immediately terminates the monitoring process and clears the downgrade start time record. Changes in the positioning DOP are also recorded during the continuous monitoring process. When the DOP exceeds a specific threshold, an accuracy degradation alarm is generated without interrupting the monitoring process.
[0081] When the floating solution state lasts until the current moment and the length of time reaches the preset trigger threshold, a benchmark calibration instruction is generated. The duration length is calculated by subtracting the start time of the reduction order from the current system time, and the calculation result is in seconds with a specific number of decimal places. The preset trigger threshold is dynamically configured according to the application scenario. For example, it is set to 3 seconds in an urban elevated road environment and 5 seconds in a tunnel exit area. The threshold setting is based on analyzing the distribution of signal recovery time in different scenarios through historical data, taking a specific percentile recovery time as the base value and then adding a safety margin. When the calculated duration is greater than or equal to the preset trigger threshold, a benchmark calibration instruction message is generated. The message contains fields such as the instruction type code, the start time of the reduction order, the current time, and the positioning precision factor value. Setting the instruction type code to a specific value indicates that a benchmark calibration is required.
[0082] The generated benchmark calibration instructions are stored in association with the reduction start time. This associated storage uses a key-value pair data structure, where the key is the timestamp string of the reduction start time and the value is the benchmark calibration instruction message content. The storage process includes: first, creating an instruction cache in non-volatile memory; converting the reduction start time into a string format as an index key; encoding the benchmark calibration instruction message in binary format; writing the key-value pair into the cache and adding a time stamp. The associated storage also maintains a specific number of recent instruction records, automatically overwriting the oldest record when a new instruction is generated. The storage process ensures that the correspondence between the reduction start time and the benchmark calibration instruction is traceable. After the storage is complete, a command ready notification signal is sent.
[0083] S3. When the benchmark calibration command is triggered, the layer height of the target object is extracted from the high-precision map. The specific implementation is as follows:
[0084] Obtain the reduced-order start time associated with the baseline calibration instruction. In specific implementation, the system retrieves the most recently generated baseline calibration instruction record from the instruction cache in non-volatile memory, which is stored as a key-value pair. The retrieval process includes: querying the storage for the record with the most recent time tag; extracting the key field of the record, which contains the reduced-order start time in string format; and parsing the string into a time object containing year, month, day, hour, minute, second, and millisecond components. For example, the obtained reduced-order start time is time data in the format of "2023-08-15 14:30:25.356," which uses the same Coordinated Universal Time reference as the timestamp of the Beidou positioning data. Data validity is verified during the acquisition process to check whether the time object is within a reasonable time range, for example, no earlier than the system startup time and no later than the current time.
[0085] Extract the target object's location coordinates corresponding to the reduction start time from the Beidou positioning data. The specific implementation includes: searching the time node closest to the reduction start time in the time-aligned Beidou positioning data set; extracting the Beidou positioning data recorded at that time node when the time deviation is less than a specific threshold, such as 50 milliseconds; and parsing the three-dimensional geographic coordinate components from the positioning data, including longitude, latitude, and elevation values. Coordinate extraction uses linear interpolation to improve accuracy. When the reduction start time falls between two sampling points, interpolation calculations are performed based on the coordinate values and time interval of the two sampling points. For example, between time points T1 and T2, the coordinate corresponding to time T is calculated by calculating the difference between the previous and next coordinates based on the time ratio. The extracted coordinate values are retained to a specific number of decimal places, such as 6 decimal places for longitude and latitude and 2 decimal places for elevation.
[0086] The target object position coordinates are converted into matching coordinates in the HD map coordinate system. The coordinate system conversion uses a parametric model that includes translation, rotation, and scale factors. The specific conversion process includes: first obtaining the coordinate system parameters defined by the HD map; loading the preset conversion parameter matrix, which is obtained by ground control point calibration; converting the geographic coordinates of the Beidou positioning data into spatial rectangular coordinates; applying the conversion parameter matrix to perform coordinate system transformation; and converting the transformation result back to the geographic coordinates of the target coordinate system. The conversion process adds elevation anomaly correction and corrects the elevation value by querying the local geoid model. For example, the coordinates of a point before conversion are a specific numerical combination, and the coordinates after conversion are another specific numerical combination. The conversion result must meet the accuracy requirements, and the plane position error must be controlled within a specific range, such as within 0.5 meters.
[0087] Query the floor height attribute field of the high-precision map based on the matching coordinates. The high-precision map is stored in a layered vector data structure. The query operation includes: using the matching coordinates as the query point; performing a spatial search in the building layer index to find the building polygon containing the query point; when the query point is inside the building polygon, extracting the attribute table of the polygon; locating the floor height attribute field in the attribute table. Spatial retrieval uses an index structure to accelerate the query, and the query accuracy sets a safety distance tolerance. For example, when a point is less than 1 meter away from the building boundary, it is considered to be inside the building. The floor height attribute field stores numerical data, indicating the vertical height from the ground to the roof of the building, in meters. For example, the floor height attribute value obtained by querying a certain office building is a specific value.
[0088] The value of the floor height attribute field is extracted as the floor height of the target object's location. The extraction process includes: verifying that the attribute value's data type is numeric; checking that the value is within a reasonable range, for example, residential buildings are typically between 3 and 100 meters; performing string cleaning when the attribute value contains unit symbols; and converting valid values into floating-point format. Special scenarios are handled: when the target object is located above the road, the floor height is the height of the elevated road surface from the ground; when located at an intersection, the average floor height of the surrounding buildings is used. The extracted floor height value is associated and stored in the target object's feature vector, establishing a corresponding relationship with the benchmark calibration instruction. The final output floor height value is used in subsequent processing flows, such as outputting a specific value to represent the floor height of the current location. The floor height value retains a specific precision, such as retaining 1 decimal place.
[0089] S4. When the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold, the high confidence compensation mode is activated. The specific implementation is as follows:
[0090] Get the layer height and the start time of the reduction of the target object. In specific implementation, the system extracts the layer height value associated with the benchmark calibration instruction from the target object feature vector storage area. The layer height value is a floating-point data obtained from the layer height attribute field of the high-precision map when executing the benchmark calibration process. The unit is meter. At the same time, the reduction start time bound to the current benchmark calibration instruction is retrieved from the instruction cache. This moment is a precise timestamp expressed in Coordinated Universal Time. The acquisition process includes data validity verification: checking whether the layer height value is within a preset reasonable range, such as between 3 meters and 300 meters; verifying whether the reduction start time maintains a reasonable timing relationship with the current system time, such as not later than the current moment and not earlier than the system startup time. When the data is invalid, the exception handling process is triggered, such as re-executing the benchmark calibration instruction acquisition process.
[0091] The duration of the reduction is calculated based on the current time and the start time of the reduction. The calculation process uses a time difference algorithm: first, obtain the current coordinated universal time provided by the Beidou timing module, with a time accuracy of milliseconds; convert the current time and the start time of the reduction to millisecond timestamps of the same time base; obtain the time difference by subtracting the millisecond timestamps; divide the time difference by the conversion factor 1000 to convert it into the duration of the reduction in seconds, and retain a specific number of decimal places in the calculation result, such as 2 decimal places. The calculation process includes boundary condition processing: when the calculation result is a negative value, it indicates a time logic error, at which point the duration value is cleared and the time synchronization calibration process is triggered; when the calculation result exceeds the maximum allowable value, such as 3600 seconds, it is automatically truncated to the maximum value and a timeout alarm is generated.
[0092] The floor height is numerically compared with a preset floor height threshold; the duration of the continuous order reduction is numerically compared with a preset time limit threshold. The floor height threshold is a critical value set based on the height distribution characteristics of urban buildings. Its setting method includes statistically analyzing historical building height data in the target area and selecting a specific percentile height value as the base threshold. It also considers the impact of different road types, such as adding a safety margin in urban expressway scenarios. A typical threshold setting is 20 meters, but in practice, it is dynamically adjusted based on geo-fences. The time limit threshold is set based on signal recovery characteristics and is determined through the following steps: collecting historical positioning order reduction event samples and statistically analyzing the distribution of normal signal recovery times; and multiplying the recovery time by a specific percentile by a safety factor as the base threshold. Numerical comparisons use a floating-point total order comparison algorithm: a greater-than relationship is determined between the floor height value and the floor height threshold; and a greater-than relationship is determined between the continuous order reduction time and the time limit threshold. The comparison operation includes tolerance processing. For example, if the floor height difference is within a specific positive or negative range, it is considered critical and requires secondary verification.
[0093] When the floor height is greater than the floor height threshold and the continuous order reduction time exceeds the time limit threshold, a high-confidence compensation mode activation instruction is generated. The judgment logic uses a parallel condition detection mechanism: the floor height comparison result and the duration comparison result are monitored simultaneously; the instruction generation is triggered only when both conditions are true. The instruction generation process includes: constructing an instruction data structure, including fields such as the instruction type identifier, trigger timestamp, measured floor height value, and continuous order reduction time value; the instruction type identifier is set to a specific code; the trigger timestamp uses the current system time; and the measured data field is filled with the latest floor height value and continuous order reduction time value. After the instruction is generated, a logical check is performed: verifying that the floor height value is indeed greater than the currently effective floor height threshold; verifying that the continuous order reduction time is indeed greater than the currently effective time limit threshold. After the check passes, the instruction is written to the instruction queue, and the system status flag is updated to the high-confidence compensation mode ready state.
[0094] When the floor height is less than or equal to the floor height threshold or the continuous order reduction time does not exceed the time limit threshold, the high confidence compensation mode is maintained in an inactive state. This branch processing includes two situations: the first situation is that the floor height does not exceed the threshold but the continuous order reduction time exceeds the limit, in which case it is determined to be ordinary positioning interference; the second situation is that the continuous order reduction time does not exceed the limit but the floor height exceeds the limit, in which case it is determined to be a short signal fluctuation. The processing flow includes: resetting the continuous order reduction time counter; keeping the high confidence compensation mode status flag as an inactive value; generating a status maintenance log record, which contains the current floor height value, continuous order reduction time value and decision timestamp. In the inactive state, the system executes conventional positioning compensation strategies, such as using the historical trajectory prediction algorithm for position correction. At the same time, the delay monitoring mechanism is started: if the floating solution state continues to reach the warning time limit, for example, 30 seconds, a secondary alarm notification is triggered.
[0095] S5. When the high confidence compensation mode is activated, the radar point cloud data is selected as the dominant benchmark. When it is not activated, the video detection frame data is selected as the dominant benchmark. Based on the preset mapping relationship between the layer height and the compensation boundary, the dynamic compensation boundary coefficient is generated to constrain the spatial fusion range of the dominant benchmark and generate the spatiotemporal compensation vector. The specific implementation is as follows:
[0096] The dominant reference data source is selected based on the activation status of the high-confidence compensation mode. The specific implementation process monitors the high-confidence compensation mode activation flag in the system status register in real time. This flag is a Boolean variable and is set by the high-confidence compensation mode activation instruction generated in step S4. When the flag is detected to be active, the point cloud data collected by the millimeter-wave radar sensor is selected as the dominant reference for position correction. This selection is based on the physical advantages of radar sensors in high-rise building environments: millimeter-waves can penetrate rain and fog and are less affected by multipath effects. Analysis of historical sensor failure records shows that in areas with floor heights exceeding a certain value, such as 20 meters, the probability of radar failure is lower than that of video sensors by a certain percentage, such as 15% to 30%. Point cloud data contains the three-dimensional spatial coordinate information of target objects and is suitable for stereoscopic positioning in high-rise areas. The data is generated by processing the raw radar signal with a point cloud clustering algorithm. Each frame of data contains a specific number of spatial point coordinates, such as 256. When the flag is inactive, the video detection frame data collected by the camera is selected as the dominant reference. The selection was based on the technical characteristics of video sensors in low-rise environments: visible light imaging has high accuracy in identifying target texture features. Measured data analysis shows that in areas below a specific floor height, such as 15 meters, the position recognition error of the video detection frame is lower than that of radar by a specific percentage, such as 10% to 25%. The detection frame data includes the target's length, width, and motion direction, making it suitable for continuous trajectory tracking on planar roads. This data is obtained by processing video frames through a target detection neural network. Data source switching is implemented using a multiplexer hardware circuit. Reference switching is completed within a specific time threshold, such as 50 milliseconds, after a state change. A log of the switching process contains the switch timestamp and reference data type.
[0097] The target object's floor height is obtained as a quantitative parameter for environmental interference intensity. The floor height data is derived from the value of the floor height attribute field output in step S3. This value is a floating-point variable with a uniform unit of meters. A mapping relationship between floor height and environmental interference intensity is established using the following method: test equipment is deployed in a typical urban area to collect measured multipath interference intensity values at locations with different floor heights. A linear regression model is then established between floor height and interference intensity. The model shows that for every specific increase in floor height (e.g., 10 meters), the interference intensity increases by a specific percentage (e.g., 0.3 standard units). After obtaining the floor height value, the system performs a normalization process: the raw floor height value is divided by a baseline floor height reference value (e.g., 30 meters) to obtain a dimensionless coefficient representing the environmental interference intensity. This processing includes data validation: if the floor height value exceeds a preset range (e.g., 0 to 300 meters), it is forced to the boundary value and a data anomaly alarm log is generated.
[0098] The dynamic compensation boundary coefficient is calculated based on a preset mapping relationship between floor height and compensation boundary. This mapping relationship is stored in a coefficient lookup table as a piecewise linear function. The function construction method involves dividing floor height into three characteristic intervals: the first interval is a low-interference zone (0 to 20 meters), the second interval is a medium-interference zone (20 to 50 meters), and the third interval is a high-interference zone (over 50 meters). Baseline compensation coefficients for each interval are calibrated through real-world road testing. The calibration process involves adjusting the compensation boundary values in various scenarios until positioning error is minimized, resulting in a baseline coefficient of 0.8 for the first interval, 1.2 for the second interval, and 1.6 for the third interval. The dynamic compensation boundary coefficient calculation process is as follows: read the current floor height value; determine the characteristic interval to which it belongs; obtain the baseline compensation coefficient for that interval; and calculate the final coefficient by linear interpolation based on the relative position of the floor height within the current interval. For example, for a floor height of 35 meters in the 20 to 50 meter interval, the calculation process is: add the baseline coefficient of 1.2 to (35-20) / (50-20)×(1.6-1.2) to obtain the interpolated result. The coefficient output range is limited to between 0.5 and 2.0. Increasing the value indicates a corresponding increase in the allowable position correction amplitude. Exception handling is performed: When the floor height data is invalid, the default coefficient of 1.0 is used and the data source re-acquisition mechanism is triggered.
[0099] The dynamic compensation boundary coefficient is used to constrain the position correction amplitude of the dominant reference during the spatial fusion process. Spatial fusion processing utilizes a weighted fusion algorithm. The correction amplitude constraint for the dominant reference is achieved through the following steps: First, the three-dimensional deviation vector between the Beidou positioning coordinates and the dominant reference coordinates is calculated. The deviation vector contains easting, northing, and elevation components. The modulus of the horizontal plane deviation vector (easting and northing components) is used as the correction reference value. The correction reference value is multiplied by the dynamic compensation boundary coefficient to obtain the maximum allowable correction amplitude threshold. If the original correction reference value exceeds the maximum correction amplitude threshold, the actual correction is compressed to within this threshold range. The specific constraint rule is as follows: If the horizontal distance between the Beidou positioning point and the dominant reference point is greater than the product of the dynamic compensation boundary coefficient and the baseline distance threshold, the maximum correction amplitude threshold is used as the final correction; otherwise, the original deviation value is used. The baseline distance threshold is determined through statistical analysis of historical positioning data, with a typical value of 3 meters. The elevation component is independently constrained, limiting the maximum elevation correction to a specific value, such as 1.5 meters. The constrained dominant reference coordinates serve as the core input for spatial fusion, ensuring that trajectory jumps caused by excessive corrections are avoided in high-altitude, high-interference environments.
[0100] A spatiotemporal compensation vector is generated based on the deviation between the dominant datum and the Beidou positioning data after a limited position correction. This vector is a four-dimensional data structure consisting of an easting offset, a northing offset, an elevation offset, and a time compensation term. The generation process includes calculating the coordinate difference between the constrained and corrected dominant datum coordinates and the current Beidou positioning coordinates; converting the coordinate difference into a three-dimensional spatial offset in the northeast celestial coordinate system; and adding a time dimension compensation. The time compensation is calculated by multiplying the continuous reduction time calculated in step S4 by a velocity attenuation factor. The specific generation rules for the spatiotemporal compensation vector are as follows: the easting offset of the spatial component is the easting component of the coordinate difference multiplied by a confidence weight; the northing offset is the northing component of the coordinate difference multiplied by a confidence weight; and the elevation offset is the elevation component of the coordinate difference. The time compensation is the continuous reduction time multiplied by a velocity attenuation factor, which is dynamically adjusted based on the rate of change of the positioning state. The resulting spatiotemporal compensation vector is stored as a structured data volume, consisting of a timestamp field, a spatial offset field, a time compensation field, and a confidence score field. This vector is transmitted via the data bus and is used to perform real-time compensation calibration on the original BeiDou positioning results. The vector update frequency is synchronized with the dominant reference data update frequency, typically a certain number of times per second, such as 10 times.
[0101] S6. Superimpose the spatiotemporal compensation vector onto the positioning data in the order reduction stage to generate calibrated spatiotemporal coordinates, associate the radar point cloud data with the video detection frame data, and convert it into a standardized traffic event stream output. The specific implementation is as follows:
[0102] Obtain the Beidou positioning data and spatiotemporal compensation vectors for the reduced-order phase. During implementation, the system reads the Beidou positioning data packet for the reduced-order phase defined by step S2 from the positioning data cache. This data packet contains a timestamp field, a longitude coordinate field, a latitude coordinate field, and a positioning precision factor field. The timestamp accuracy is milliseconds, and the coordinates use the internationally accepted WGS84 geographic coordinate system. The spatiotemporal compensation vector is derived from the structured data output generated by step S5. This data contains an easting offset component, a northing offset component, an elevation offset component, and a time compensation component. The physical units of each component are standard international units. The data acquisition process performs validity verification operations: checking whether the timestamp of the Beidou positioning data is after the start of the reduction and no earlier than a specific time range of the current time, such as a time window of no more than 30 minutes; and verifying whether the confidence score field value of the spatiotemporal compensation vector is higher than a preset validity threshold, such as a score threshold of 0.7. When invalid data is detected, the data re-acquisition mechanism is triggered. The specific operation is to send a data regeneration request instruction to step S5 and wait for a response timeout. For example, if there is no response within 2000 milliseconds, the backup compensation vector is activated.
[0103] The space-time compensation vector is superimposed onto the reduced-order Beidou positioning data to generate calibrated space-time coordinates. The superposition calculation utilizes a process combining spatial vector addition and time compensation. First, the Beidou positioning latitude and longitude coordinates are converted to three-dimensional coordinates in a spatial rectangular coordinate system using a coordinate conversion function, employing a standard ellipsoid projection algorithm. The easting, northing, and elevation offset components of the space-time compensation vector are then axially superimposed onto the three-dimensional coordinates in the same coordinate system. The superposition calculation formula states that the calibrated coordinate value equals the original coordinate value plus the corresponding axial offset component. Time dimension processing adds the time compensation component to the original Beidou timestamp to generate a calibrated timestamp. The calculation process incorporates data plausibility constraints: if the elevation offset causes the calculated elevation to exceed a specific percentage of the maximum terrain elevation in the geographic information database, such as 120%, the elevation is clipped using the terrain elevation boundary value at that location. The calibrated spatial rectangular coordinates are then converted to standard latitude and longitude format using an inverse conversion function, ultimately forming a calibrated space-time coordinate data structure consisting of a calibrated timestamp field, a calibrated longitude field, and a calibrated latitude field.
[0104] The calibration spatiotemporal coordinates are spatially associated with the radar point cloud data. This association is performed using a nearest neighbor search algorithm: a three-dimensional cube search region is constructed with the calibration spatiotemporal coordinate's spatial location as the center point, with a specific side length, such as a 3-meter search radius. The radar point cloud data frame output in step S5 is retrieved for all spatial point cloud clusters that fall within this cube region. The Euclidean distance between the geometric center point of each point cloud cluster and the calibration spatial location point is calculated, and the point cloud cluster with the smallest distance is selected for association. Association conditions include: the spatial distance between the point cloud cluster's geometric center and the calibration location must be less than an association threshold, such as a 1.5-meter distance threshold; and the absolute difference between the point cloud cluster's timestamp and the calibration timestamp must be less than a time tolerance window, such as a ±100-millisecond window. If association fails, a progressive expansion search is initiated: the search cube size is gradually expanded by a fixed step size, for example, by 0.5-meter side length increments, until the maximum search boundary, such as a 5-meter side length limit, is reached. The unique identifiers of successfully associated point cloud clusters are written into the association record table, along with the actual spatial deviation values.
[0105] The calibrated spatiotemporal coordinates are timestamp-aligned and associated with the video detection frame data. This timestamp alignment utilizes a sliding window matching mechanism: Using the calibration timestamp field value of the calibrated spatiotemporal coordinates as the reference time point, a symmetrical time alignment window is established with a window size set to twice the video frame sampling interval, for example, a 66 millisecond window width for a video frame rate of a certain number of times per second. A dataset of all video detection frames whose timestamps fall within this window is retrieved from the video stream. The absolute time difference between each detection frame's timestamp and the reference time point is calculated. The detection frame with the smallest time difference is selected for association. Association validation rules include: the horizontal projection distance between the detection frame's center point coordinates and the calibration position must be less than a spatial matching threshold, such as a 2-meter planar distance threshold; and the detection frame's aspect ratio must conform to the physical constraints of the target object. For example, the aspect ratio of a vehicle target must be within a specific range, such as 1.5 to 3.0. Association result records, containing the detection frame's global identifier and actual time offset value, are stored in an association database.
[0106] Target size attributes are extracted from the associated radar point cloud data, and target trajectory attributes are extracted from the associated video detection frame data. The target size attribute extraction method performs a minimum bounding cube calculation on the associated point cloud clusters to generate a three-dimensional attribute vector containing length, width, and height dimensions. The length dimension is defined as the physical dimension of the longest side of the cube, the width as the second-longest side, and the height as the shortest side. The target trajectory attribute extraction method obtains the coordinate sequence of the center points of the current associated video detection frame and the associated detection frames of the three previous frames. The displacement vectors of the center points of adjacent frames are calculated, consisting of horizontal and vertical components. The displacement vectors are divided by the inter-frame time interval to obtain the instantaneous velocity vector. The average acceleration scalar value is calculated from the displacement vectors of three consecutive frames. The extraction process includes data preprocessing: the size vector is filtered using a median filter algorithm to eliminate outliers, with a filter window size of a specific value, such as five historical samples. The velocity vector is smoothed using a Kalman filter algorithm, with the state equation parameters dynamically adjusted based on the target type.
[0107] The fusion of target size attributes and target trajectory attributes generates a standardized traffic event stream containing event type, timestamp, and spatial coordinates. This fusion process is performed using a multi-attribute decision-making model. First, the target object type is determined based on the target size attribute vector. For example, if the length dimension exceeds a certain threshold (e.g., 6 meters) and the height dimension is less than a certain threshold (e.g., 2 meters), it is identified as a large vehicle. The traffic event type is then determined based on the target trajectory attributes. For example, if the instantaneous speed value is less than a certain threshold (e.g., 5 meters per second) and the acceleration value is less than a certain threshold (e.g., -3 meters per second squared), it is identified as a sudden braking event. The standardized traffic event stream data structure consists of three core fields: an event type code field, which uses a preset enumeration value to represent different event types; a timestamp field, which directly inherits the value of the calibrated timestamp field of the calibrated spatiotemporal coordinates; and a spatial coordinate field, which uses the longitude and latitude fields of the calibrated spatiotemporal coordinates. Event generation rules include: a lane change event is generated when the target width dimension change rate exceeds a certain proportional threshold (e.g., 15%) and the velocity direction angle changes by more than a certain angle threshold (e.g., 30 degrees); and a lift event is generated when the height dimension value continues to increase and the horizontal velocity value is less than a certain threshold (e.g., 1 meter per second). The final event stream data is encapsulated in JSON structured format and output to the event processing interface of the traffic management platform through the TCP / IP protocol. The data update frequency is consistent with the generation frequency of the calibrated space-time coordinates, for example, a transmission rate of 10 times per second.
[0108] The technical solution formed by the collaborative efforts of steps S1 through S6 in this embodiment addresses the issue of vehicle-based positioning degradation by using building height as a quantitative parameter for environmental interference intensity and establishing a dynamic compensation boundary coefficient based on this parameter. This overcomes the technical bias of conventional positioning compensation that relies on fixed thresholds or signal strength. By establishing a dynamic mapping relationship based on the physical correlation between building height and multipath interference, spatially characterized modeling of interference intensity is achieved.
[0109] The generation mechanism for spatiotemporal compensation vectors integrates dual compensation for spatial offset and time delay. The time compensation is achieved through the dynamic coupling of continuously reduced-order time and velocity attenuation factors, addressing the time-domain cumulative error caused by satellite signal obstruction. The association logic for calibrating spatiotemporal coordinates and multi-source sensor data employs a dual constraint mechanism for spatial position matching and timestamp alignment. Specifically, a geometric center nearest neighbor search is employed for radar point clouds, and a sliding window dynamic alignment is employed for video data. This differentiated association strategy overcomes the inherent drawback of heterogeneous data, which suffers from inconsistent spatiotemporal benchmarks.
[0110] The standardized traffic event stream generation process cross-validates physical dimensional features with kinematic characteristics through a collaborative decision-making mechanism that combines target size and trajectory attributes, reducing the false alarm rate in traffic event identification. The entire technology chain is highly operational at the implementation level. All parameters, such as floor height data, are derived from real-time output from the onboard environmental perception system. Compensation coefficients are computationally lightweight using piecewise linear functions. Boundary condition processing mechanisms address real-world scenarios such as data anomalies and association failures, forming a complete technical closed loop.
[0111] Example 2: Figure 2 The present invention provides a structural schematic diagram of a traffic data intelligent collection system based on Beidou multimodal fusion, which includes:
[0112] Multi-source perception module, used to obtain Beidou positioning data, radar point cloud data and video detection frame data of the target object;
[0113] The downgrade monitoring module is used to monitor the positioning status change information of Beidou positioning data, and triggers the benchmark calibration instruction when the Beidou signal is downgraded from a fixed solution to a floating solution;
[0114] The floor height extraction module is used to extract the floor height of the target object from the high-precision map when the benchmark calibration instruction is triggered;
[0115] A mode activation module is used to activate the high confidence compensation mode when the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold;
[0116] The benchmark compensation module selects radar point cloud data as the dominant benchmark when the high-confidence compensation mode is activated, and selects video detection frame data as the dominant benchmark when it is not activated. It generates dynamic compensation boundary coefficients based on the preset mapping relationship between layer height and compensation boundary to constrain the spatial fusion range of the dominant benchmark and generate a spatiotemporal compensation vector.
[0117] The event generation module is used to superimpose the spatiotemporal compensation vector onto the positioning data in the reduced-order stage to generate calibrated spatiotemporal coordinates, associate the radar point cloud data with the video detection frame data, and convert it into a standardized traffic event stream output.
[0118] The calculations involved in the embodiments are all dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to actual conditions.
[0119] It should be noted that the present invention can be deployed on the device itself to implement embedded applications, and can also be run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.
[0120] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wireless or wired transmission. Wired transmission methods include optical fiber, twisted pair, coaxial cable, etc.; wireless transmission methods include infrared, microwave, etc. The computer-readable storage medium can be any available medium accessible by a computer, or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.
[0121] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0123] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0124] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0125] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0126] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0127] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for intelligently collecting traffic data based on Beidou multimodal fusion, characterized in that: include: S1. Obtain Beidou positioning data, radar point cloud data, and video detection frame data of the target object; S2. Monitor the positioning status change information of Beidou positioning data, and trigger the benchmark calibration instruction when the Beidou signal is degraded from a fixed solution to a floating solution; S3. When the benchmark calibration command is triggered, the layer height of the target object is extracted from the high-precision map; S4. When the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold, the high confidence compensation mode is activated; S5: When the high confidence compensation mode is activated, the radar point cloud data is selected as the dominant benchmark; when it is not activated, the video detection frame data is selected as the dominant benchmark; based on the preset mapping relationship between the layer height and the compensation boundary, the dynamic compensation boundary coefficient is generated to constrain the spatial fusion range of the dominant benchmark and generate the spatiotemporal compensation vector; S6. Superimpose the spatiotemporal compensation vector onto the positioning data in the reduced-order stage to generate calibrated spatiotemporal coordinates, associate the radar point cloud data with the video detection frame data, and convert it into a standardized traffic event stream output.
2. The method for intelligently collecting traffic data based on Beidou multimodal fusion according to claim 1 is characterized in that: Obtain Beidou positioning data, radar point cloud data, and video detection frame data of the target object, including: Collect Beidou positioning data of target objects in real time through Beidou positioning terminals; Scan the target object with millimeter-wave radar to generate radar point cloud data and record the timestamp of the radar point cloud data; Capture the video stream of the target object through the video surveillance device, extract the video detection frame data based on the target detection model, and record the timestamp of the video detection frame data; Align the timestamps of Beidou positioning data, radar point cloud data, and video detection frame data to the same time base.
3. The method for intelligently collecting traffic data based on Beidou multimodal fusion according to claim 2 is characterized in that: Monitor the positioning status change information of Beidou positioning data, and trigger the benchmark calibration command when the Beidou signal is degraded from a fixed solution to a floating solution, including: Real-time analysis of positioning status words from Beidou positioning data; When the positioning status word indicates that the positioning mode changes from fixed solution to floating solution, the state change time point is recorded as the starting time of the reduction order; Continuously monitor the changes in the positioning status word of Beidou positioning data based on the start time of the reduction order; When the floating solution state lasts until the current moment and the length of time reaches the preset trigger threshold, a reference calibration instruction is generated; The generated reference calibration instruction is associated with the start time of the order reduction and stored.
4. The method for intelligently collecting traffic data based on Beidou multimodal fusion according to claim 3 is characterized in that: When the benchmark calibration command is triggered, the layer height of the target object is extracted from the HD map, including: Obtaining the reduced-order start time associated with the reference calibration instruction; Extract the target object position coordinates corresponding to the start time of order reduction from Beidou positioning data; Convert the target object's position coordinates into matching coordinates in the high-precision map coordinate system; Query the layer height attribute field of the HD map based on the matching coordinates; Extract the value of the layer height attribute field as the layer height where the target object is located.
5. The method for intelligently collecting traffic data based on Beidou multimodal fusion according to claim 4 is characterized in that: When the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold, the high confidence compensation mode is activated, including: Get the target object's layer height and the start time of order reduction; Calculate the duration of the scale reduction based on the current time and the start time of the scale reduction; Compare the floor height with the preset floor height threshold; compare the continuous downgrade time with the preset time limit threshold; When the layer height is greater than the layer height threshold and the continuous order reduction time exceeds the time limit threshold, a high confidence compensation mode activation instruction is generated; When the floor height is less than or equal to the floor height threshold or the continuous order reduction time does not exceed the time limit threshold, the high confidence compensation mode is maintained in the inactive state.
6. The method for intelligently collecting traffic data based on Beidou multimodal fusion according to claim 5 is characterized in that: When high confidence compensation mode is activated, radar point cloud data is selected as the dominant benchmark; when it is not activated, video detection frame data is selected as the dominant benchmark; Based on the preset mapping relationship between floor height and compensation boundary, dynamic compensation boundary coefficients are generated to constrain the spatial fusion range of the dominant benchmark and generate spatiotemporal compensation vectors, including: According to the high confidence compensation mode activation state: When the high confidence compensation mode is activated, the radar point cloud data is selected as the dominant benchmark based on the low failure probability characteristics of the radar sensor in a high-rise environment; When the high confidence compensation mode is not activated, the video detection frame data is selected as the dominant benchmark based on the motion continuity characteristics of the video sensor in the low-level environment; Obtain the floor height of the target object as a characterization parameter of the environmental interference entropy value; Calculate the dynamic compensation boundary coefficient based on the preset mapping relationship between the floor height and the compensation boundary; Use dynamic compensation boundary coefficients to limit the position correction amplitude of the dominant reference during the spatial fusion process; A spatiotemporal compensation vector is generated based on the deviation between the dominant reference and the Beidou positioning data after limiting the position correction amplitude.
7. The method for intelligently collecting traffic data based on Beidou multimodal fusion according to claim 6 is characterized in that: As the floor height increases, the entropy value of environmental interference increases, and the dynamic compensation boundary coefficient increases synchronously.
8. The method for intelligently collecting traffic data based on Beidou multimodal fusion according to claim 6 is characterized in that: The time-space compensation vector is superimposed on the positioning data in the order reduction stage to generate calibrated time-space coordinates. The radar point cloud data and video detection frame data are then associated and converted into a standardized traffic event stream output, including: Obtain Beidou positioning data and time-space compensation vectors in the reduced-order stage; The time-space compensation vector is superimposed on the BeiDou positioning data in the reduced-order stage to generate the calibrated time-space coordinates; Associating the calibrated spatiotemporal coordinates with the radar point cloud data; Align the calibrated spatiotemporal coordinates with the video detection frame data by timestamp; Extract target size attributes from the associated radar point cloud data, and extract target motion trajectory attributes from the video detection frame data; The target size attribute and target motion trajectory attribute are fused to generate a standardized traffic event stream containing event type, timestamp, and spatial coordinates.
9. A traffic data intelligent collection system based on Beidou multimodal fusion, used to implement the traffic data intelligent collection method based on Beidou multimodal fusion according to any one of claims 1 to 8, characterized in that: include: Multi-source perception module, used to obtain Beidou positioning data, radar point cloud data and video detection frame data of the target object; The downgrade monitoring module is used to monitor the positioning status change information of Beidou positioning data, and triggers the benchmark calibration instruction when the Beidou signal is downgraded from a fixed solution to a floating solution; The floor height extraction module is used to extract the floor height of the target object from the high-precision map when the benchmark calibration instruction is triggered; A mode activation module is used to activate the high confidence compensation mode when the floor height is greater than the floor height threshold and the continuous degradation time represented by the positioning state change information exceeds the time limit threshold; The benchmark compensation module is used to select the radar point cloud data as the dominant benchmark when the high confidence compensation mode is activated, and to select the video detection frame data as the dominant benchmark when it is not activated; Based on the preset mapping relationship between the floor height and the compensation boundary, a dynamic compensation boundary coefficient is generated to constrain the spatial fusion range of the dominant benchmark and generate a spatiotemporal compensation vector; The event generation module is used to superimpose the spatiotemporal compensation vector onto the positioning data in the reduced-order stage to generate calibrated spatiotemporal coordinates, associate the radar point cloud data with the video detection frame data, and convert it into a standardized traffic event stream output.
Citation Information
Patent Citations
Beidou-based plant transportation multi-dimensional digital dynamic supervision method and system
CN118690938A
Traffic target monitoring and tracking system and method based on monocular vision daytime scene reconstruction
CN119600550A