Multi-source perception data fusion and trajectory reconstruction method for mixed traffic flow
Patent Information
- Application Number
- CN202610823216.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-06-09
AI Technical Summary
这使得卡尔曼滤波器将伪位移误判为真实机动,引发协方差矩阵频繁误调整,最终融合轨迹呈现锯齿状高频震荡,严重消耗带宽,难以达到厘米级平滑
[0022] By extracting the gait fundamental frequency from the inertial acceleration sequence across modes and mapping it to the visual sampling rate to construct temporal filtering coefficients, the geometric center coordinates are filtered to eliminate periodic deformation interference that is in sync with the gait frequency. This breaks the assumption that the geometric center and the physical centroid are rigidly coincident, and eliminates the sawtooth high-frequency oscillation of the Kalman fusion trajectory from the root cause.
Smart Images

Figure CN122365405B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multi-source sensor fusion localization and target tracking technology, specifically a method for fusion of multi-source sensing data and trajectory reconstruction of mixed traffic flow. Background Technology
[0002] With the rapid development of intelligent transportation and autonomous driving, there is an urgent need for high-precision continuous positioning of pedestrians in mixed traffic flows. Single sensors struggle to balance accuracy, frequency response, and interference resistance, making the fusion of inertial measurement units (IMUs) and visual cameras the mainstream approach. A typical architecture involves a cascaded object detection network and a Kalman filter: the geometric center of the two-dimensional bounding box is used as the estimate of the pedestrian's physical centroid, which is then used as the measurement input to correct the inertial prior state and output a fused trajectory.
[0003] However, during normal walking, the stride and arm swing cause periodic, breathing-like deformations in the visual bounding box, resulting in a pseudo-displacement of the geometric center relative to the true centroid that resonates with the gait. For a long time, this pseudo-displacement has been mistakenly treated as algorithmic noise, with attempts made to suppress it through superficial measures such as deepening the network and introducing attention mechanisms, without recognizing its biomechanical root cause. This causes the Kalman filter to misinterpret the pseudo-displacement as genuine maneuvering, leading to frequent misadjustments of the covariance matrix. Ultimately, the fused trajectory exhibits jagged, high-frequency oscillations, severely consuming bandwidth and making it difficult to achieve centimeter-level smoothness.
[0004] Furthermore, fixed-time-window spectral analysis suffers from a rigid contradiction between frequency resolution and time response: an excessively long window leads to sluggish maneuver response, while an excessively short window results in insufficient frequency extraction accuracy. Therefore, eliminating gait deformation pseudo-displacements, overcoming the rigid assumption of homogeneous geometric centers, and achieving dynamic adaptive adjustment of time-frequency parameters have become urgent technical problems to be solved. Summary of the Invention
[0005] The purpose of this application is to provide a method for fusing and reconstructing trajectory data from multiple sources of mixed traffic flow. The aim is to use cross-modal filtering to remove pseudo-displacements caused by gait, thereby fundamentally eliminating jagged oscillations in the trajectory and making the fused trajectory smooth and stable.
[0006] The objective of this application can be achieved through the following technical solution: Firstly, a method for fusing and reconstructing trajectory data from multiple sources of hybrid traffic flow, comprising the following steps:
[0007] Obtain the three-dimensional acceleration sequence of the target object and its RGB video stream, and extract the geometric center coordinate sequence of the target object from the RGB video stream;
[0008] Based on the frequency domain energy distribution characteristics of the three-dimensional acceleration sequence within a preset sliding time window, the target fundamental frequency parameter characterizing the periodic kinematic changes of the target object is obtained;
[0009] The dynamic sampling rate of the RGB video stream within a preset sliding time window is obtained, and a temporal filtering coefficient for the morphological fluctuation of the target object is constructed in combination with the target fundamental frequency parameter.
[0010] The geometric center coordinate sequence is subjected to time-domain filtering based on the time-domain filtering coefficients to obtain a corrected center coordinate sequence, and a measurement vector is obtained by mapping based on the corrected center coordinate sequence.
[0011] Based on the three-dimensional acceleration sequence, kinematic state prediction is performed to obtain the prior estimated state quantity of the target object, and based on the measurement vector, posterior state update is performed on the prior estimated state quantity to obtain the optimal spatial position of the target object.
[0012] The continuously obtained optimal spatial positions are spliced in time to generate a reconstructed trajectory sequence of the target object, and feedback optimization parameters characterizing the confidence of the reconstructed trajectory sequence are obtained, and the preset sliding time window is updated based on them.
[0013] Secondly, the hybrid traffic flow multi-source sensing data fusion and trajectory reconstruction system includes the following modules:
[0014] The data acquisition module is used to acquire the three-dimensional acceleration sequence of the target object and its RGB video stream, and to extract the geometric center coordinate sequence of the target object from the RGB video stream;
[0015] The fundamental frequency extraction module is used to obtain the target fundamental frequency parameters characterizing the periodic changes in the kinematics of the target object based on the frequency domain energy distribution characteristics of the three-dimensional acceleration sequence within a preset sliding time window.
[0016] The filtering construction module is used to obtain the dynamic sampling rate of the RGB video stream within a preset sliding time window, and to construct temporal filtering coefficients for the morphological fluctuations of the target object in combination with the target fundamental frequency parameters;
[0017] The vector mapping module is used to perform time-domain filtering on the geometric center coordinate sequence based on the time-domain filtering coefficients to obtain a corrected center coordinate sequence, and to map a measurement vector based on the corrected center coordinate sequence.
[0018] The position generation module is used to perform kinematic state prediction based on the three-dimensional acceleration sequence to obtain the prior estimated state quantity of the target object, and to perform posterior state update on the prior estimated state quantity based on the measurement vector to obtain the optimal spatial position of the target object.
[0019] The feedback optimization module is used to temporally stitch together the continuously obtained optimal spatial positions to generate a reconstructed trajectory sequence of the target object, obtain feedback optimization parameters characterizing the confidence of the reconstructed trajectory sequence, and update the preset sliding time window based on them.
[0020] Thirdly, a computer storage medium stores computer-executable instructions, which, when executed, implement the hybrid traffic flow multi-source sensing data fusion and trajectory reconstruction method described in the first aspect.
[0021] Compared with the prior art, the beneficial effects of this application are:
[0022] By extracting the gait fundamental frequency from the inertial acceleration sequence across modes and mapping it to the visual sampling rate to construct temporal filtering coefficients, the geometric center coordinates are filtered to eliminate periodic deformation interference that is in sync with the gait frequency. This breaks the assumption that the geometric center and the physical centroid are rigidly coincident, and eliminates the sawtooth high-frequency oscillation of the Kalman fusion trajectory from the root cause.
[0023] Using the intensity of transient maneuvers in the reconstructed trajectory as a feedback optimization parameter, the sliding time window is dynamically and adaptively updated. This improves the frequency extraction accuracy during smooth travel and enhances the response speed during sudden maneuvers, which is beneficial for achieving stable convergence and maintaining accuracy of the fused trajectory under complex mixed traffic flows. Attached Figure Description
[0024] Figure 1 This is a schematic diagram illustrating the steps of the hybrid traffic flow multi-source sensing data fusion and trajectory reconstruction method of this application;
[0025] Figure 2 This is a schematic diagram of the modules of the hybrid traffic flow multi-source sensing data fusion and trajectory reconstruction system of this application. Detailed Implementation
[0026] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but only to illustrate selected embodiments of this application.
[0027] Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item has been defined in one figure, it does not need to be further defined and explained in subsequent figures. The terms first, second, etc. are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0028] In existing multi-source fusion localization systems for mixed traffic flows, a common approach is to use a deep learning-based target detection model to extract the 2D bounding box of a pedestrian from an RGB video stream, and then directly use the geometric center of this bounding box as the measurement input to a Kalman filter. This approach assumes that the geometric center is linearly aligned with the pedestrian's actual physical centroid and remains rigidly stationary. However, pedestrians inevitably exhibit regular strides and arm swings during normal walking, resulting in a periodic, breathing-like deformation of the visual bounding box contour with an amplitude of approximately 10cm to 20cm. This causes a pseudo-displacement of the geometric center relative to the physical centroid, which is strictly in sync with the gait cycle. When the system inputs this pseudo-displacement as a true measurement into the Kalman filter, its frequency characteristics significantly deviate from the zero-mean Gaussian white noise assumption required by the Kalman filter system. The filter misinterprets this as the target object's actual high-frequency maneuvering, triggering frequent misadjustments of the covariance matrix. Ultimately, the output fused trajectory exhibits a sawtooth-like high-frequency oscillation.
[0029] For example, in a mixed traffic flow scenario on urban roads, when a pedestrian walks naturally across the camera's field of view at a stride frequency of approximately 1.5Hz, the horizontal width and vertical height of the visual bounding box will periodically contract and expand at a frequency of 1.5Hz due to arm swings and stride movements. This causes the geometric center of the bounding box to shift periodically by approximately 3cm to 8cm relative to the pedestrian's hip joint center. If this geometric center is still used as the measurement input for the Kalman filter, the fused output position sequence will exhibit high-frequency sawtooth oscillations with a frequency of approximately 1.5Hz and an amplitude of approximately 0.43m superimposed on the actual translational trajectory, severely affecting the operation of downstream path planning and trajectory prediction modules.
[0030] If the above problems are not addressed, the high-frequency oscillations of the fused trajectory will continuously interfere with the downstream decision-making module's judgment of pedestrian intentions, potentially causing the autonomous driving system to make erroneous emergency braking or detours when encountering pedestrians. This not only reduces passenger comfort but also introduces safety risks such as rear-end collisions. Furthermore, the fixed spectral analysis time window cannot simultaneously consider frequency resolution and time response when pedestrian motion states dynamically switch, limiting the system's robustness under complex conditions. To address these issues, this application first considers how to break the rigid assumption that the geometric center is equivalent to the physical center of mass and establish a quantifiable deformation pseudo-displacement stripping mechanism.
[0031] This application creatively utilizes the vertical acceleration signal from an inertial measurement unit across modalities to extract the heel-grounding fundamental frequency of the target object as a feature parameter characterizing its gait period. This fundamental frequency is then mapped to a visual discrete sampling system to construct a digital notch filter, performing temporal filtering on the geometric center coordinates of the bounding box. Furthermore, addressing the issue of uneven performance under different motion states within a fixed time window, this application uses the intensity of transient maneuvers calculated based on the reconstructed trajectory sequence as a feedback optimization parameter to dynamically update the preset sliding time window, ultimately forming a complete technical solution that deeply integrates microscopic feature extraction, macroscopic coupled filtering, and closed-loop feedback optimization.
[0032] Therefore, such as Figure 1 As shown, this application provides a method for fusing multi-source sensing data and reconstructing trajectories of hybrid traffic flow, including the following steps:
[0033] The first step involves acquiring the three-dimensional acceleration sequence and its RGB video stream of the target object, and extracting the geometric center coordinate sequence of the target object from the RGB video stream. The three-dimensional acceleration sequence refers to the set of acceleration time-series signals synchronously acquired by an inertial measurement unit on three mutually orthogonal sensor coordinate axes. Specifically, this can be achieved using a MEMS accelerometer array worn on a pedestrian's waist or in a portable terminal, combined with timestamp calibration technology. Its function is to provide the raw physical input for subsequent extraction of the target's fundamental frequency parameters and kinematic state prediction.
[0034] An RGB video stream refers to a sequence of image frames continuously output by color cameras deployed at the edge nodes of a traffic scene. Specifically, it can be implemented using industrial cameras or roadside cameras operating at a fixed capture frame rate. Its purpose is to provide visual perception data for subsequent extraction of the geometric center coordinate sequence of target objects. The geometric center coordinate sequence is a two-dimensional vector sequence composed of the coordinates of the geometric center pixels of the two-dimensional bounding boxes corresponding to the target objects in the RGB video stream, arranged in temporal order. This can be achieved by using a deep learning-based object detection model to infer the coordinates of the center point of each video frame and calculate the center point coordinates of its output bounding box. Its role is to serve as the raw carrier of visual measurements and to provide processing objects for subsequent temporal filtering.
[0035] Step two involves obtaining the target fundamental frequency parameter, characterizing the periodic changes in the kinematics of the target object, based on the frequency domain energy distribution characteristics of the three-dimensional acceleration sequence within a preset sliding time window. The preset sliding time window refers to a dynamically adjustable time segment that moves forward along the time axis and continuously refreshes the data content. This can be achieved using a circular buffer queue combined with timestamp indexing, establishing a dynamic balance between frequency resolution and time response. The frequency domain energy distribution characteristics refer to the power distribution of the three-dimensional acceleration sequence at different frequency components after Fourier transform. This can be achieved by performing spectral analysis on the signal within the window using a Fast Fourier Transform algorithm. The target fundamental frequency parameter is a scalar frequency value that characterizes the periodicity of kinematic behaviors such as heel strike, stride, and arm swing during walking. This can be achieved by finding the global maximum of the power spectral density within the physically feasible step frequency range.
[0036] Step three involves obtaining the dynamic sampling rate of the RGB video stream within a preset sliding time window, and constructing temporal filtering coefficients for the morphological fluctuations of the target object based on the target fundamental frequency parameters. The dynamic sampling rate refers to the average frequency at which the RGB video stream actually generates valid image frames within the current preset sliding time window. Specifically, this can be achieved by counting the total number of video frames in which the target detection model successfully outputs valid bounding boxes within the window and dividing that number by the window duration. The temporal filtering coefficients are constant vectors used to drive the second-order infinite impulse response digital notch filter to perform temporal difference operations. Specifically, this can be achieved by substituting the target fundamental frequency parameters and the dynamic sampling rate into the digital angular frequency mapping formula and combining it with the pole-zero configuration principle of the preset pole radius factor.
[0037] Step four involves performing time-domain filtering on the geometric center coordinate sequence based on the time-domain filtering coefficients to obtain a corrected center coordinate sequence, and then mapping the measurement vector based on the corrected center coordinate sequence. The corrected center coordinate sequence refers to the pure coordinate sequence obtained by removing periodic deformation interference with the same frequency as the target fundamental frequency parameter from the geometric center coordinate sequence using a digital notch filter. Specifically, this can be achieved by decoupling the geometric center coordinate sequence spatially into independent x-axis and y-axis channel sequences and iteratively calculating them using second-order difference equations. The measurement vector refers to the observation data carrier formed by linear mapping of the corrected center coordinate sequence, which is necessary for the Kalman filter to perform posterior state updates. Specifically, this can be achieved by substituting the corrected center coordinates into a preset observation matrix and performing a projection operation.
[0038] Step 5 involves performing kinematic state prediction based on the three-dimensional acceleration sequence to obtain the prior estimated state quantity of the target object, and then performing a posterior state update on the prior estimated state quantity based on the measurement vector to obtain the optimal spatial position of the target object. The prior estimated state quantity refers to the target object state vector obtained solely from the three-dimensional acceleration sequence through kinematic integration, without incorporating visual measurement information for correction. Specifically, this can be achieved by performing a quadratic time integration on the three-dimensional acceleration sequence. The optimal spatial position refers to the target object's spatial coordinates that are probabilistically closest to the true position at the current observation time, obtained after comprehensively considering the uncertainties of both the prior estimated state quantity and the measurement vector. This can be achieved using a two-step prediction-update iteration using a classic Kalman filter system.
[0039] Step six involves temporally concatenating the continuously obtained optimal spatial positions to generate a reconstructed trajectory sequence of the target object. Feedback optimization parameters characterizing the confidence level of the reconstructed trajectory sequence are then obtained, and the preset sliding time window is updated based on these parameters. The reconstructed trajectory sequence refers to a set of spatiotemporal trajectory data of the target object, formed by temporally concatenating the optimal spatial positions generated at continuous observation times, and directly accessible to downstream modules. Specifically, this can be achieved by assigning a unique identifier to each target object and maintaining its own dedicated double-ended queue buffer. The feedback optimization parameters are scalar indices characterizing the intensity of the target object's transient maneuvering, extracted from the reconstructed trajectory sequence through discrete-time high-order difference operations. Specifically, this can be achieved by calculating the magnitude of the third-order time derivative of the optimal spatial position (i.e., the transient jerk vector) and taking a moving average.
[0040] The core innovation of this application lies in the following: by cross-modal frequency domain mapping, the gait fundamental frequency of the target object is extracted from the one-dimensional vertical acceleration signal of the inertial measurement unit, and the fundamental frequency is mapped to the visual discrete sampling system to construct a digital notch filter, and time-domain filtering is performed on the geometric center coordinates of the bounding box. This scheme breaks the rigid assumption that the geometric center is equivalent to the physical centroid, and fundamentally eliminates the periodic pseudo-displacement caused by gait deformation, so that the high-frequency sawtooth oscillation of the Kalman fusion trajectory is completely eliminated, and the adaptive dynamic adjustment of time-frequency parameters is realized through a feedback optimization mechanism based on the intensity of transient maneuvering.
[0041] I. The process of extracting the geometric center coordinate sequence;
[0042] At the hardware level, inertial measurement units (IMUs) typically output acceleration data continuously at a fixed sampling rate of around 100Hz. However, due to fluctuations in edge node computing power and the inference time of the target detection model, the effective output frame rate of RGB video streams often dynamically varies between 35Hz and 45Hz, resulting in significant differences in the timestamp tempo of the two types of data. To eliminate the phase error caused by these differences in subsequent frequency domain mapping, this step uses the capture timestamp of the video frame as a reference and performs nearest linear interpolation or cubic spline interpolation in the three-dimensional acceleration sequence. This ensures that the interpolated acceleration data corresponds one-to-one with the capture time of each video frame, thereby obtaining a three-dimensional acceleration sequence synchronized with the RGB video stream. Furthermore, since the sensor coordinate system of the IMU is not strictly aligned with the world coordinate system, the sensor attitude needs to be calculated based on the gravitational component distribution of the accelerometer in the initial static state. The three-axis acceleration is then projected from the sensor coordinate system to the world coordinate system using an attitude rotation matrix, ultimately separating the axial acceleration sequence perpendicular to the horizontal ground direction. This axial acceleration sequence can reflect the longitudinal mechanical impact caused by the pedestrian's heel hitting the ground to the greatest extent in a physical sense, and is the optimal data carrier for subsequent gait fundamental frequency extraction.
[0043] The preset target detection model can employ a lightweight target detection network based on convolutional neural networks (such as YOLOv5s), which has been pre-trained on a large-scale pedestrian dataset and deployed at edge nodes. Whenever a new RGB image frame arrives, the target detection model performs a single forward inference on that image, outputting the coordinates of the two-dimensional bounding boxes corresponding to all candidate targets in the image. Each two-dimensional bounding box is determined by the coordinates of its top-left and bottom-right corners, or equivalently by the coordinates of its center point and its width and height. For the specific target object currently being tracked, this step selects the corresponding two-dimensional bounding box from all candidate bounding boxes according to the target association matching algorithm, and calculates its geometric center pixel coordinates using the arithmetic mean of the top-left and bottom-right corner coordinates as the geometric center coordinates of the two-dimensional bounding box. Arranging the geometric center pixel coordinates calculated from consecutive video frames in chronological order yields the geometric center coordinate sequence, which will be used as the processing object for subsequent temporal filtering.
[0044] II. The process of extracting the target fundamental frequency parameters;
[0045] Because the axial acceleration sequence, when at rest, is superimposed with the DC component of gravitational acceleration and the low-frequency slowly varying component caused by sensor temperature drift, these components are concentrated around 0Hz in the spectrum, causing significant spectral leakage and masking in the 1.0Hz to 2.5Hz frequency band of interest corresponding to human walking. Therefore, this step inputs the axial acceleration sequence into a digital high-pass filter with a cutoff frequency of 0.5Hz to attenuate and remove the gravitational acceleration component and baseline drift component, obtaining a mean-free vertical acceleration sequence. This sequence mathematically rigorously reflects the high-frequency mechanical impact vibration caused by the heel striking the ground during walking. Based on this, the vertical acceleration sequence within a preset sliding time window is extracted, and a Fast Fourier Transform is performed on it to obtain a complex frequency domain sequence. The square of its amplitude is then taken and normalized according to the window length to obtain the power spectral density sequence corresponding to each discrete frequency component. The power spectral density sequence quantifies the energy contribution distribution of the vertical acceleration sequence at different frequencies in an energy sense, and is the core data foundation for subsequent extraction of the target fundamental frequency parameters.
[0046] Due to the physiological and kinesiological laws of the human body, the gait frequency of a normal pedestrian typically falls within the range of 1.0Hz to 2.5Hz. Frequency components outside this range cannot physically be generated by actual gait behavior. Therefore, this step sets this range as a preset gait frequency interval and performs peak search on the power spectral density sequence only within this interval to avoid misidentifying pseudo-peaks generated by other low-frequency swaying or high-frequency jitter as the gait fundamental frequency. Within this interval, the global maximum point of the power spectral density sequence is searched, and the power spectral density value corresponding to this maximum point is compared with a preset environmental noise floor reference value. The environmental noise floor reference value characterizes the background energy level of the three-dimensional acceleration sequence in this frequency band when the target object is standing still or hardly performing any stride behavior, and can be determined by offline system calibration. When the power spectral density value corresponding to the global maximum point is greater than the environmental noise floor reference value, it indicates that there is indeed a significant periodic vibration caused by the actual gait behavior within the current preset sliding time window. At this time, the frequency value corresponding to the global maximum point is taken as the target fundamental frequency parameter. Conversely, when the power spectral density value corresponding to the global maximum point is not greater than the environmental noise floor reference value, the target fundamental frequency parameter can be set to zero, which triggers the bypass pass-through logic of subsequent filtering processing.
[0047] III. The process of constructing time-domain filtering coefficients;
[0048] Since the effective output frame rate of the RGB video stream is affected by multiple factors such as fluctuations in edge node computing power, changes in object detection model inference time, and network transmission jitter, it is not a strictly fixed physical constant. If the nominal frame rate is directly used as the sampling rate for subsequent frequency domain mapping, phase errors that do not match the actual system state will be introduced. This step dynamically calculates the visual sampling rate through real-time statistics: between the start and end timestamps of the current preset sliding time window, the total number of video frames in which the object detection model successfully outputs valid bounding boxes is counted. This total number of video frames is divided by the time length of the preset sliding time window (in seconds) to obtain the dynamic sampling rate (in Hz) of the RGB video stream within the current window. This dynamic sampling rate strictly characterizes the actual processing rhythm of the visual perception node under the current system state, providing a benchmark for the subsequent accurate mapping of physical frequencies to discrete digital angular frequencies.
[0049] According to the fundamental principles of digital signal processing, the expression of frequency signals in the continuous physical world within a discrete sampling system is constrained by the Nyquist sampling theorem. Their continuous frequency values must be mapped to discrete digital angular frequencies before they can be incorporated into the computational system of the difference equations. When the target fundamental frequency parameter is not equal to zero (i.e., significant gait behavior does exist within the current window), the ratio obtained by dividing the target fundamental frequency parameter by the dynamic sampling rate is used as the normalized frequency, and then multiplied by... This yields the digital center angular frequency in the discrete domain (its physical meaning is the phase rotation of the discrete signal corresponding to the target fundamental frequency parameter within a unit sampling period), and its algebraic expression is: ;
[0050] in, Indicates the center angular frequency of the digital signal. This represents the target fundamental frequency parameter. This represents the dynamic sampling rate. This mapping operation mathematically and rigorously aligns the clock reference from the inertial physical time domain to the visual discrete sampling domain, ensuring that the zeros of the subsequent notch filter are precisely configured at the discrete frequency points corresponding to the physical gait fundamental frequency. When the target fundamental frequency parameter is equal to zero, the operation of this sub-step is skipped, and the bypass pass-through logic of the subsequent filtering process is triggered, that is, the corrected center coordinates are directly equal to the geometric center coordinates.
[0051] This step employs a classic second-order infinite impulse response digital notch filter as the mathematical model for the time-domain filter. The transfer function of this filter in the Z-domain is defined as the ratio of quadratic polynomials in both the numerator and denominator about the discrete unit delay operator. Specifically, the numerator polynomial has a pair of conjugate complex zeros at the unit circle corresponding to the digital center angular frequency, ensuring that the signal gain at that frequency is strictly zero. The denominator polynomial has a pair of conjugate complex poles in phase with the zeros but with slightly smaller magnitudes, thus ensuring the filter's stability and precisely controlling the stopband width. The closer the pole radius factor is to 1, the narrower the notch filter's stopband and the less interference it causes to signals outside the zero frequency. Based on this pole-zero configuration principle, the five one-dimensional constants used to drive the time-domain differential operation can be calculated from the digital center angular frequency and the preset pole radius factor. The eigenvector formed by these five constants is given by the following expression:
[0052] , , , , ;
[0053] in, , , , , These are the five one-dimensional constant terms that make up the eigenvector. Indicates the center angular frequency of the digital signal. This represents the preset pole radius factor. The preferred value of the preset pole radius factor is between 0.90 and 0.95, to balance the stopband width and transient response speed of the notch filter. If the value is too small (much less than 0.90), the stopband will be too wide, which will attenuate the real motion signal near the gait fundamental frequency. If the value is too large (too close to 1), the stopband will be too narrow, and the ability to cover the small fluctuations in human gait frequency will be insufficient.
[0054] IV. The process of obtaining the corrected center coordinate sequence;
[0055] Since the mathematical system of the second-order infinite impulse response digital notch filter is designed for one-dimensional scalar time series, directly performing global filtering on a two-dimensional vector would evolve into a complex multivariable matrix difference equation. Therefore, this step decouples the geometric center coordinate sequence, composed of two-dimensional pixel coordinates, into x-axis and y-axis channel sequences in space. This reduces the two-dimensional filtering problem to two independent one-dimensional filtering subproblems, perfectly fitting the standard digital filtering formula system. Physically, the movement of a pedestrian in a visual image is three-dimensional. The body deformation caused by gait manifests horizontally as lateral contour stretching due to strides and vertically as longitudinal contour contraction due to center-of-gravity fluctuations. Although both are governed by the gait fundamental frequency parameter, they have different amplitudes and phases. Therefore, decoupling them physically ensures that the two-axis filtering does not interfere with each other, avoiding trajectory distortion caused by severe y-axis jerking errors contaminating the relatively stable x-axis data.
[0056] After decoupling to obtain the x-axis channel sequence and the y-axis channel sequence, this step extracts the current sampling time from the current end of each sequence. Initial coordinates below And read the previous sampling time from the system cache queue. The first historical coordinates Compared with the previous two sampling times The second historical coordinates below These historical coordinates are continuously refreshed after the system starts up, with each frame of filtering operation completed, forming the input data context necessary for the filter to perform causal calculations.
[0057] The extracted initial coordinates, first historical coordinates, and second historical coordinates, along with the corrected coordinates from the previous sampling time and the two previous sampling times stored in the system buffer queue, and the obtained time-domain filter coefficients, are substituted into the difference equation of the second-order infinite impulse response digital notch filter for iterative calculation to obtain the corrected coordinates at the current sampling time. The difference equations for the x-axis and y-axis channels are given by the following expressions:
[0058] ;
[0059] ;
[0060] in, The current sampling time Corrected coordinates below , , These are the current sampling times. Previous sampling time The first two sampling times The initial coordinates below are the coordinates of the geometric center pixel before filtering. The previous sampling time Corrected coordinates below The first two sampling times Corrected coordinates below.
[0061] Once the difference equations complete the calculation of the corrected coordinates at the current sampling time, the calculated corrected coordinates, along with the current initial coordinates, are pushed to the tail of the system's buffer queue. Simultaneously, the oldest historical state at the head of the queue is popped, ensuring that the iterative calculations at the next sampling time have complete context input. During system initialization, since historical coordinates and historical corrected coordinates do not yet exist, all historical items can be preset to zero or the arithmetic mean of the initial coordinates. After the transient process of several sampling periods ends, the filter will enter steady-state operation. Finally, by recombining the corrected coordinates calculated at consecutive sampling times in chronological order, a corrected center coordinate sequence can be obtained. This sequence rigorously eliminates periodic deformation interference at the same frequency as the target's fundamental frequency parameters in the frequency domain, while losslessly preserving the true macroscopic translational signal and high-frequency random noise components of the target object, making it a high-confidence visual measurement carrier.
[0062] V. The process of obtaining prior calculated state variables and measurement vectors;
[0063] When performing posterior state updates, the Kalman filter requires measurement information provided by the observation device as input, and requires that this measurement information be linked to the system state variables through a linear mapping relationship, which is represented by a preset observation matrix. In the application scenario of this invention, the state vector of the target object contains two types of components: position and velocity. Since the visual perception node can only directly observe the position information of the target object, the specific form of the preset observation matrix is a selection matrix for extracting the position components from the state vector. Its function is to filter out the position dimensions that are comparable to the visual measurements from the complete multidimensional state vector. This step corrects the center coordinate sequence to the corrected coordinates at the current sampling time. Substituting the data into the preset observation matrix and performing a linear mapping operation, a two-dimensional measurement vector composed of visual position information is obtained. This measurement vector will be used as the actual observation carrier in the posterior state update stage and fed into the Kalman filter.
[0064] According to classical kinematic integral relations, the change in the position of a target object over time is equal to the integral of its velocity over time, and the change in its velocity is equal to the integral of its acceleration over time. Therefore, this step uses the optimal spatial position and optimal spatial velocity at the previous sampling time as initial conditions to perform a time integration operation on the three-dimensional acceleration sequence between the current sampling time and the previous sampling time, obtaining the predicted velocity component at the current sampling time; further, a time integration operation is performed on the predicted velocity component (or equivalently, a second integration is performed on the acceleration), obtaining the predicted position component at the current sampling time. In engineering implementation, since the three-dimensional acceleration sequence has been solved to the world coordinate system, the integration operation can be directly performed on its two directional components in the horizontal plane to obtain the predicted position and predicted velocity of the target object on the horizontal ground. The obtained predicted position component and predicted velocity component are arranged and combined according to the predefined order of the state vector to obtain the corresponding prior estimated state quantity. This prior estimated state quantity physically represents the target object state estimate obtained by dead reckoning relying solely on inertial sensors. Since the inertial integration error will accumulate and diverge over time, it must be corrected a posteriori by subsequent visual measurements.
[0065] VI. The process of obtaining the optimal spatial location;
[0066] The innovation residual is a core physical quantity in the Kalman filter system that measures the deviation between the current actual measurement and the system's prior prediction. This step first substitutes the prior calculated state variables into a preset observation matrix and performs a linear mapping operation to obtain the projection values of the prior calculated state variables in the observation space (i.e., the position coordinates of the target object that should have been observed by the visual sensor, predicted only based on inertial calculations). Then, the difference between the obtained measurement vector and this projection value is calculated, and the resulting difference sequence is the target innovation residual sequence. This sequence statistically reflects the direction and magnitude of the visual measurement's correction to the inertial prediction. Next, this step combines a preset prior error covariance and a preset measurement noise covariance to construct a gain matrix according to the standard Kalman gain calculation formula. The algebraic expression of the gain matrix is as follows:
[0067] ;
[0068] in, Indicates the current sampling time The gain matrix below, This represents the preset prior error covariance (which characterizes the system's degree of uncertainty about the prior estimated state variables). and These are the preset observation matrix and its transpose matrix (which transforms the state space and the observation space through mutual projection). This represents the preset measurement noise covariance (which characterizes the level of measurement error fluctuation of the visual perception node under the current operating conditions).
[0069] From a physics perspective, the value of the gain matrix reflects the system's dynamic trade-off between the reliability of actual measurements and prior predictions: when the confidence level of visual measurements is extremely high (i.e., ... When the value is very small, the gain matrix approaches the maximum correction mode that completely replaces the prior observation projection with the actual measurement, making the estimate of the optimal spatial location closely follow the visual measurement; conversely, when the confidence of the visual measurement is low (such as due to occlusion or decreased detection confidence), As the numerical value increases, the gain matrix approaches zero, making the optimal spatial location primarily dependent on prior calculations of state variables. In practical engineering implementations, the measurement noise covariance... It is not a fixed constant, but can be dynamically adjusted based on multiple indicators such as the detection confidence output by the target detection model, the intersection-union density of the two-dimensional bounding box and other targets in the scene, and the visual environment lighting conditions, thereby further improving the robustness of the system under complex working conditions.
[0070] The gain matrix obtained in this step A matrix multiplication operation is performed with the target information residual sequence to obtain the state correction vector at the current sampling time. This state correction vector physically represents the optimal correction amount provided by visual measurements that should be superimposed on the prior-calculated state quantities. It is worth noting that, due to the gain matrix... The matrix is a rectangular matrix (the number of rows equals the dimension of the state vector, and the number of columns equals the dimension of the observation vector). After the two-dimensional target information residual sequence is amplified and distributed by this matrix, it will be transformed into a complete state space dimension correction vector. This vector contains not only the correction component in the position dimension but also the correction component in the velocity dimension. This is because the prior error covariance matrix implicitly stores the covariance correlation between the position uncertainty and the velocity uncertainty in its structure, so that the position and velocity state components can be corrected simultaneously by relying only on the input of position measurement information. This is the most ingenious mathematical property of the Kalman filter system.
[0071] Furthermore, the aforementioned state correction vector is summed with the prior calculated state quantity to obtain the posterior optimal state quantity at the current sampling moment. This posterior optimal state quantity comprehensively considers the uncertainties of both inertial prediction and visual measurement, and is the complete state description of the target object that is probabilistically closest to the true value at the current observation moment. Extracting the two components in the position dimension yields the optimal spatial position of the target object; simultaneously, the two components in the velocity dimension constitute the optimal spatial velocity. Both are stored in the system state cache and used as the initial conditions for performing kinematic state prediction at the next sampling moment.
[0072] Through the two-step iterative process of Kalman filtering described above, the prior calculated state variables are corrected a posteriori based on the visual measurement vector purified by notch filtering. Since the target information residual sequence at this time has eliminated the periodic deformation interference that is in sync with the gait, its statistical characteristics are highly consistent with the zero-mean Gaussian white noise assumption required by the Kalman filtering system. This avoids frequent misadjustment of the covariance matrix caused by pseudo-displacement, and ultimately makes the optimal spatial position sequence of the target object closely track the real motion trajectory on a macroscopic level, and strictly eliminates the high-frequency sawtooth oscillation caused by gait deformation pseudo-displacement on a microscopic level.
[0073] VII. The update process of the preset sliding time window;
[0074] Based on the principle of calculating higher-order derivatives of discrete data in numerical analysis, the third-order time derivative of a discrete position sequence sampled at equal time intervals can be accurately calculated using the four-point backward difference formula. This step extracts the current sampling time from the reconstructed trajectory sequence in chronological order. Its optimal spatial location at the previous three sampling times is denoted as follows: , , , The transient vector of the target object at the current sampling time is calculated according to the following expression. :
[0075] ;
[0076] in, Indicates the target object at the current sampling time The transient vector below; , , , These represent the current sampling time. Previous sampling time The first two sampling times The first three sampling times The optimal spatial location; This represents the time interval between adjacent sampling times.
[0077] From a physical perspective, the transient vector, often referred to as the jerk vector in cybernetics, is essentially the instantaneous rate of change of an object's acceleration. It is the most sensitive physical indicator for measuring whether an object is undergoing sudden stops, starts, sharp turns, or other violent maneuvers. Under normal, stable driving conditions, a pedestrian's acceleration changes gradually, and the transient vector amplitude is very small. However, when a pedestrian undergoes sudden, violent maneuvers such as evasive maneuvers, the transient vector amplitude increases significantly and instantaneously.
[0078] Because transient vectors at a single sampling moment are susceptible to misjudgment due to instantaneous sensor noise, this step employs a moving average method to smooth the amplitude of transient vectors over several consecutive sampling moments, thereby obtaining stable and reliable feedback optimization parameters. Specifically, the Euclidean norm (i.e., amplitude) of the transient vector calculated at each sampling moment is obtained, and the resulting amplitude is represented as a scalar value. This amplitude scalar completely removes the directional information of the transient vector in a physical sense, retaining only its intensity information. This allows for the unified quantification and evaluation of violent maneuvers in different directions, avoiding evaluation distortion caused by the mutual cancellation of motions in different directions. Subsequently, the amplitudes of a predetermined number of continuously generated transient vectors (e.g., 10 sampling periods) are extracted, and an arithmetic average is performed on them. The resulting average value is the feedback optimization parameter characterizing the intensity of the transient maneuver of the target object. This feedback optimization parameter, in a physical sense, strictly reflects the overall intensity of the target object's motion within the recent time window; it is a non-negative scalar indicator.
[0079] Given that floating-point exponentiation is a computationally intensive basic operation, and that the feedback optimization parameters of the target object remain at extremely low levels under static or extremely stable movement conditions, making window length adjustments insignificant, this step introduces a preset update threshold (i.e., a steady-state dead zone threshold) as the trigger condition for window updates. Only when the feedback optimization parameters exceed the preset update threshold is the exponential mapping operation of the window length performed; otherwise, the window length remains unchanged at the preset upper limit of the time window. This dead-zone control mechanism not only significantly reduces the computational overhead of the system under normal operating conditions but also avoids invalid window length jitter caused by sensor noise. When the feedback optimization parameters do indeed exceed the preset update threshold, the updated preset sliding time window is calculated according to the following exponential mapping function. :
[0080] ;
[0081] in, This indicates the updated preset sliding time window; This indicates the maximum value of the preset time window; This indicates the lower limit of the preset time window; It is a natural constant; This is a preset decay adjustment coefficient (its value is greater than zero, used to control the rate of exponential decay). This represents the feedback optimization parameters; This indicates the preset update threshold.
[0082] From the perspective of dynamic control logic, the above exponential mapping function has the following three excellent properties: First, when the feedback optimization parameter is exactly equal to the preset update threshold, the exponential term... The exponent is zero, and the function value is strictly equal to the upper limit of the time window, ensuring perfect continuity in window switching between steady-state and maneuvering conditions without any jumps. Secondly, as the feedback optimization parameters gradually increase (i.e., the target object maneuvers more violently), the exponent term decays rapidly, and the updated preset sliding time window shortens exponentially and smoothly converges to the lower limit of the time window, enabling the system to respond to changes in motion state and obtain the latest gait frequency characteristics at an extremely fast speed in sudden maneuvering scenarios. Thirdly, the value of the preset decay adjustment coefficient controls the sensitivity of the window to changes in the intensity of random maneuvering—the larger the value, the more sensitive the window is to the response of the feedback optimization parameters, and vice versa.
[0083] The updated preset sliding time window is applied to subsequent tasks; that is, in subsequent calculations such as target fundamental frequency parameter extraction and dynamic sampling rate calculation, the updated window length is uniformly used as the new sliding time window length. This establishes a complete, trajectory quality-driven, multi-scale time-frequency adaptive closed-loop feedback mechanism.
[0084] When the target object is in a stable moving condition, the feedback optimization parameters are kept at a low level, the preset sliding time window is extended to the upper limit, and the extraction accuracy of the target fundamental frequency parameters is maximized; when the target object undergoes a sudden maneuver, the feedback optimization parameters increase instantaneously, the preset sliding time window is shortened exponentially to the lower limit, and the dynamic response speed of the system is maximized; thus, an intelligent trade-off based on motion state is achieved between the contradictory physical rigid constraints of frequency resolution and time response.
[0085] In another implementation, such as Figure 2 As shown, this application also provides a system for multi-source sensing data fusion and trajectory reconstruction of hybrid traffic flow, including the following modules:
[0086] The data acquisition module is used to acquire the three-dimensional acceleration sequence of the target object and its RGB video stream, and extract the geometric center coordinate sequence of the target object from the RGB video stream. At the hardware level, the data acquisition module includes an inertial measurement unit associated with the target object and a color camera deployed in the traffic scene, and integrates a target detection inference submodule for extracting two-dimensional bounding boxes from the RGB video stream; its function is to provide raw physical input and visual perception data for subsequent module processing.
[0087] The fundamental frequency extraction module is used to obtain the target fundamental frequency parameters characterizing the periodic changes in the kinematics of the target object based on the frequency domain energy distribution characteristics of the three-dimensional acceleration sequence within a preset sliding time window. The fundamental frequency extraction module includes an attitude calculation submodule, a high-pass filtering submodule, and a spectral peak detection submodule; its function is to accurately lock the gait fundamental frequency of the target object from the one-dimensional vertical acceleration signal, providing a center frequency reference value for subsequent construction of a digital notch filter.
[0088] The filtering construction module is used to obtain the dynamic sampling rate of the RGB video stream within a preset sliding time window, and to construct temporal filtering coefficients for the morphological fluctuations of the target object in combination with the target fundamental frequency parameters. The filtering construction module includes a frame rate statistics submodule and a coefficient calculation submodule; its function is to legally map the gait frequency in physical space to the digital angular frequency in the discrete visual sampling system, and to calculate the five one-dimensional constant terms of the second-order infinite impulse response digital notch filter.
[0089] The vector mapping module is used to perform temporal filtering on the geometric center coordinate sequence based on the temporal filtering coefficients to obtain a corrected center coordinate sequence, and to map the obtained measurement vector based on the corrected center coordinate sequence. The vector mapping module includes a coordinate decoupling submodule, a difference iteration submodule, and an observation projection submodule; its function is to accurately remove periodic pseudo-displacements caused by gait deformation in the geometric center coordinate sequence, and to map the obtained corrected center coordinate sequence into the measurement vector required by the Kalman filter.
[0090] The position generation module is used to perform kinematic state prediction based on the three-dimensional acceleration sequence to obtain the prior estimated state variables of the target object, and to perform posterior state update on the prior estimated state variables based on the measurement vector to obtain the optimal spatial position of the target object. The position generation module includes a kinematic integration submodule, an innovation residual calculation submodule, a gain matrix construction submodule, and a posterior state update submodule; its function is to achieve optimal fusion of prior prediction and visual measurement through a classic two-step iterative process of Kalman filtering.
[0091] The feedback optimization module is used to temporally stitch together the continuously obtained optimal spatial positions to generate a reconstructed trajectory sequence of the target object, obtain feedback optimization parameters characterizing the confidence of the reconstructed trajectory sequence, and update the preset sliding time window based on these parameters. The feedback optimization module includes a trajectory stitching submodule, a transient vector calculation submodule, an amplitude averaging submodule, and a window exponential mapping submodule; its function is to construct a multi-scale time-frequency adaptive closed-loop feedback mechanism driven by the quality of the reconstructed trajectory.
[0092] In another embodiment, this application also provides a computer storage medium storing computer-executable instructions, which, when executed, implement the hybrid traffic flow multi-source sensing data fusion and trajectory reconstruction method.
[0093] The computer storage medium includes, but is not limited to, disk storage, optical disk storage, flash memory, solid-state drives, hard disk drives, random access memory, and other forms of industrial-grade storage devices. The processor can be implemented using a central processing unit (CPU), embedded microcontroller, digital signal processor, or application-specific integrated circuit (ASIC) in an edge computing node, or it can be implemented using a combination of a general-purpose processor and a graphics processor in a cloud server cluster. When the processor loads and executes the computer-executable instructions, it can sequentially complete the entire process of operations, including acquiring three-dimensional acceleration sequences and RGB video streams, extracting geometric center coordinate sequences, extracting target fundamental frequency parameters, calculating dynamic sampling rates, constructing temporal filtering coefficients, acquiring corrected center coordinate sequences, mapping measurement vectors, acquiring prior calculated state quantities, acquiring optimal spatial positions, generating reconstructed trajectory sequences, and updating preset sliding time windows.
[0094] From the perspective of ergonomics and urban public transportation design, this method significantly reduces high-frequency sawtooth oscillations in the fused trajectory by stripping away gait deformation pseudo-displacements and dynamically adapting to the time-frequency characteristics of pedestrian movement. This makes downstream pedestrian behavior analysis (such as gait phase recognition and avoidance intention prediction) more accurate and reliable. This method can be widely applied to the research and development and implementation of urban intelligent transportation sensing equipment (such as roadside intelligent cameras and multi-source sensor integrated machines), outdoor public transportation facilities (such as bus stops and pedestrian crossing early warning systems), and pedestrian safety assistance intelligent products (such as wearable navigation terminals for the blind and shared mobility terminals), providing design support for improving the smoothness and stability of pedestrian trajectory reconstruction in mixed traffic flow environments.
[0095] The above embodiments are only used to illustrate the technical methods of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of this application without departing from the spirit and scope of the technical methods of this application.
Claims
1. A method for fusing and reconstructing trajectory data from multiple sources of mixed traffic flow, characterized in that, Includes the following steps: Obtain the three-dimensional acceleration sequence of the target object and its RGB video stream, and extract the geometric center coordinate sequence of the target object from the RGB video stream; Based on the frequency domain energy distribution characteristics of the three-dimensional acceleration sequence within a preset sliding time window, the target fundamental frequency parameter characterizing the periodic kinematic changes of the target object is obtained; The dynamic sampling rate of the RGB video stream within a preset sliding time window is obtained, and a temporal filtering coefficient for the morphological fluctuation of the target object is constructed in combination with the target fundamental frequency parameter. The geometric center coordinate sequence is subjected to time-domain filtering based on the time-domain filtering coefficients to obtain a corrected center coordinate sequence, and a measurement vector is obtained by mapping based on the corrected center coordinate sequence. Based on the three-dimensional acceleration sequence, kinematic state prediction is performed to obtain the prior estimated state quantity of the target object, and based on the measurement vector, posterior state update is performed on the prior estimated state quantity to obtain the optimal spatial position of the target object. The continuously obtained optimal spatial positions are spliced in time to generate a reconstructed trajectory sequence of the target object, and feedback optimization parameters characterizing the confidence of the reconstructed trajectory sequence are obtained, and the preset sliding time window is updated based on them. The process of obtaining the target fundamental frequency parameters includes: Timestamp matching interpolation is performed on the three-dimensional acceleration sequence to obtain a three-dimensional acceleration sequence that is time-synchronized with the RGB video stream, and coordinate system solution is performed on it to separate the axial acceleration sequence perpendicular to the horizontal ground direction; The axial acceleration sequence is filtered to obtain the vertical acceleration sequence after removing its gravitational acceleration component. The vertical acceleration sequence within a preset sliding time window is subjected to Fourier transform to obtain the power spectral density sequence corresponding to each frequency. Global maxima are extracted from the power spectral density sequence within a preset gait frequency range. When the power spectral density value corresponding to the global maxima is greater than the preset environmental noise floor reference value, the frequency value corresponding to the power spectral density sequence of the global maxima is taken as its target fundamental frequency parameter.
2. The method for fusion and trajectory reconstruction of multi-source sensing data of hybrid traffic flow according to claim 1, characterized in that, The process of extracting the geometric center coordinate sequence includes: The continuous video frames of the RGB video stream are input into a preset target detection model to output a two-dimensional bounding box of the target object. The geometric center pixel coordinates of the two-dimensional bounding box are extracted to form a geometric center coordinate sequence.
3. The method for fusion and trajectory reconstruction of multi-source sensing data of hybrid traffic flow according to claim 1, characterized in that, The process of constructing time-domain filter coefficients includes: The total number of video frames contained in the RGB video stream within a preset sliding time window is obtained, and then divided by the time length of the preset sliding time window to obtain the dynamic sampling rate. When the target fundamental frequency parameter is not equal to zero, the digital center angular frequency in the discrete domain is obtained based on the target fundamental frequency parameter and the dynamic sampling rate. ; Based on the digital center angular frequency and the preset pole radius factor, an eigenvector consisting of five one-dimensional constant terms is obtained and used as the time-domain filtering coefficient. , , , , ; in, This represents the target fundamental frequency parameter. This represents the dynamic sampling rate. , , , , These are the five one-dimensional constant terms that make up the eigenvector. This represents the preset pole radius factor.
4. The method for fusion and trajectory reconstruction of multi-source sensing data in hybrid traffic flow according to claim 3, characterized in that, The process of obtaining the corrected center coordinate sequence includes: The geometric center coordinate sequence is decoupled spatially into independent x-axis channel sequences and y-axis channel sequences, and the current sampling time is extracted based on both. Initial coordinates below and its previous sampling time The first historical coordinates below Compared with the previous two sampling times The second historical coordinates below ; Get the current sampling time Corrected coordinates below The corrected coordinates at consecutive sampling times are temporally reassembled to obtain the corrected center coordinate sequence; ; ; in, The previous sampling time Corrected coordinates below The first two sampling times Corrected coordinates below.
5. The method for fusion and trajectory reconstruction of multi-source sensing data of hybrid traffic flow according to claim 1, characterized in that, The process of obtaining prior-derived state variables includes: The corrected center coordinate sequence is substituted into the preset observation matrix to perform a linear mapping operation to generate a measurement vector; The three-dimensional acceleration sequence is integrated over time to obtain the predicted position and velocity components of the target object at the current sampling time, and the two are combined into the corresponding prior state variables.
6. The method for fusion and trajectory reconstruction of multi-source sensing data of hybrid traffic flow according to claim 5, characterized in that, The process of obtaining the optimal spatial location includes: The difference between the measurement vector and the projection of the prior calculated state variable in the observation space is calculated to obtain the target information residual sequence, and a gain matrix is constructed by combining it with the preset prior error covariance. ; in, This represents the pre-defined prior error covariance. and The preset observation matrix and its transpose are given. This represents the preset measurement noise covariance; The gain matrix The product of the target information residual sequence and the prior calculated state variables are summed to obtain the posterior optimal state variables containing the optimal spatial position and optimal spatial velocity.
7. The method for fusion and trajectory reconstruction of multi-source sensing data of hybrid traffic flow according to claim 1, characterized in that, The process of updating the preset sliding time window includes: Based on the reconstructed trajectory sequence, extract information including the current sampling time. Optimal spatial location at multiple consecutive sampling times. And obtain the transient vector of the target object at the current sampling time. ; in, , , These represent the current sampling time. Previous sampling time The first two sampling times The first three sampling times The optimal spatial location is as follows. The time interval between adjacent sampling times; The average value of the amplitudes of a predetermined number of continuously generated transient vectors is obtained as a feedback optimization parameter characterizing the severity of the transient maneuver of the target object. When the feedback optimization parameter is greater than the preset update threshold At that time, obtain the updated preset sliding time window. ; ; in, The preset attenuation adjustment coefficient, , These represent the upper and lower limits of the preset time window, respectively, and the updated preset sliding time window will be applied to subsequent tasks.
8. A system for fusing and reconstructing trajectory data from multiple sources of mixed traffic flow, characterized in that: Includes the following modules: The data acquisition module is used to acquire the three-dimensional acceleration sequence of the target object and its RGB video stream, and to extract the geometric center coordinate sequence of the target object from the RGB video stream; The fundamental frequency extraction module is used to obtain the target fundamental frequency parameters characterizing the periodic changes in the kinematics of the target object based on the frequency domain energy distribution characteristics of the three-dimensional acceleration sequence within a preset sliding time window. The filtering construction module is used to obtain the dynamic sampling rate of the RGB video stream within a preset sliding time window, and to construct temporal filtering coefficients for the morphological fluctuations of the target object in combination with the target fundamental frequency parameters; The vector mapping module is used to perform time-domain filtering on the geometric center coordinate sequence based on the time-domain filtering coefficients to obtain a corrected center coordinate sequence, and to map a measurement vector based on the corrected center coordinate sequence. The position generation module is used to perform kinematic state prediction based on the three-dimensional acceleration sequence to obtain the prior estimated state quantity of the target object, and to perform posterior state update on the prior estimated state quantity based on the measurement vector to obtain the optimal spatial position of the target object. The feedback optimization module is used to temporally stitch together the continuously obtained optimal spatial positions to generate a reconstructed trajectory sequence of the target object, obtain feedback optimization parameters characterizing the confidence of the reconstructed trajectory sequence, and update the preset sliding time window based on them. The process of obtaining the target fundamental frequency parameters includes: Timestamp matching interpolation is performed on the three-dimensional acceleration sequence to obtain a three-dimensional acceleration sequence that is time-synchronized with the RGB video stream, and coordinate system solution is performed on it to separate the axial acceleration sequence perpendicular to the horizontal ground direction; The axial acceleration sequence is filtered to obtain the vertical acceleration sequence after removing its gravitational acceleration component. The vertical acceleration sequence within a preset sliding time window is subjected to Fourier transform to obtain the power spectral density sequence corresponding to each frequency. Global maxima are extracted from the power spectral density sequence within a preset gait frequency range. When the power spectral density value corresponding to the global maxima is greater than the preset environmental noise floor reference value, the frequency value corresponding to the power spectral density sequence of the global maxima is taken as its target fundamental frequency parameter.
9. A computer storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, they implement the method for multi-source sensing data fusion and trajectory reconstruction of hybrid traffic flow as described in any one of claims 1-7.
Citation Information
Patent Citations
Gait sub-phase prediction method and device based on deep learning model fusion
CN119700094A
Unmanned aerial vehicle flight path positioning method and system based on multi-source information fusion
CN121252823A