Storage Compression and Dynamic Synchronization Method for Multi-Source Heterogeneous Vehicle Data

By preprocessing, time-frequency conversion, clock calibration and feature matching of multi-source heterogeneous driving data, data inconsistency problem is solved, high-precision data synchronization and storage are achieved, and data reliability and consistency are improved.

CN119862184BActive Publication Date: 2025-07-04JARVIS INTELLIGENCE (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510345202.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-04
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

There are inconsistencies in the storage and synchronization process of multi-source heterogeneous driving data, resulting in low accuracy and reliability of data synchronization, making it difficult to control data synchronization errors, affecting the data alignment accuracy and consistency in time.

Method used

By preprocessing the original driving data, adaptively selecting the time-frequency transformation method to mine redundant information, calibrate the clock using the clock synchronization algorithm, establish feature matching relationships, perform data fine-tuning, and store and update data in the multi-source heterogeneous driving database.

Benefits of technology

It improves the accuracy and consistency of data in time, enhances the stability and reliability of data synchronization, reduces storage space requirements, and ensures accurate synchronization of data in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862184B_ABST
    Figure CN119862184B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of driving data processing, and specifically discloses a method for storing, compressing, and dynamically synchronizing multi-source heterogeneous driving data. First, the original driving data is preprocessed to convert it into a unified format; according to the characteristics of the data, a time-frequency transformation method is adaptively selected to deeply mine the redundant information of the data, and the mined redundant information is quantized; then, a clock synchronization algorithm is used to accurately calibrate the clocks of each data source. Through the adaptive time-frequency transformation, the present invention can better capture the characteristics of the data, more accurately mine the redundant information in the frequency domain, further improve the compression ratio on the premise of ensuring data quality, and reduce the storage space required; by adopting the clock synchronization algorithm, the error of data synchronization can be controlled within a smaller range, improving the alignment accuracy of the data in time and ensuring higher consistency of the data from different data sources in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle data processing, and in particular relates to a storage compression and dynamic synchronization method for multi-source heterogeneous vehicle data. Background Art

[0002] Multi-source heterogeneous driving data is data related to vehicle driving that is collected from a variety of different types of data sources and has different data structures and formats. These data sources include but are not limited to vehicle-mounted sensors, driving recorders, navigation systems, roadside sensors in intelligent transportation systems, and data from other vehicles or infrastructure obtained by the Internet of Vehicles platform. In order to effectively store and transmit these large-scale data, these data need to be compressed.

[0003] At present, the storage and compression method for multi-source heterogeneous driving data is usually to convert the driving data from the time domain to the frequency domain, compress it using the redundant information in the frequency domain, and then synchronize and update the data of each data source according to the set time interval; however, since multi-source heterogeneous data may come from different sensors and devices, there are differences in their collection methods, sampling rates and accuracy, which may lead to inconsistency problems in the data compression process, resulting in low accuracy and reliability of data synchronization. Therefore, we need to propose a storage compression and dynamic synchronization method for multi-source heterogeneous driving data to solve the above problems, so that it can control the error of data synchronization within a smaller range, improve the accuracy of data alignment in time, ensure that the data from different data sources are more consistent in time, and improve the accuracy and reliability of data synchronization. Summary of the invention

[0004] The purpose of the present invention is to provide a storage compression and dynamic synchronization method for multi-source heterogeneous driving data, to control the error of data synchronization within a smaller range, to improve the temporal alignment accuracy of data, to ensure that data from different data sources are more consistent in time, to improve the accuracy and reliability of data synchronization, so as to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The storage compression and dynamic synchronization method of multi-source heterogeneous driving data includes the following steps:

[0007] S1. Preprocessing the original driving data to convert the original driving data into a unified format;

[0008] S2. According to the characteristics of the data, the time-frequency transformation method is adaptively selected to deeply mine the redundant information of the data, and the mined redundant information is quantified;

[0009] S3. Use the clock synchronization algorithm to precisely calibrate the clocks of each data source;

[0010] S4. Extract representative feature data from the data of different data sources;

[0011] S5. Establish the correspondence between the data of different data sources by calculating the similarity between different features for feature matching;

[0012] S6. Fine-tune the data according to the feature matching results to ensure the accurate synchronization of the data in time;

[0013] S7. Establish a multi-source heterogeneous vehicle operation database and store the compressed data on different nodes;

[0014] S8. Establish metadata information for the data of each data source to facilitate querying, retrieving, and managing the data;

[0015] S9. Set a data update trigger mechanism and update the data in the database.

[0016] Preferably, in step S1, the process of preprocessing the original vehicle operation data is as follows:

[0017] S11. Perform moving average filtering on the original vehicle operation data to obtain the filtered data. The formula for moving average filtering is as follows:

[0018] ,

[0019] Where, And , is the number of time series data, is the number of data points participating in the moving average calculation, is the filtering result after moving average filtering at position , is the j-th data point in the original time series data, is the position index of the current data point in the entire data sequence;

[0020] S12. Sort the filtered data from smallest to largest, and then calculate the positions of the quartiles respectively. The calculation formula is as follows:

[0021] ,

[0022] , where, is the position of the upper quartile, is the position of the lower quartile, is the number of the filtered data;

[0023] S13. Calculate the lower quartile according to the positions of the quartiles and the upper quartile , and the calculation formulas are as follows:

[0024] If and are integers, then , ;

[0025] If or is not an integer, then set or to have a decimal part of a and an integer part of b, and then use the quartile calculation formula to calculate and respectively. The calculation formula is:

[0026] , where is the lower quartile or the upper quartile , and are the data values at the th and th positions after sorting respectively;

[0027] S14. Subtract the lower quartile from the upper quartile to obtain the interquartile range IQ, set the normal data range to , and delete the data outside the normal data range;

[0028] S15. Use the FFmpeg library function to uniformly convert the video files into the mp4 format;

[0029] S16. Perform linear normalization on the numerical data to convert the data into a unified format. The normalization formula is as follows:

[0030] , where is the output value after normalization, is the numerical data, is the minimum value of the numerical data, is the maximum value of the numerical data.

[0031] Preferably, in step S2, the process of mining redundant information of the data is as follows:

[0032] A1. Obtain frequency domain information data through Fourier transform or wavelet transform;

[0033] A2. Calculate the mean of the frequency-domain coefficients corresponding to different frames in the frequency domain. The mean calculation formula is:

[0034] , where M is the total number of data frames, is the frequency-domain coefficient of the -th frame, is the frequency-domain coordinate, is the mean value at coordinate ;

[0035] A3. Calculate the variance at the coordinate according to the mean . The variance calculation formula is:

[0036] , where is the variance at the coordinate ;

[0037] A4. Set the quantization interval as [c, d] and the quantization level as L. Calculate the quantization step Log. The calculation formula is:

[0038] , where is the lower limit value of the quantization interval, is the upper limit value of the quantization interval;

[0039] A5. Calculate the quantized output value according to the quantization step Log. The calculation formula is:

[0040] , where is the frequency-domain coefficient, is the operation of mapping continuous frequency-domain coefficients to a finite number of discrete values, is the lower limit value of the quantization interval, is the quantized output value.

[0041] Preferably, in step S3, the process of calibrating the clocks of each data source is as follows:

[0042] S31. Initialize the clocks of each data source according to the function of the data source's clock as the master clock or the slave clock;

[0043] S32. The master clock periodically sends synchronization messages to the slave clocks in the network. The synchronization messages include the local timestamp when the master clock sends the synchronization message;

[0044] S33. When the slave clock receives the synchronization message, record the local time when the message is received;

[0045] S34. After receiving the synchronization message, the slave clock sends a delay request message to the master clock and records the local time when the delay request message is sent. ;

[0046] S35. After receiving the delay request message, the master clock sends a delay response message to the slave clock, and the delay response message contains the local time of the master clock when the delay request message is received. ;

[0047] S36. The slave clock calculates the clock offset Offset and network transmission delay Delay between the slave clock and the master clock according to the received timestamp information. The calculation formulas are as follows:

[0048] ,

[0049] ;

[0050] S37. The slave clock adjusts its local clock according to the calculated clock offset Offset to synchronize the slave clock with the master clock.

[0051] Preferably, in step S4, the process of feature data extraction is as follows:

[0052] S41. Determine the specific moment for which features are to be extracted according to the timestamp of the data. The determination of the specific moment satisfies , where is the time series of known data, is the data for which the specific moment is to be selected, is the index for traversing and summing the data points near the specific moment ;

[0053] S42. Arrange the data near the specific moment in ascending order, calculate the median of the arranged data, and take the median as the feature data.

[0054] Preferably, in step S5, the process of feature matching is as follows:

[0055] S51. Centered on the specific moment , set the time window to be , where is the size of the time window;

[0056] S52. Calculate the similarity between the feature data of different data sources according to the feature data of different data sources within the time window. The similarity calculation formula is:

[0057] , where r is the dimension of the feature vector, is the i-th element of the feature vector A, is the i-th element of the feature vector B, is the similarity between the feature vector A and the feature vector B;

[0058] S53. According to the calculated similarity value, select the feature data pairs with high similarity to establish the corresponding relationship between the feature data pairs.

[0059] Preferably, in step S6, the process of micro-adjusting the data is as follows:

[0060] S61. According to the result of feature matching, compare the timestamps of the data from different data sources and calculate the time deviation between different data sources;

[0061] S62. Make adjustments according to the time deviation to minimize the time deviation between different data sources, and update the data records of the relevant data sources. When the time deviation is small, use the interpolation method for adjustment; when the time deviation is large, it is necessary to further check the data accuracy or re-perform feature matching.

[0062] Preferably, in step S7, the process of establishing the multi-source heterogeneous vehicle driving database is as follows:

[0063] S71. Determine the overall architecture of the multi-source heterogeneous vehicle driving database, and the overall architecture includes the storage method and the data organization form;

[0064] S72. Determine the number of nodes and node configurations according to the characteristics and scale of the data;

[0065] S73. Store the compressed data on different nodes according to the planned architecture and storage method.

[0066] Preferably, in step S8, when establishing the metadata information for the data of each data source, first define a detailed metadata structure for the data of each data source, then collect the metadata information of each data source during the data acquisition or storage process, and finally use a relational database to store the collected metadata information in the metadata management module of the database to ensure an effective association relationship is established between the metadata and the actual data.

[0067] Preferably, in step S9, the process of setting the data update trigger mechanism is as follows:

[0068] S91. According to the application requirements for real-time vehicle status monitoring, driving behavior analysis or intelligent transportation management, determine the frequency and timeliness requirements for data update;

[0069] S92. Set the conditions for event triggering, and monitor the occurrence of relevant events in real time. Determine whether the event triggering conditions are met. When the event triggering conditions are not met, continue to monitor the relevant events in real time. When the event triggering conditions are met, obtain the latest data from each data source;

[0070] S93. Clean and transform the obtained latest data, and then synchronize the updated data in chronological order according to the timestamp information of the data, ensuring that the data is inserted into the database in the normal chronological order. If there are deviations in the timestamps of the data from different data sources, calibration and adjustment are required to minimize the deviation;

[0071] S94. Insert the synchronized updated data into the corresponding location in the database. If there are duplicate or conflicting data when inserting the data, overwrite the duplicate or conflicting data.

[0072] The method for storing, compressing and dynamically synchronizing multi-source heterogeneous vehicle driving data proposed by the present invention has the following advantages compared with the prior art:

[0073] 1. Through the adaptive time-frequency transformation of the present invention, the characteristics of the data can be better captured, the redundant information in the frequency domain can be more accurately mined, the compression ratio can be further improved on the premise of ensuring the data quality, and the storage space required can be reduced; by adopting the clock synchronization algorithm, the error of data synchronization can be controlled within a smaller range, the alignment accuracy of the data in time can be improved, and the consistency of the data from different data sources in time can be ensured to be higher.

[0074] 2. Through the mechanism of feature matching of the present invention, the data synchronization problem caused by factors such as sensor acquisition differences and network transmission delays can be overcome to a certain extent, the stability and reliability of data synchronization are enhanced, and accurate data synchronization can be ensured even in a complex vehicle driving environment. Description of the Drawings

[0075] Figure 1 Shows a flowchart according to an embodiment of the present invention;

[0076] Figure 2 Shows a flowchart for preprocessing the original vehicle driving data according to an embodiment of the present invention;

[0077] Figure 3 Shows a flowchart for mining redundant information of data according to an embodiment of the present invention;

[0078] Figure 4 Shows a flowchart for calibrating the clocks of each data source according to an embodiment of the present invention. Detailed Embodiments

[0079] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0080] The present invention provides a method for storing, compressing and dynamically synchronizing multi-source heterogeneous driving data as shown in Figures 1-4 and includes the following steps:

[0081] S1. Preprocess the original driving data to convert the original driving data into a unified format;

[0082] As shown in Figure 2 , the process of preprocessing the original driving data is as follows:

[0083] S11. Perform a moving average filtering process on the original driving data to obtain filtered data. The formula for the moving average filtering process is as follows:

[0084] ,

[0085] where and , is the number of time series data, is the number of data points participating in the moving average calculation, is the filtering result at position after moving average filtering, is the j-th data point in the original time series data, is the position index of the current data point in the entire data sequence;

[0086] S12. Sort the filtered data from smallest to largest, and then calculate the positions of the quartiles respectively. The calculation formula is as follows:

[0087] ,

[0088] , where is the position of the upper quartile, is the position of the lower quartile, is the number of filtered data;

[0089] S13. Calculate the lower quartile and the upper quartile respectively according to the positions of the quartiles. The calculation formula is as follows:

[0090] If and are integers, then , ;

[0091] If or is not an integer, then set or to have a decimal part of a and an integer part of b, and then use the quartile calculation formula to calculate and respectively. The calculation formula is:

[0092] , where is the lower quartile or the upper quartile , and are the data values at the -th and -th positions after sorting respectively;

[0093] S14. Subtract the lower quartile from the upper quartile to obtain the interquartile range IQ, and set the normal data range to , and delete the data that exceeds the normal data range;

[0094] S15. Use the FFmpeg library function to uniformly convert the video files into the mp4 format to improve the compatibility of video data on different devices and software platforms;

[0095] S16. Perform linear normalization processing on the numerical data to convert the data into a unified format. The normalization processing formula is as follows:

[0096] , where is the output value after normalization processing, is the numerical data, is the minimum value of the numerical data, is the maximum value of the numerical data;

[0097] By performing moving average filtering on the original driving data, the data can be made smoother and more accurate, reducing the impact of random interference on data analysis. Using the quartile method to find and delete outliers can reduce the misguidance of incorrect data on the overall analysis, ensuring that the data can represent normal driving conditions. Converting the video files into a format enables better sharing, transmission, and processing of the data on different devices, software, and platforms, improving the universality of the data. Normalization processing can make data of different magnitudes and distributions comparable and consistent, which is beneficial for subsequent data processing.

[0098] S2. Adaptively select a time-frequency transformation method according to the characteristics of the data to deeply mine the redundant information of the data, and quantify the mined redundant information;

[0099] The time-frequency transformation method is set to Fourier transform or wavelet transform. Among them, the Fourier transform is applicable to vehicle speed data with obvious periodicity. The vehicle speed data in the time domain is converted to the frequency domain through the Fourier transform to analyze the frequency components in the frequency domain and find the features corresponding to the periodic characteristics. The wavelet transform is applicable to brake data with mutation characteristics. By continuously performing the next-level discrete wavelet transform on the low-frequency approximation coefficients, the characteristics of the signal at different scales can be analyzed to capture the mutation information in the brake data;

[0100] As Figure 3 shown, the process of mining the redundant information of the data is as follows:

[0101] A1. Obtain frequency domain information data through Fourier transform or wavelet transform;

[0102] Among them, the Fourier transform formula is:

[0103] , where is the frequency domain data after Fourier transform, is the vehicle speed value collected at the discrete time point p, p = 0, 1,..., N - 1, that is, the vehicle speed data sequence from the 0th time point to the (N - 1)th time point, and N is the total number of points for collecting vehicle speed data, is the imaginary unit, satisfying , is the complex exponential term in the DFT transform, which is used to convert the time domain signal a to the frequency domain;

[0104] The formula for wavelet transform is as follows:

[0105] ,

[0106] , where = 0, 1,..., , N is the total number of points of the brake signal data, is the low-frequency component coefficient at the th position after DWT decomposition; is the high-frequency component coefficient at the th position after DWT decomposition, is the index variable used to identify the low-frequency approximation coefficient and the high-frequency detail coefficient , is the sample data of even numbers in the original signal, The sample data of odd numbers in the original discrete signal;

[0107] A2. Calculate the mean of the frequency-domain coefficients corresponding to different frames in the frequency domain. The mean calculation formula is:

[0108] , where M is the total number of data frames, is the frequency-domain coefficient of the th frame, is the frequency-domain coordinate, is the mean value at coordinate

[0109] A3. Calculate the variance at the coordinate according to the mean . The variance calculation formula is:

[0110] , where is the variance at the coordinate ;

[0111] A4. Set the quantization interval as [c, d] and the quantization level as L, and calculate the quantization step Log. The calculation formula is:

[0112] , where is the lower limit value of the quantization interval, is the upper limit value of the quantization interval;

[0113] A5. Calculate the quantized output value according to the quantization step Log. The calculation formula is:

[0114] , where is the frequency-domain coefficient, is the operation of mapping continuous frequency-domain coefficients to a finite number of discrete values, is the lower limit value of the quantization interval, is the quantized output value;

[0115] S3. Use the clock synchronization algorithm to accurately calibrate the clocks of each data source;

[0116] As Figure 3 shown, the process of calibrating the clocks of each data source is as follows:

[0117] S31. Initialize the clocks of each data source according to the function of the data source's clock as the master clock or the slave clock;

[0118] S32. The master clock periodically sends synchronization messages to the slave clocks in the network. The synchronization messages include the local timestamp when the master clock sends the synchronization message;

[0119] S33. When receiving a synchronization message from the slave clock, record the local time when the message is received ;

[0120] S34. After receiving the synchronization message, the slave clock sends a delay request message to the master clock and records the local time when the delay request message is sent ;

[0121] S35. After receiving the delay request message, the master clock sends a delay response message to the slave clock, and the delay response message contains the local time of the master clock when the delay request message is received ;

[0122] S36. The slave clock calculates the clock offset Offset and network transmission delay Delay between the slave clock and the master clock according to the received timestamp information, and the calculation formulas are as follows:

[0123] ,

[0124] ;

[0125] S37. The slave clock adjusts the local clock according to the calculated clock offset Offset to synchronize the slave clock with the master clock;

[0126] Through the above calibration method, the error of data synchronization can be controlled within a smaller range, the alignment accuracy of data in time can be improved, and the consistency of data from different data sources in time can be ensured to be higher.

[0127] S4. Extract representative feature data from the data of different data sources;

[0128] As Figure 4 shown, the process of feature data extraction is as follows:

[0129] S41. Determine the specific moment for which features are to be extracted according to the timestamp of the data. The determination of the specific moment satisfies , where is the time series of known data, is the data for which the specific moment is to be selected, is the index for traversing and summing the data points near the specific moment ;

[0130] S42. Arrange the data near the specific moment in ascending order, calculate the median of the arranged data, and take the median as the feature data;

[0131] S5. Establish the correspondence between data from different data sources by calculating the similarity between different features for feature matching;

[0132] The process of feature matching is as follows:

[0133] S51. Centered on a specific moment set the time window to be , where is the time window size;

[0134] S52. Calculate the similarity between the feature data of different data sources according to the feature data of different data sources within the time window. The similarity calculation formula is:

[0135] , where r is the dimension of the feature vector, is the i-th element of the feature vector A, is the i-th element of the feature vector B, is the similarity between the feature vector A and the feature vector B;

[0136] S53. According to the calculated similarity values, select the feature data pairs with high similarity to establish the correspondence between the feature data pairs.

[0137] S6. Fine-tune the data according to the feature matching results to ensure the accurate synchronization of the data in time;

[0138] The process of data fine-tuning is as follows:

[0139] S61. According to the results of feature matching, compare the timestamps of the data from different data sources and calculate the time deviation between different data sources;

[0140] S62. Make adjustments according to the time deviation to minimize the time deviation between different data sources and update the data records of the relevant data sources. When the time deviation is small, use the interpolation method for adjustment; when the time deviation is large, it is necessary to further check the data accuracy or re-perform feature matching;

[0141] Through the mechanism of feature matching, it can overcome the data synchronization problems caused by sensor acquisition differences and network transmission delays to a certain extent, enhance the stability and reliability of data synchronization, and ensure the accurate synchronization of data even in a complex driving environment.

[0142] S7. Establish a multi-source heterogeneous driving database, store the compressed data on different nodes, and improve the reliability and scalability of storage;

[0143] The process of establishing a multi-source heterogeneous driving database is as follows:

[0144] S71. Determine the overall architecture of the multi-source heterogeneous vehicle database. The overall architecture includes the storage method and the data organization form. Among them, the storage method includes file storage and relational database storage; the data organization form includes classification by data source and classification by time sequence;

[0145] S72. Determine the number of nodes and node configurations according to the characteristics and scale of the data;

[0146] S73. Store the compressed data on different nodes according to the planned architecture and storage method;

[0147] S8. Establish metadata information for the data of each data source to facilitate data query, retrieval and management. The metadata information includes the data source, collection time, sampling rate and sampling accuracy;

[0148] When establishing metadata information for the data of each data source, first define a detailed metadata structure for the data of each data source (the metadata structure includes fields such as data source, collection time, sampling rate and accuracy), then collect the metadata information of each data source during the data collection or storage process, and finally use a relational database to store the collected metadata information in the metadata management module of the database to ensure an effective association relationship is established between the metadata and the actual data, facilitating query and retrieval.

[0149] S9. Set a data update trigger mechanism and update the data in the database.

[0150] The process of setting the data update trigger mechanism is as follows:

[0151] S91. Determine the frequency and timeliness requirements for data update according to the application requirements for real-time vehicle status monitoring, driving behavior analysis or intelligent transportation management;

[0152] S92. Set the conditions for event triggering, and monitor the occurrence of relevant events in real time to determine whether the event triggering conditions are met. When the event triggering conditions are not met, continue to monitor the relevant events in real time. When the event triggering conditions are met, obtain the latest data from each data source;

[0153] S93. Clean and transform the obtained latest data, and then synchronize the updated data in chronological order according to the timestamp information of the data to ensure that the data is inserted into the database in the normal chronological order for subsequent analysis and query. If there are deviations in the timestamps of the data from different data sources, calibration and adjustment are required to minimize the deviation;

[0154] S94. Insert the synchronized updated data into the corresponding position in the database. If there are duplicate or conflicting data during data insertion, the duplicate or conflicting data will be overwritten.

[0155] The characteristics of data can be better captured through adaptive time-frequency transformation, redundant information in the frequency domain can be more accurately mined, and on the premise of ensuring data quality, the compression ratio can be further increased and the space required for storage can be reduced; by adopting a clock synchronization algorithm, the error of data synchronization can be controlled within a smaller range, the alignment accuracy of data in time can be improved, and the consistency of data from different data sources in time can be ensured to be higher.

[0156] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for storing, compressing, and dynamically synchronizing multi-source heterogeneous vehicle data, characterized in that: It includes the following steps: S1. Preprocess the original vehicle running data and convert it into a unified format; S2. According to the characteristics of the data, adaptively select a time-frequency transformation method to deeply mine the redundant information of the data, and perform quantization processing on the mined redundant information; S3. Use a clock synchronization algorithm to accurately calibrate the clocks of each data source; S4. Extract representative feature data from the data of different data sources; S5. Establish a correspondence between the data of different data sources by calculating the similarity between different features for feature matching; S6. Fine-tune the data according to the feature matching results to ensure accurate synchronization of the data in time; S7. Establish a multi-source heterogeneous vehicle running database and store the compressed data on different nodes; S8. Establish metadata information for the data of each data source to facilitate querying, retrieving, and managing the data; S9. Set a data update trigger mechanism and update the data in the database; Among them, in step S2, the process of mining the redundant information of the data is as follows: A1. Obtain frequency-domain information data through Fourier transform or wavelet transform; A2. Calculate the mean value of the frequency-domain coefficients corresponding to different frames in the frequency domain. The mean value calculation formula is: , where M is the total number of data frames, is the frequency domain coefficient of the frame, is the frequency domain coordinate, and A3. Calculate the variance at the coordinate according to the mean value. The variance calculation formula is: ​ , where is the variance at the coordinate . A4. Set the quantization interval as [c, d] and the quantization level as L, and calculate the quantization step Log. The calculation formula is: , where is the lower limit value of the quantization interval, is the upper limit value of the quantization interval; A5. Calculate the quantized output value according to the quantization step Log. The calculation formula is: , where is the frequency domain coefficient, is an operation that maps continuous frequency domain coefficients to a finite number of discrete values, is the lower limit value of the quantization interval, is the quantized output value.

2. The storage compression and dynamic synchronization method for multi-source heterogeneous vehicle data according to claim 1, wherein: In step S1, the process of preprocessing the original vehicle running data is as follows: S11. Perform moving average filtering on the original vehicle running data to obtain the filtered data. The moving average filtering formula is as follows: , Among them, and , is the number of time series data, is the number of data points participating in the moving average calculation, is the filtering result at position after moving average filtering, is the j-th data point in the original time series data, is the position index of the current data point in the entire data sequence; S12. Sort the filtered data from small to large, and then calculate the positions of the quartiles respectively. The calculation formula is as follows: , , where is the position of the upper quartile, is the position of the lower quartile, is the number of filtered data; S13. Calculate the lower quartile and the upper quartile respectively according to the positions of the quartiles. and the upper quartile , and the calculation formulas are as follows: If and are integers, then , ; If and are not integers, then set or to have a fractional part of a and an integer part of b, and then use the quartile calculation formula to calculate and respectively. The calculation formula is: , where is the lower quartile or the upper quartile , and are respectively the data values at the -th position after sorting; S14. Subtract the lower quartile from the upper quartile to obtain the interquartile range IQ, and set the normal data range as , and delete the data that exceeds the normal data range; S15. Use the FFmpeg library function to uniformly convert the video file into the mp4 format; S16. Perform linear normalization on the numerical data to convert the data into a unified format. The normalization formula is as follows: , where is the output value after normalization processing, is numerical data, is the minimum value of the numerical data, is the maximum value of the numerical data.

3. The method for storing, compressing and dynamically synchronizing multi-source heterogeneous vehicle running data according to claim 1, characterized in that: In step S3, the process of calibrating the clocks of each data source is as follows: S31. Initialize the clocks of each data source according to the function of the data source clock as the master clock or the slave clock; S32. The master clock periodically sends a synchronization message to the slave clocks in the network. The synchronization message includes the local timestamp of the master clock when sending the synchronization message. ; S33. When receiving a synchronization message from the clock, record the local time when the message is received ; S34. After receiving the synchronization message, the slave clock sends a delay request message to the master clock and records the local time when the delay request message is sent ; S35. After receiving the delay request message, the master clock sends a delay response message to the slave clock, and the local time of the master clock when the delay request message is received is included in the delay response message ; S36. Calculate the clock offset Offset and network transmission delay Delay between the slave clock and the master clock according to the received timestamp information of each time. The calculation formula is as follows: , ; S37. The slave clock adjusts its local clock according to the calculated clock offset Offset to synchronize the slave clock with the master clock.

4. The storage compression and dynamic synchronization method for multi-source heterogeneous vehicle running data according to claim 1, characterized in that: In step S4, the process of extracting feature data is as follows: S41. Determine a specific moment for feature extraction according to the timestamp of the data, and the determination of the specific moment satisfies , where is the time series of known data, is the data for which the specific moment is to be selected, is the index for traversing and summing the data points near the specific moment ; S42. Arrange the data near a specific moment in ascending order, calculate the median of the arranged data, and take this median as the characteristic data. ​ 5. The storage compression and dynamic synchronization method for multi-source heterogeneous vehicle running data according to claim 1, characterized in that: In step S5, the process of performing feature matching is as follows: S51. Centered around a specific moment , set the time window to be , where is the size of the time window; S52. Calculate the similarity between the feature data of different data sources within the time window. The similarity calculation formula is: , where is the dimension of the feature vector, is the i-th element of the feature vector A, is the i-th element of the feature vector B, is the similarity between the feature vector A and the feature vector B; S53. According to the calculated similarity value, select the feature data pairs with high similarity to establish the correspondence between the feature data pairs.

6. The storage compression and dynamic synchronization method for multi-source heterogeneous vehicle data according to claim 1, characterized in that: In step S6, the process of fine-tuning the data is as follows: S61. According to the feature matching results, compare the timestamps of the data of different data sources and calculate the time deviation between different data sources; Adjust according to the time deviation to minimize the time deviation between different data sources, and update the data records of relevant data sources. When the time deviation is small, use the interpolation method for adjustment; when the time deviation is large, further check the data accuracy or re-perform feature matching.

7. The method for storing, compressing and dynamically synchronizing multi-source heterogeneous vehicle running data according to claim 1, wherein: In step S7, the establishment process of the multi-source heterogeneous vehicle operation database is as follows: S71. Determine the overall architecture of the multi-source heterogeneous vehicle operation database, where the overall architecture includes the storage method and the data organization form; S72. Determine the number of nodes and node configurations according to the characteristics and scale of the data; S73. Store the compressed data on different nodes according to the planned architecture and storage method.

8. The storage compression and dynamic synchronization method for multi-source heterogeneous vehicle data according to claim 1, characterized in that: In step S8, when establishing metadata information for the data of each data source, first define a detailed metadata structure for the data of each data source, then collect the metadata information of each data source during the data acquisition or storage process, and finally use a relational database to store the collected metadata information in the metadata management module of the database to ensure an effective association relationship is established between the metadata and the actual data.

9. The method for storing, compressing and dynamically synchronizing multi-source heterogeneous vehicle running data according to claim 1, wherein: In step S9, the process of setting the data update trigger mechanism is as follows: S91. Determine the frequency and timeliness requirements for data update according to the application requirements for real-time vehicle status monitoring, driving behavior analysis, or intelligent transportation management; S92. Set the event trigger conditions and monitor the occurrence of relevant events in real time to determine whether the event trigger conditions are met. When the event trigger conditions are not met, continue to monitor the relevant events in real time. When the event trigger conditions are met, obtain the latest data from each data source; S93. Clean and transform the obtained latest data, and then synchronize the updated data in chronological order according to the timestamp information of the data to ensure that the data is inserted into the database in the normal chronological order. If there are deviations in the timestamps of the data from different data sources, calibration and adjustment are required to minimize the deviation; S94. Insert the synchronized updated data into the corresponding position in the database. If there are duplicate or conflicting data during the data insertion, overwrite the duplicate or conflicting data.

Citation Information

Patent Citations

  • Intelligent traffic data processing method and system based on edge calculation

    CN119479316A