A highway multi-source heterogeneous charging data real-time collection and fusion method
By predicting traffic flow through machine learning, and combining chaotic dynamics and game theory to optimize computing power allocation and dynamically schedule the computing power of edge nodes, the problem of insufficient computing power in highway toll collection systems under high traffic scenarios has been solved, and the data consistency and transmission reliability have been improved.
Patent Information
- Application Number
- CN202510523839.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In existing highway toll collection systems, the lack of a dynamic response mechanism for edge node computing power allocation during high-volume or abnormal traffic conditions leads to insufficient computing power or wasted resources, affecting the real-time performance and consistency of data processing.
By predicting traffic flow changes through machine learning, analyzing computing power requirements using chaotic dynamics, dynamically scheduling computing power at edge nodes, optimizing fusion strategies based on trust and game theory, and utilizing holographic data compression and error correction mechanisms for data transmission, the system optimizes bandwidth allocation and transmission paths.
It improves the efficiency and robustness of highway toll data processing, solves the problem of insufficient computing power in high-traffic scenarios, and enhances data consistency and transmission reliability.
Smart Images

Figure CN120407654B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of highway toll data management, and more specifically, to a method for real-time collection and fusion of multi-source heterogeneous toll data on highways. Background Art
[0002] Highway toll collection systems are a vital component of modern transportation infrastructure. They collect real-time traffic records and toll collection data through multiple sources, including gantries, lanes, and toll booths, to support functions such as traffic flow monitoring, toll management, and data analysis. With the expansion of highway networks and the increase in vehicle traffic, the scale and complexity of data collection have significantly increased. Edge-deployed computing centers are responsible for the initial fusion of heterogeneous multi-source data before transmitting it to the cloud for in-depth processing. However, due to device heterogeneity, differences in collection frequency, and the dynamic traffic environment, edge nodes must achieve efficient data cleaning, feature extraction, and initial fusion within limited computing power, while ensuring that the fusion strategy can adapt to real-time changes in data source quality. This scenario places high demands on edge computing capabilities and the dynamic nature of the fusion strategy.
[0003] In existing technologies, data collection and preliminary fusion usually adopt static rules or simple load balancing methods. For example, data cleaning based on preset thresholds and fixed fusion strategies are more common when processing gantry, lane and toll station data. However, these methods have significant defects: the edge node computing power allocation lacks a dynamic response mechanism to traffic fluctuations and nonlinear demands, resulting in insufficient computing power or waste of resources in high-load scenarios, affecting the real-time and consistency of preliminary fusion; these problems are particularly prominent under high traffic or abnormal traffic patterns (such as holiday peaks), which directly restricts the efficiency and quality of highway toll data processing. There is an urgent need for a technical solution that can dynamically optimize computing power allocation and adaptively adjust fusion strategies to improve edge processing capabilities. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways, comprising the following steps:
[0006] Real-time data collection of highway toll station gantries, lane traffic record data, and toll collection data is collected. After cleaning, the data is obtained from the collection end and transmitted to the edge computing center through data links and network equipment for preliminary fusion to obtain a preliminary fused data set.
[0007] The computing power center deployed at the edge of each toll station is recorded as an edge node. Future traffic flow trends are predicted through machine learning, and a traffic forecast distribution is generated. Based on this traffic forecast distribution, the computing power demand profile of each edge node is extracted. The machine learning includes: collecting historical and real-time traffic data, normalizing it, performing time series analysis using a time series prediction model, mapping the traffic data to a high-dimensional feature representation, adjusting parameters through an optimizer, and outputting future traffic forecast values; and calculating the computing power requirements of each edge node based on the traffic forecast values to form a computing power demand profile.
[0008] A chaotic dynamics model is used to analyze the nonlinear fluctuation characteristics of the computing power demand profile, identify the chaotic attractors of high-load nodes, generate the overall computing power boundary of the cluster, and dynamically schedule computing power based on a chaos control strategy. If the computing power demand exceeds the preset proportion of the node's available computing power, the prediction deviation is corrected through cross-node computing power borrowing and online learning.
[0009] When conflicts in key fields are detected in multi-source data, the fusion strategy is adjusted based on dynamic trust and game theory optimization.
[0010] Preferably, the chaotic attractor of a high-load node refers to a set of trajectories that evolve over time and remain stable when computing power demand fluctuates, reflecting the long-term behavior pattern of computing power; the overall computing power boundary of the cluster refers to the maximum computing power limit that the entire edge node cluster can provide within a given time period, taking into account the computing power resources and network bandwidth constraints of each node.
[0011] Preferably, a preliminary fusion data set is obtained through feature extraction and strategy selection, and the preliminary fusion process includes the following steps:
[0012] Step S11, data source preparation and feature extraction: using metadata analysis tools, extract the data source features of each data source in the highway multi-source heterogeneous toll collection data, including data structure, field type, data distribution, and collection frequency;
[0013] Step S12, real-time strategy selection and execution: Build a strategy library containing multi-source heterogeneous data fusion methods, match the extracted data source features with the methods in the strategy library, and generate candidate fusion strategies; select the optimal strategy and execute it based on historical fusion results and data quality indicators;
[0014] Step S13, data fusion: execute the selected fusion strategy to integrate multi-source heterogeneous data into a data set in a unified format.
[0015] Preferably, the preliminary fused data set is transmitted to the highway toll network data center deployed at the cloud edge through data links and network equipment, and deep fusion is performed to output the fused highway toll data; the deep fusion includes position alignment, time alignment, and outlier correction.
[0016] Preferably, when information conflicts occur in multi-source heterogeneous data, a fusion strategy adjustment step based on dynamic trust and game theory optimization is triggered, specifically including:
[0017] In real-time strategy selection and execution, the trust score of each data source is calculated in real time based on the integrity, consistency, and timeliness of the data collected. The Nash equilibrium model in game theory is used to dynamically adjust the fusion weight of each data source.
[0018] If the trust level of a data source is lower than the preset threshold, the corresponding weight is reduced through a zero-sum game strategy, and an adversarial verification mechanism is introduced to verify the impact of the weight adjustment on the fusion consistency;
[0019] Combined with the computing power status fed back by edge nodes in real time, the payment matrix of the game model is dynamically updated to ensure the coordinated optimization of weight adjustment and computing power allocation;
[0020] The adjusted fusion results are verified for consistency with the trust change log, and an analysis report including the game optimization effect is generated.
[0021] Preferably, the trust score is obtained in the following manner:
[0022] Dynamic integrity quality index DCIQ is quantitatively expressed based on field completeness rate FCR, missing record percentage MRP and missing field distribution standard deviation;
[0023] The multi-source consistency quality index (MCIQ) is quantitatively expressed based on the matching rate of non-missing record fields, the conflict ratio of non-missing records, and the standard deviation of the conflict field distribution.
[0024] The real-time quality index RTI is quantitatively expressed based on the update frequency UF, expected update frequency, average delay time DL and delay distribution standard deviation;
[0025] The dynamic integrity quality index, multi-source consistency quality index and real-time timeliness quality index are jointly analyzed to obtain the trust score CQI.
[0026] Preferably, the method further includes a transmission network optimization step, specifically including:
[0027] Step S31: Optimizing bandwidth allocation and traffic scheduling through a heuristic optimization algorithm;
[0028] Step S32, adaptive neural network topology optimization: deploying a network topology simulated by a neural network in the network between the edge nodes and the cloud, mimicking the adaptive characteristics of neuronal connections in the human brain; using a generative adversarial network to dynamically generate the optimal network path and respond to topology changes in real time;
[0029] Step S33, holographic data compression and transmission: Holographically encode the data collected at the edge node, map the multi-source heterogeneous data to a high-dimensional holographic space, and generate a compressed data packet;
[0030] Step S34: Combine redundant coding with a multi-path transmission protocol to fragment the data packet and transmit it in parallel through multiple paths, and then reassemble it in the cloud.
[0031] Preferably, when the edge node generates a holographically encoded data packet, the redundant coding and error correction mechanism is used to bind the data packet and the error correction bit into an entangled pair; during the multi-path transmission process, the packet loss rate of each path is monitored in real time. If the packet loss rate of a certain path exceeds the preset value, the distribution ratio of the entangled pairs is dynamically adjusted, and the error correction bit is preferentially transmitted to the low packet loss path.
[0032] Preferably, in the holographic data compression and transmission step:
[0033] Different holographic encoding strategies are used for the image data and numerical data of the acquisition end. After the image data is mapped to the high-dimensional holographic space, sparse representation technology is used to reduce redundancy and generate a compressed image feature matrix. The numerical data is vector quantized to generate a compressed feature vector, retaining key charging information.
[0034] After the edge node completes the holographic encoding, it integrates the compressed image feature matrix and numerical feature vector into a unified data packet, and records the compression rate and encoding time of each data type;
[0035] During reassembly in the cloud, a deep learning model is used to decode the compressed data packets and restore the original data structure. At the same time, the data packet loss rate is calculated through integrity verification. If the loss rate exceeds a preset threshold, a retransmission mechanism is triggered to retrieve the lost data packets from the edge node. The deep learning model is trained based on historical transmission data and adaptively adjusts decoding parameters to cope with changes in data distribution in different sections and time periods.
[0036] Preferably, in the deep fusion process, an implicit time drift correction step based on path topology and multi-scale time drift compensation is included, including:
[0037] Construct a highway path topology map, use the longitude and latitude coordinates of toll booths and gantries to generate sequential constraints for vehicle travel paths, and calculate theoretical travel time ranges based on historical speed distribution and road speed limits;
[0038] The distribution of timestamps recorded by each device is analyzed using short-term sliding windows and long-term sliding windows. The short-term window length ranges from 1 to 5 minutes, and the long-term window length ranges from 30 to 60 minutes. The mean shift and variance fluctuation characteristics of the time drift are extracted.
[0039] Based on the path topology constraints and the extracted time drift features, a reinforcement learning algorithm based on Q-learning is used to dynamically correct the timestamp deviation of each device, and the rationality of vehicle speed and data alignment consistency are used as reward functions.
[0040] The technical effects and advantages of the present invention are as follows:
[0041] (1) The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways provided by the present invention collects data from gantries, lanes and toll stations in real time, performs preliminary fusion after cleaning, uses machine learning to predict traffic flow, combines chaotic dynamics to analyze fluctuations in computing power demand, and dynamically schedules edge node computing power to solve the problem of insufficient computing power in high-traffic scenarios. At the same time, the trust score of the data source is calculated based on the integrity, consistency and timeliness indicators, and game theory (Nash equilibrium and zero-sum game) is used to dynamically adjust the fusion weight to improve data consistency, thereby significantly improving fusion efficiency and robustness.
[0042] (2) The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways provided by the present invention effectively solves the delay, packet loss and security problems in the transmission of highway toll data through the dynamic allocation of bandwidth and traffic prediction by heuristic optimization algorithm, dynamic adjustment of transmission path by adaptive neural network, holographic data compression combined with redundant coding and error correction mechanism, and zero-trust architecture to enhance security verification; holographic compression technology maps multi-source heterogeneous data to high-dimensional space, solving the data transmission delay problem existing in the process of highway toll data transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flow chart of the method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways of the present invention.
[0044] Figure 2 This is a structural block diagram of the preliminary fusion process of the present invention.
[0045] Figure 3 This is a flow chart of the transmission network optimization of the present invention. DETAILED DESCRIPTION
[0046] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0047] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0048] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0049] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0050] Example 1, see Figure 1 Flowchart of the real-time collection and fusion method of highway multi-source heterogeneous toll collection data and Figure 2 The present invention provides a method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways, including the following steps:
[0051] Step 1: Collect highway toll station gantry, lane traffic record data, and toll collection data in real time. After cleaning, obtain highway collection end data, and transmit it to the edge computing center through data links and network equipment for preliminary fusion to obtain a preliminary fused data set. Explain that highway toll station gantry, lane traffic record data, and toll collection data constitute highway multi-source heterogeneous toll collection data.
[0052] Step 2: Transmit the preliminary fused data set to the highway toll network data center deployed at the cloud edge via data links and network equipment, perform deep fusion, and output the fused highway toll data. The deep fusion includes obtaining high-quality, unified-format highway toll data based on position alignment, time alignment, and outlier correction, thus compensating for the lack of computing power at the edge nodes and improving data consistency and availability.
[0053] Step 3: Store the fused highway toll data in a centralized database or data warehouse to support data query and statistical analysis. For example, the fused highway toll data can be indexed by time, toll station, or vehicle type to facilitate the generation of traffic flow reports or toll statistics.
[0054] Step 4: Monitor the operating status of the data link and edge computing center in real time, and trigger an alarm if an anomaly is detected.
[0055] In this implementation, it is necessary to further explain that a preliminary fusion data set is obtained through feature extraction and strategy selection. The preliminary fusion process includes the following steps:
[0056] Step S11, data source preparation and feature extraction: using metadata analysis tools, extract the data source features of each data source in the highway multi-source heterogeneous toll collection data, including data structure, field type, data distribution, and collection frequency;
[0057] Step S12, real-time strategy selection and execution: Build a strategy library containing multi-source heterogeneous data fusion methods, match the extracted data source features with the methods in the strategy library, and generate candidate fusion strategies; select the optimal strategy and execute it based on historical fusion results and data quality indicators;
[0058] Step S13, data fusion: execute the selected fusion strategy to integrate multi-source heterogeneous data into a data set in a unified format.
[0059] In this embodiment, it is necessary to further explain that the computing power center deployed at the edge of each toll station is recorded as an edge node; the method includes the step of allocating computing power to the edge nodes, including the following steps:
[0060] Step S21: Predict the traffic flow change trend for the next hour through machine learning, generate a traffic flow prediction distribution, and extract the computing power requirement profile of each edge node based on the traffic flow prediction distribution, wherein the profile is based on the data processing load per unit time;
[0061] The detailed implementation process includes: collecting real-time traffic data from highway toll station gantries, lanes, and toll station terminals, integrating historical data (recommended to be at least six months) including vehicle pass counts, average speeds, and timestamps as a training set, and incorporating contextual variables such as holidays and weather conditions; normalizing the data to map traffic values to the [0, 1] interval to prevent dimensional differences from affecting model training;
[0062] Perform statistical processing on the predicted values to calculate the mean and standard deviation of traffic every 5 minutes. Use kernel density estimation (KDE) to smooth the predicted curve and generate a continuous traffic distribution function. Store the results as a time series for subsequent computing power demand calculations.
[0063] Define the data processing load per unit time. Based on system benchmarks, assume that each traffic record requires 0.1 unit of computing power (in CPU cycles or FLOPS). Calculate the computing power required for each edge node. For each 5-minute period, the demand is the product of the traffic forecast mean and the computing power per record. Organize the results by node and time to form a computing power demand profile.
[0064] Step S22: Analyze the nonlinear fluctuation characteristics of the computing power demand profile using a chaotic dynamics model, identify chaotic attractors of high-load nodes, and generate the overall computing power boundary of the cluster. Explanation: The chaotic attractor of a high-load node refers to a set of stable trajectories exhibited by high-load nodes (such as toll stations that process a large number of vehicle traffic records) that evolve over time during fluctuations in computing power demand, reflecting the long-term behavior pattern of computing power. The overall computing power boundary of the cluster refers to the maximum computing power that the entire edge node cluster (i.e., the computing power center of all toll stations) can provide within a given time period, taking into account constraints such as the computing power resources and network bandwidth of each node.
[0065] The detailed implementation process includes: taking the generated computing power demand profile as input, organizing it by edge nodes and time series; using Python's chaospy library or MATLAB's chaos analysis module, building a chaotic model based on the logistic map, normalizing the computing power demand to the interval [0, 1], setting control parameters to generate a chaotic sequence, and running 1000 iterations to ensure that the sequence enters a stable chaotic state; calculating the Lyapunov exponent; if it is a positive value (for example, 0.5), confirming that the demand fluctuation has chaotic characteristics; reconstructing the phase space using delayed embedding technology; counting the peak computing power demand of all nodes, calculating the total available computing power of the cluster, and integrating it by time period to form a dynamic boundary curve;
[0066] Step S23: If the computing power demand exceeds the preset value, for example, exceeds the node's available computing power by 20%, it is determined to be a high load type. The computing power is dynamically scheduled based on the chaos control strategy. The prediction deviation is corrected through cross-node computing power borrowing and online learning. The borrowing ratio is adaptively adjusted according to the stability of the chaotic attractor.
[0067] Step S24: Monitor the scheduling effect in real time. If the consistency error is less than 5%, feed the scheduling strategy back to the machine learning model to optimize the prediction for the next cycle.
[0068] In this implementation, it is necessary to further explain that when information conflicts occur in multi-source heterogeneous data, or when conflicts are detected in key fields of multi-source data (e.g., timestamp differences exceed a reasonable threshold), the fusion strategy adjustment steps based on dynamic trust and game theory optimization are triggered, specifically including:
[0069] In real-time strategy selection and execution, the trust score of each data source is calculated in real time based on the integrity, consistency, and timeliness of the data collected. The Nash equilibrium model in game theory is used to dynamically adjust the fusion weight of each data source.
[0070] If the trust level of a data source is lower than the preset threshold, the corresponding weight is reduced through a zero-sum game strategy, and an adversarial verification mechanism is introduced to verify the impact of the weight adjustment on the fusion consistency;
[0071] Combined with the computing power status fed back by edge nodes in real time, the payment matrix of the game model is dynamically updated to ensure the coordinated optimization of weight adjustment and computing power allocation;
[0072] The adjusted fusion results are verified for consistency with the trust change log, and an analysis report including the game optimization effect is generated.
[0073] Furthermore, the integrity, consistency and timeliness indicators of the data at the collection end are quantified into the dynamic integrity quality index, the multi-source consistency quality index and the real-time timeliness quality index respectively; the dynamic integrity quality index DCIQ is quantitatively expressed based on the field completeness rate FCR, the missing record percentage MRP and the standard deviation of the missing field distribution; the multi-source consistency quality index MCIQ is quantitatively expressed based on the non-missing record field matching rate, the non-missing record conflict ratio and the standard deviation of the conflict field distribution; the real-time timeliness quality index RTI is quantitatively expressed based on the update frequency UF, the expected update frequency, the average delay time DL and the standard deviation of the delay distribution; the dynamic integrity quality index, the multi-source consistency quality index and the real-time timeliness quality index are jointly analyzed to obtain the trust score CQI.
[0074] In one possible embodiment, the integrity index is used to measure the missing and empty values of key fields (such as timestamps, license plates, and traffic records) in the data collected by the acquisition end. The integrity index is quantified by calculating the completeness rate of required fields for each data packet, that is, the ratio of the actual filled-in value of the key fields to the required filled-in value; using the sliding window technique, the percentage of missing records within a certain time window is calculated;
[0075] In one possible embodiment, the consistency index is used to evaluate the degree of matching between multi-source data in the data collection end in terms of logical relationships, data formats, and semantics. The consistency index is quantified by: for the same vehicle and the same traffic event, field matching is performed in different data sources, and the matching rate of field values is calculated; and the proportion of records with conflicts or inconsistent data formats is counted and analyzed to evaluate data consistency.
[0076] In a possible embodiment, the timeliness index is used to measure the delay time of the data from the collection end to the computing power center, as well as the real-time update frequency of the data. The real-time update frequency of the data refers to the number of times the data is updated within a specific time, reflecting the freshness and timeliness of the data; the delay time is quantified by recording the timestamp at each link of data collection, transmission, and processing, and determining the delay time by calculating the time difference between each link; the real-time update frequency of the data is quantified by setting a time window, recording the number of times the data is updated within the window, and calculating the ratio of the number of updates to the window length to obtain the update frequency of the data at the collection end.
[0077] For ease of understanding, the embodiment of the present invention provides a specific method for quantifying the trust score as follows:
[0078] By formula The dynamic integrity quality index DCIQ is obtained; the field completeness rate FCR is calculated as the ratio of the number of fields actually filled in the time window to the number of fields that should be filled in; the missing record percentage MRP is calculated as the ratio of the number of missing records in the time window to the total number of records; σ m issing indicates the standard deviation of the missing field distribution;
[0079] By formula Get the multi-source consistency quality index MCIQ and the non-missing record field matching rate FMR non-missing The calculation method is the ratio of the number of matching fields in non-missing records to the total number of compared fields; the non-missing record conflict ratio CRP non-missing The calculation method is the ratio of the number of conflicting records in non-missing records to the total number of non-missing records; the standard deviation of the conflict field distribution σ c onflict is calculated as the standard deviation of the number of conflicting fields in nonmissing records;
[0080] Get the expected update frequency UF e xp, average delay time DL, delay distribution standard deviation σ delay , through the formula The real-time quality index RTI is obtained, and the update frequency UF is calculated as the ratio of the number of updates in the time window to the window length;
[0081] The Dynamic Integrity Quality Index (DCIQ), Multi-Source Consistency Quality Index (MCIQ) and Real-Time Quality Index (RTIQ) are jointly analyzed to obtain the trust score CQI;
[0082]
[0083] Explanation: The numerator is used to integrate the contributions of the three and emphasize the weak link effect; if any of the indicators (DCIQ, MCIQ, RTIQ) is too low, the CQI will drop significantly; in the denominator, the constant 3 (the maximum sum of the three) ensures that the denominator range is reasonable, and the square root term is used to nonlinearly penalize low-value indicators.
[0084] Furthermore, the steps of optimizing time window selection based on traffic patterns include: using historical traffic data, speed data and real-time gantry traffic counts, and adopting a random forest algorithm to train a traffic pattern classifier to identify peak, trough and congestion patterns; dynamically adjusting the time window length according to the traffic pattern to match a short window length for peak periods, for example, the peak period window length is 1 to 5 minutes, the trough period window length is 30 to 60 minutes, and the congestion period window length is 10 to 20 minutes.
[0085] Summary: The embodiment of the present invention collects data from gantries, lanes, and toll booths in real time, performs preliminary fusion after cleaning, uses machine learning to predict traffic flow, combines chaotic dynamics to analyze fluctuations in computing power demand, and dynamically schedules edge node computing power to solve the problem of insufficient computing power in high-traffic scenarios. At the same time, the data source trust score is calculated based on integrity, consistency, and timeliness indicators, and game theory (Nash equilibrium and zero-sum game) is used to dynamically adjust the fusion weight to improve data consistency. After the fusion result passes the consistency verification, a preliminary data set is output and feedback is provided to optimize the prediction model. Traffic mode classification is also introduced to dynamically adjust the time window to adapt to different traffic scenarios.
[0086] Example 2, in order to solve the existing data transmission delay problem, is applied to: transmitting the data to the edge-deployed computing center through data links and network equipment, and transmitting the initially integrated data from the collection end to the cloud-edge-deployed highway toll network data center, refer to Figure 3 The transmission network optimization flow chart of the method further includes a transmission network optimization step, specifically including the following:
[0087] Step S31: Optimize bandwidth allocation and traffic scheduling using a heuristic optimization algorithm. The specific implementation process is as follows: Deploy a heuristic optimization algorithm to optimize bandwidth allocation in the data link between the edge node and the cloud; Based on historical traffic flow data and real-time road conditions (such as peak holiday periods and accident-prone periods), use machine learning to predict traffic flow trends for the next 5-30 minutes, analyze multiple variables (such as vehicle speed, toll station distribution, and network load) in parallel, and generate a dynamic bandwidth scheduling strategy to address the computational bottleneck of traditional deep learning in ultra-large-scale variable optimization;
[0088] Step S32, adaptive neural network topology optimization: Deploy a network topology simulated by a neural network in the network between the edge node and the cloud (the highway toll network data center deployed at the cloud edge), imitating the adaptive characteristics of human brain neuron connections; use a generative adversarial network to dynamically generate the optimal network path and respond to topology changes in real time;
[0089] Step S33, holographic data compression and transmission: Holographic encoding is performed on the acquisition end data at the edge node, and multi-source heterogeneous data (such as images recorded by the gantry and the numerical values of the charging data) is mapped to a high-dimensional holographic space to generate a compressed data packet;
[0090] Step S34: Combine redundant coding with a multi-path transmission protocol to fragment the data packet and transmit it in parallel through multiple paths, and then reassemble it in the cloud.
[0091] Furthermore, in holographic data compression and transmission, an adaptive error correction mechanism is used to improve the robustness of multi-path parallel transmission. The specific implementation process is as follows: when generating a holographic encoded data packet at the edge node, the redundant coding and error correction check mechanism is used to bind the data packet and the error correction bit into an entangled pair; during the multi-path transmission process, the packet loss rate of each path is monitored in real time. If the packet loss rate of a certain path exceeds a preset value, such as 3%, the distribution ratio of the entangled pairs is dynamically adjusted, and the error correction bit is preferentially transmitted to the low packet loss path; during cloud reassembly, the lost data is restored based on the entanglement state measurement technology, and the error correction success rate can reach more than 98%; the adaptive error correction mechanism combined with the entanglement heuristic protocol breaks through the limitations of traditional error correction methods (such as Reed-Solomon code) in dynamic network environments, and improves the transmission reliability of highway toll data in highly dynamic and multi-interference scenarios;
[0092] Furthermore, real-time security optimization under the zero-trust architecture: during data transmission, a zero-trust network architecture is implemented, and each data stream needs to pass dynamic authentication (such as a digital signature based on the vehicle ID) and integrity check (such as hash verification); combined with lightweight encryption, security is guaranteed without increasing latency; the real-time security optimization steps under the zero-trust architecture include: in dynamic authentication, a unique time-sensitive digital signature is generated based on the vehicle ID and toll station location, and the signature is encrypted using a lightweight encryption algorithm (where time sensitivity is achieved by appending the current timestamp to milliseconds) to ensure that the signature is valid within a specific time window such as 5 seconds); in the integrity check, the hash value of the data stream is calculated in real time and compared with the verification value pre-generated by the edge node. If an anomaly is found, the anomaly event is recorded and an alarm is triggered; for example, if the two are inconsistent, it is determined that the data may have been tampered with, the time of the anomaly event, the data stream identifier and difference details are recorded, and an alarm is triggered to notify the computing power center.
[0093] It should be further explained in the embodiment of the present invention that the holographic data compression and transmission step further includes:
[0094] Different holographic encoding strategies are used for the image data and numerical data of the acquisition end. After the image data is mapped to the high-dimensional holographic space, sparse representation technology is used to reduce redundancy and generate a compressed image feature matrix. The numerical data is vector quantized to generate a compressed feature vector, retaining key charging information.
[0095] After the edge node completes the holographic encoding, it integrates the compressed image feature matrix and numerical feature vector into a unified data packet, and records the compression rate and encoding time of each data type;
[0096] During reassembly in the cloud, a deep learning model is used to decode the compressed data packets and restore the original data structure. At the same time, the data packet loss rate is calculated through integrity verification. If the loss rate exceeds a preset threshold (such as 5%), a retransmission mechanism is triggered to retrieve the lost data packets from the edge node. The deep learning model is trained based on historical transmission data and adaptively adjusts decoding parameters to cope with changes in data distribution in different sections and time periods.
[0097] In summary, the embodiments of the present invention effectively solve the delay, packet loss and security problems in highway toll data transmission by optimizing bandwidth allocation and traffic prediction, dynamically adjusting the transmission path through an adaptive neural network, combining holographic data compression with an entanglement error correction mechanism, and strengthening security verification through a zero-trust architecture; holographic compression technology maps multi-source heterogeneous data to a high-dimensional space, solving the data transmission delay problem that exists in the process of highway toll data transmission.
[0098] Example 3: The difference between this embodiment of the present invention and Examples 1 and 2 is that the deep fusion includes the following steps:
[0099] Data alignment: Obtain the preliminarily fused acquisition data with timestamps transmitted by each edge node, and align them based on the location of each edge node (such as the latitude and longitude of the toll station) and timestamp;
[0100] Outlier correction: Fix data errors using statistical rules or machine learning methods, including correcting license plate recognition bias and completing missing fields;
[0101] Verify the logical relationship between multi-source data. For records that fail verification, take corrective measures (such as calibration, interpolation, and confidence level selection) first. If correction is not possible, remove the conflicting records and record the cause of the anomaly.
[0102] Furthermore, the process of verifying the logical relationship between the multi-source data at least includes verifying the rationality of the toll amount and the travel time, and eliminating conflicting records;
[0103] Verify that the vehicle identification (license plate number) matches the toll records (the license plate number of the same vehicle should be consistent in the gantry, lane and toll station records);
[0104] Verify the consistency of the passage time and geographic location (based on the timestamp and the latitude and longitude of the toll station, verify whether the time sequence of vehicle passage is consistent with the actual path);
[0105] Verify the rationality of the travel speed and time interval (the time interval between adjacent toll stations should be reasonably matched with the vehicle travel speed and distance).
[0106] It is necessary to further explain in the embodiment of the present invention that, in the deep fusion process, an implicit time drift correction step based on path topology and multi-scale time drift compensation is included, specifically including:
[0107] Construct a highway path topology map, use the longitude and latitude coordinates of toll booths and gantries to generate sequential constraints for vehicle travel paths, and calculate theoretical travel time ranges based on historical speed distribution and road speed limits;
[0108] The distribution of timestamps recorded by each device is analyzed using short-term sliding windows and long-term sliding windows. The short-term window length ranges from 1 to 5 minutes, and the long-term window length ranges from 30 to 60 minutes. The mean shift and variance fluctuation characteristics of the time drift are extracted.
[0109] Based on the path topology constraints and the extracted time drift features, a Q-learning-based reinforcement learning algorithm is used to dynamically correct the timestamp deviation of each device, with the rationality of vehicle speed and data alignment consistency as the reward function;
[0110] Furthermore, each edge node is collaboratively optimized by sharing time drift characteristics and correction strategies to build a global time drift correction network. The correction step size is dynamically adjusted between edge nodes through a joint reward function based on reinforcement learning (combining local consistency and global consistency), prioritizing the correction of drift deviations in high-traffic nodes, thereby improving the time synchronization accuracy of the entire network and the consistency of fused data.
[0111] Explanation: The joint reward function is a function used in reinforcement learning to evaluate the correction effect. It combines local consistency (data alignment accuracy of a single node, such as timestamp deviation < 1 second) and global consistency (overall accuracy of data alignment of all nodes, such as license plate matching rate > 95%) to comprehensively optimize time correction. High-traffic nodes refer to edge nodes that process a large number of vehicle passage records within a specific time period. For example, during peak holiday periods, a toll station processes 200 records per minute, far higher than the average of 50.
[0112] The corrected timestamp is applied to the data alignment step, and the detected time drift anomaly is fed back to the equipment maintenance module through the network alarm system.
[0113] Summary: High-quality fusion is achieved through data alignment (based on location and timestamp), outlier correction (statistical rules or machine learning to repair license plate deviations and fill missing fields), and multi-source data logical relationship verification (including the rationality of toll amounts, license plate matching, passage order, and speed consistency). For records that fail verification, calibration or interpolation is prioritized. If correction is not possible, they are removed and the anomaly is recorded. In addition, the method introduces implicit time drift correction based on path topology and multi-scale time drift compensation, constructs a path topology graph, uses short-term (1-5 minutes) and long-term (30-60 minutes) sliding windows to extract drift features, and uses Q-learning reinforcement learning to dynamically correct timestamp deviations. Further, through collaborative optimization of edge nodes, drift features and correction strategies are shared to construct a global time drift correction network. Combined with a joint reward function (integrating local and global consistency), the correction step size is dynamically adjusted, high-traffic nodes are corrected first, and time synchronization accuracy is improved.
[0114] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A real-time collection and fusion method for multi-source heterogeneous toll collection data on highways, characterized by: include: Real-time data collection of highway toll station gantries, lane traffic record data, and toll collection data is collected. After cleaning, the data is obtained from the collection end and transmitted to the edge computing center through data links and network equipment for preliminary fusion to obtain a preliminary fused data set. The computing center deployed at the edge of each toll station is recorded as an edge node. Future traffic flow trends are predicted through machine learning to generate a traffic forecast distribution. Based on this traffic forecast distribution, the computing power demand profile of each edge node is extracted. The machine learning includes: collecting historical and real-time traffic data, normalizing it, performing time series analysis using a time series prediction model, mapping the traffic data to a high-dimensional feature representation, adjusting parameters through an optimizer, and outputting future traffic forecast values; and calculating the computing power requirements of each edge node based on the traffic forecast values to form a computing power demand profile. A chaotic dynamics model is used to analyze the nonlinear fluctuation characteristics of the computing power demand profile, identify the chaotic attractors of high-load nodes, generate the overall computing power boundary of the cluster, and dynamically schedule computing power based on a chaos control strategy. If the computing power demand exceeds the preset proportion of the node's available computing power, the prediction deviation is corrected through cross-node computing power borrowing and online learning. When conflicts in key fields are detected in multi-source data, the fusion strategy is adjusted based on dynamic trust and game theory optimization.
2. The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways according to claim 1 is characterized in that: The chaotic attractor of a high-load node refers to a set of stable trajectories exhibited by a high-load node that evolves over time amid fluctuations in computing power demand, reflecting the long-term behavior pattern of computing power. The overall computing power boundary of the cluster refers to the maximum computing power that the entire edge node cluster can provide within a given time period, taking into account the computing power resources and network bandwidth constraints of each node.
3. The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways according to claim 1 is characterized in that: A preliminary fusion data set is obtained through feature extraction and strategy selection. The preliminary fusion process includes the following steps: Step S11, data source preparation and feature extraction: using metadata analysis tools, extract the data source features of each data source in the highway multi-source heterogeneous toll collection data, including data structure, field type, data distribution, and collection frequency; Step S12, real-time strategy selection and execution: Build a strategy library containing multi-source heterogeneous data fusion methods, match the extracted data source features with the methods in the strategy library, and generate candidate fusion strategies; select the optimal strategy and execute it based on historical fusion results and data quality indicators; Step S13, data fusion: execute the selected fusion strategy to integrate multi-source heterogeneous data into a data set in a unified format.
4. The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways according to claim 1 is characterized in that: include: The preliminary fused data set is transmitted to the highway toll network data center deployed at the cloud edge through data links and network equipment, and deep fusion is performed to output the fused highway toll data; the deep fusion includes position alignment, time alignment, and outlier correction.
5. The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways according to claim 3 is characterized in that: When information conflicts occur among heterogeneous data from multiple sources, the fusion strategy adjustment steps based on dynamic trust and game theory optimization are triggered, including: In real-time strategy selection and execution, the trust score of each data source is calculated in real time based on the integrity, consistency, and timeliness of the data collected. The Nash equilibrium model in game theory is used to dynamically adjust the fusion weight of each data source. If the trust level of a data source is lower than the preset threshold, the corresponding weight is reduced through a zero-sum game strategy, and an adversarial verification mechanism is introduced to verify the impact of the weight adjustment on the fusion consistency; Combined with the computing power status fed back by edge nodes in real time, the payment matrix of the game model is dynamically updated to ensure the coordinated optimization of weight adjustment and computing power allocation; The adjusted fusion results are verified for consistency with the trust change log, and an analysis report including the game optimization effect is generated.
6. The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways according to claim 5 is characterized in that: The trust score is obtained as follows: Dynamic integrity quality index DCIQ is quantitatively expressed based on field completeness rate FCR, missing record percentage MRP and missing field distribution standard deviation; The multi-source consistency quality index (MCIQ) is quantitatively expressed based on the matching rate of non-missing record fields, the conflict ratio of non-missing records, and the standard deviation of the conflict field distribution. The real-time quality index RTI is quantitatively expressed based on the update frequency UF, expected update frequency, average delay time DL and delay distribution standard deviation; The dynamic integrity quality index, multi-source consistency quality index and real-time timeliness quality index are jointly analyzed to obtain the trust score CQI.
7. A method for real-time collection and fusion of multi-source heterogeneous toll collection data for highways according to any one of claims 1 to 4, characterized in that: The method further includes a transmission network optimization step, specifically comprising: Step S31: Optimizing bandwidth allocation and traffic scheduling through a heuristic optimization algorithm; Step S32, adaptive neural network topology optimization: deploying a network topology simulated by a neural network in the network between the edge nodes and the cloud, mimicking the adaptive characteristics of neuronal connections in the human brain; using a generative adversarial network to dynamically generate the optimal network path and respond to topology changes in real time; Step S33, holographic data compression and transmission: Holographically encode the data collected at the edge node, map the multi-source heterogeneous data to a high-dimensional holographic space, and generate a compressed data packet; Step S34: Combine redundant coding with a multi-path transmission protocol to fragment the data packet and transmit it in parallel through multiple paths, and then reassemble it in the cloud.
8. The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways according to claim 7 is characterized in that: When a holographically encoded data packet is generated at the edge node, the redundant coding and error correction mechanism is used to bind the data packet and the error correction bit into an entangled pair. During the multi-path transmission process, the packet loss rate of each path is monitored in real time. If the packet loss rate of a certain path exceeds the preset value, the distribution ratio of the entangled pairs is dynamically adjusted, and the error correction bit is preferentially transmitted to the low-packet-loss path.
9. The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways according to claim 7 is characterized in that: Holographic data compression and transmission steps: Different holographic encoding strategies are used for the image data and numerical data of the acquisition end. After the image data is mapped to the high-dimensional holographic space, sparse representation technology is used to reduce redundancy and generate a compressed image feature matrix. The numerical data is vector quantized to generate a compressed feature vector, retaining key charging information. After the edge node completes the holographic encoding, it integrates the compressed image feature matrix and numerical feature vector into a unified data packet, and records the compression rate and encoding time of each data type; During cloud reassembly, a deep learning model is used to decode the compressed data packets and restore the original data structure. At the same time, the data packet loss rate is calculated through integrity verification. If the loss rate exceeds a preset threshold, a retransmission mechanism is triggered to retrieve the lost data packets from the edge node. The deep learning model is trained based on historical transmission data and adaptively adjusts decoding parameters to cope with changes in data distribution in different sections and time periods.
10. The method for real-time collection and fusion of multi-source heterogeneous toll collection data on highways according to claim 1 is characterized in that: In the deep fusion process, implicit time drift correction steps based on path topology and multi-scale time drift compensation are included, including: Construct a highway path topology map, use the longitude and latitude coordinates of toll booths and gantries to generate sequential constraints for vehicle travel paths, and calculate theoretical travel time ranges based on historical speed distribution and road speed limits; The distribution of timestamps recorded by each device is analyzed using short-term sliding windows and long-term sliding windows. The short-term window length ranges from 1 to 5 minutes, and the long-term window length ranges from 30 to 60 minutes. The mean shift and variance fluctuation characteristics of the time drift are extracted. Based on the path topology constraints and the extracted time drift features, a reinforcement learning algorithm based on Q-learning is used to dynamically correct the timestamp deviation of each device, and the rationality of vehicle speed and data alignment consistency are used as reward functions.
Citation Information
Patent Citations
Heterogeneous computing power platform design method, platform and resource scheduling method
CN117421108A
Smart city traffic management method and system based on multi-source data fusion
CN117912251A