Expressway multi-source heterogeneous charging data real-time acquisition and fusion method

Through machine learning, predicting traffic flow, combining chaotic dynamics and game theory to optimize computing power distribution, the problem of insufficient computing power and data consistency of highway toll systems in high-flow scenarios is solved, and efficient and safe data fusion and transmission are achieved.

CN120407654AActive Publication Date: 2025-08-01广西计算中心有限责任公司

Patent Information

Application Number
CN202510523839.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-01
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the prior art, in the highway toll system under high traffic or abnormal traffic mode, the computing power distribution of edge nodes lacks a dynamic response mechanism, resulting in insufficient computing power or waste of resources, affecting the real-time and consistency of data processing.

Method used

Through machine learning, predict traffic changes, combine chaotic dynamics to analyze computing power requirements, dynamically schedule computing power at edge nodes, and optimize fusion strategies based on trust and game theory, holographic data compression and error correction verification mechanisms are used to transmit data to realize dynamic computing power distribution and safe transmission.

Benefits of technology

It improves the integration efficiency and robustness of highway toll data processing, solves the problem of insufficient computing power in high-traffic scenarios, and optimizes data transmission delay and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407654A_ABST
    Figure CN120407654A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time collection and fusion method for highway multi-source heterogeneous toll data, and particularly relates to the technical field of highway toll data management, comprising the following steps: transmitting data to a computing power center deployed at the edge through a data link and network equipment for preliminary fusion to obtain a preliminary fusion data set; marking a computing power center deployed at the edge of each toll station as an edge node, predicting a future traffic flow change trend through machine learning, generating flow prediction distribution, extracting a computing power demand contour of each edge node based on the flow prediction distribution, analyzing nonlinear fluctuation characteristics of the computing power demand contour by using a chaos dynamics model, and calculating the traffic flow of each toll station. Chaos attractors of high-load nodes are identified, a cluster overall computing power boundary is generated, and computing power is dynamically scheduled based on a chaos control strategy; when a multi-source data fusion conflict is detected, a fusion strategy is optimized and adjusted based on the dynamic trust degree and the game theory; and the problem of insufficient computing power or resource waste in a high-load scene is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of highway toll data management. More specifically, the present invention relates to a method for real-time collection and fusion of multi-source heterogeneous toll data on expressways. Background Art

[0002] The expressway toll system is an important part of modern transportation infrastructure. It collects passing records and toll data in real time through multi-source devices such as gantries, lanes, and toll stations to support functions such as traffic flow monitoring, toll management, and data analysis. With the expansion of the expressway network and the increase in vehicle flow, the scale and complexity of data collection have increased significantly. The computing center deployed at the edge is responsible for the preliminary fusion of multi-source heterogeneous data and then transmits it to the cloud for in-depth processing. However, due to the influence of device heterogeneity, acquisition frequency differences, and dynamic traffic environments, edge nodes need to achieve efficient data cleaning, feature extraction, and preliminary fusion with limited computing power, while ensuring that the fusion strategy can adapt to the real-time changes in the quality of data sources. This scenario poses high requirements for the computing power of edge computing and the dynamics of fusion strategies.

[0003] In the prior art, data collection and preliminary fusion usually adopt static rules or simple load balancing methods. For example, data cleaning based on preset thresholds and fixed fusion strategies are common when processing data from gantries, lanes, and toll stations. However, these methods have significant defects: the computing power allocation of edge nodes lacks a dynamic response mechanism to traffic fluctuations and non-linear demands, resulting in insufficient computing power or resource waste in high-load scenarios, affecting the real-time performance and consistency of preliminary fusion. These problems are particularly prominent in high-traffic or abnormal traffic patterns (such as holiday peaks), directly restricting the efficiency and quality of expressway toll data processing. There is an urgent need for a technical solution that can dynamically optimize computing power allocation and adaptively adjust the fusion strategy to improve the processing ability at the edge. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method for real-time collection and fusion of multi-source heterogeneous toll data on expressways to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solution: A method for real-time collection and fusion of multi-source heterogeneous toll data on expressways, comprising the following steps:

[0006] Real-time collect the passing record data and toll data of gantries, lanes of expressway toll stations, and after cleaning, obtain the data at the collection end, and transmit it through a data link and network equipment to a computing center deployed at the edge for preliminary fusion to obtain a preliminary fusion data set;

[0007] The computing center deployed at the edge of each toll station is recorded as an edge node. By using machine learning to predict the future change trend of vehicle flow, generate a traffic flow prediction distribution, and extract the computing power demand profiles of each edge node based on the traffic flow prediction distribution. Machine learning includes: collecting historical and real-time traffic flow data, after normalization, performing time series analysis using a time series prediction model, mapping the traffic flow data to a high-dimensional feature representation, and adjusting parameters through an optimizer to output future traffic flow prediction values; calculating the computing power requirements of each edge node based on the traffic flow prediction values to form a computing power demand profile;

[0008] Using a chaotic dynamics model to analyze the non-linear fluctuation characteristics of the computing power demand profile, identify the chaotic attractor of high-load nodes, generate the overall cluster computing power boundary, and dynamically schedule computing power based on a chaos control strategy. Among them, if the computing power demand exceeds a preset ratio of the available computing power of the node, cross-node computing power borrowing and online learning are used to correct the prediction deviation;

[0009] When it is detected that there are conflicts in the keyword fields of multi-source data, the fusion strategy is optimized and adjusted based on dynamic trust and game theory.

[0010] Preferably, the chaotic attractor of high-load nodes refers to the set of trajectories that stably exist over time and are exhibited by high-load nodes during the fluctuation of computing power demand, reflecting the long-term behavior pattern of computing power; the overall cluster computing power boundary refers to the maximum computing power upper limit that the entire edge node cluster can provide within a given time period, taking into account the computing power resources and network bandwidth constraints of each node.

[0011] Preferably, a preliminary fusion data set is obtained through feature extraction and strategy selection. The preliminary fusion process includes the following steps:

[0012] Step S11, data source preparation and feature extraction: Using a meta-data analysis tool, extract the data source features of each data source in the multi-source heterogeneous toll collection data of the highway, including data structure, field type, data distribution, and collection frequency;

[0013] Step S12, real-time strategy selection and execution: Construct a strategy library containing multi-source heterogeneous data fusion methods, match the extracted data source features with the methods in the strategy library to generate candidate fusion strategies; select the optimal strategy and execute it according to the historical fusion effect and data quality indicators;

[0014] Step S13, data fusion: Execute the selected fusion strategy to integrate the multi-source heterogeneous data into a data set in a unified format.

[0015] Preferably, the preliminary fusion data set is transmitted to the highway toll network data center deployed at the cloud edge through a data link and network devices, and deep fusion is performed to output the fused highway toll data; the deep fusion includes position alignment, time alignment, and outlier correction.

[0016] Preferably, when there is information conflict in multi-source heterogeneous data, a fusion strategy adjustment step based on dynamic trust degree and game theory optimization is triggered, which specifically includes:

[0017] In real-time strategy selection and execution, based on the integrity, consistency, and timeliness indicators of the data at the acquisition end, the trust degree scores of each data source are calculated in real-time, and the fusion weights of each data source are dynamically adjusted using the Nash equilibrium model in game theory;

[0018] If the trust degree of a certain data source is lower than the preset threshold, the corresponding weight is reduced through a zero-sum game strategy, and an adversarial verification mechanism is introduced to verify the impact of the weight adjustment on the fusion consistency;

[0019] Combined with the computing power status feedback by the edge node in real-time, the payoff matrix of the game model is dynamically updated to ensure the coordinated optimization of weight adjustment and computing power allocation;

[0020] The consistency verification is performed on the adjusted fusion result and the trust degree change log, and an analysis report including the game optimization effect is generated.

[0021] Preferably, the method for obtaining the trust degree score is as follows:

[0022] Based on the field completion rate FCR, the percentage of missing records MRP, and the standard deviation of the missing field distribution, the dynamic integrity quality index DCIQ is quantitatively represented;

[0023] Based on the non-missing record field matching rate, the non-missing record conflict ratio, and the standard deviation of the conflict field distribution, the multi-source consistency quality index MCIQ is quantitatively represented;

[0024] Based on the update frequency UF, the expected update frequency, the average delay time DL, and the standard deviation of the delay distribution, the real-time timeliness quality index RTI is quantitatively represented;

[0025] The dynamic integrity quality index, the multi-source consistency quality index, and the real-time timeliness quality index are jointly analyzed to obtain the trust degree score CQI.

[0026] Preferably, the method further includes a transmission network optimization step, which specifically includes:

[0027] Step S31, optimizing the bandwidth allocation and traffic scheduling through a heuristic optimization algorithm;

[0028] Step S32, Adaptive Neural Network Topology Optimization: In the network between the edge nodes and the cloud, deploy a network topology simulated by a neural network to mimic the adaptive characteristics of the connection of human brain neurons; use a generative adversarial network to dynamically generate an optimal network path and respond to topological changes in real time;

[0029] Step S33, Holographic Data Compression and Transmission: At the edge node, perform holographic encoding on the data from the acquisition end, map multi-source heterogeneous data to a high-dimensional holographic space, and generate compressed data packets;

[0030] Step S34, Combine redundant encoding with a multi-path transmission protocol, fragment the data packets and transmit them in parallel through multiple paths for recombination at the cloud.

[0031] Preferably, when generating holographic encoded data packets at the edge node, use redundant encoding and error correction verification mechanisms to bind the data packets and error correction bits into entangled pairs; during the multi-path transmission process, monitor the packet loss rate of each path in real time. If the packet loss rate of a certain path exceeds the preset value, dynamically adjust the allocation ratio of the entangled pairs and preferentially transmit the error correction bits to the low packet loss path.

[0032] Preferably, in the step of holographic data compression and transmission:

[0033] Adopt different holographic encoding strategies for the image data and numerical data of the data from the acquisition end respectively. After the image data is mapped to the high-dimensional holographic space, use sparse representation technology to reduce redundancy and generate a compressed image feature matrix; the numerical data generates a compressed feature vector through vector quantization to retain key toll information;

[0034] After completing holographic encoding at the edge node, integrate the compressed image feature matrix and numerical feature vector into a unified data packet, and record the compression rate and encoding time of each data type;

[0035] When recombining at the cloud, use a deep learning model to decode the compressed data packet to restore the original data structure. At the same time, calculate the loss rate of the data packet through integrity verification. If the loss rate exceeds the preset threshold, trigger a retransmission mechanism to obtain the lost data packet from the edge node again; the deep learning model is trained based on historical transmission data and adaptively adjusts the decoding parameters to cope with the data distribution changes in different road sections and time periods.

[0036] Preferably, during the deep fusion process, it includes an implicit time drift correction step based on path topology and multi-scale time drift compensation, including:

[0037] Construct a highway path topology map, use the longitude and latitude coordinates of toll stations and gantries to generate the sequential constraints of the vehicle passing path, and calculate the theoretical passing time range according to the historical speed distribution and road speed limit;

[0038] Analyze the timestamp distributions recorded by each device through short - period sliding windows and long - period sliding windows respectively. The short - period window length is from 1 to 5 minutes, and the long - period window length is from 30 to 60 minutes. Extract the mean offset and variance fluctuation characteristics of time drift.

[0039] Based on the path topology constraints and the extracted time - drift characteristics, use a reinforcement learning algorithm based on Q - learning to dynamically correct the timestamp deviation of each device, with the rationality of vehicle passing speed and data alignment consistency as the reward function.

[0040] The technical effects and advantages of the present invention:

[0041] (1) The real - time acquisition and fusion method for multi - source heterogeneous toll data on highways provided by the present invention collects data from gantries, lanes, and toll stations in real - time, performs preliminary fusion after cleaning, uses machine learning to predict traffic flow, combines chaotic dynamics analysis to calculate the fluctuation of computing power requirements, and dynamically schedules the computing power of edge nodes to solve the problem of insufficient computing power in high - traffic scenarios. At the same time, based on integrity, consistency, and timeliness indicators, calculate the trustworthiness scores of data sources, and use game theory (Nash equilibrium and zero - sum game) to dynamically adjust the fusion weights to improve data consistency, significantly enhancing the fusion efficiency and robustness.

[0042] (2) The real - time acquisition and fusion method for multi - source heterogeneous toll data on highways provided by the present invention effectively solves the problems of delay, packet loss, and security in the transmission of highway toll data through a heuristic optimization algorithm for dynamic bandwidth allocation and traffic prediction, an adaptive neural network for dynamically adjusting the transmission path, holographic data compression combined with redundancy coding and error - correction check mechanisms, and a zero - trust architecture for strengthening security verification. The holographic compression technology maps multi - source heterogeneous data to a high - dimensional space to solve the problem of data transmission delay in the process of highway toll data transmission. Brief Description of the Drawings

[0043] Figure 1 It is a flowchart of the real - time acquisition and fusion method for multi - source heterogeneous toll data on highways of the present invention.

[0044] Figure 2 It is a block diagram of the preliminary fusion process structure of the present invention.

[0045] Figure 3 It is a flowchart of the transmission network optimization of the present invention. Detailed Embodiments

[0046] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0047] At the same time, it should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0048] The following description of at least one exemplary embodiment is actually merely illustrative and in no way limits the present application and its application or use.

[0049] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the specification.

[0050] Example 1, referring to Figure 1 the flowchart of the real-time acquisition and fusion method for multi-source heterogeneous toll data on expressways and Figure 2 the block diagram of the preliminary fusion process structure of , the present invention provides a real-time acquisition and fusion method for multi-source heterogeneous toll data on expressways, including the following steps:

[0051] Step 1: Real-time collect the gantry, lane passing record data, and toll data of expressway toll stations, clean them to obtain the acquisition-end data of the expressway, and transmit them to the computing center deployed at the edge through the data link and network devices for preliminary fusion to obtain a preliminary fusion data set; Explanation: The multi-source heterogeneous toll data on expressways consists of the gantry, lane passing record data, and toll data of expressway toll stations.

[0052] Step 2: Transmit the preliminary fusion data set to the expressway toll network data center deployed at the cloud edge through the data link and network devices, and perform deep fusion to output the fused expressway toll data; The deep fusion includes obtaining high-quality, uniformly formatted expressway toll data based on location alignment, time alignment, and outlier correction, making up for the insufficient computing power of edge nodes, and improving data consistency and availability.

[0053] Step 3: Store the fused expressway toll data in a centralized database or data warehouse to support data query and statistical analysis. For example, index the fused expressway toll data by time, toll station, or vehicle type to facilitate generating traffic flow reports or toll statistics.

[0054] Step 4: Real-time monitor the operating status of the data link and the edge computing center, and trigger an alarm if an abnormality is detected.

[0055] In this implementation, what needs to be further explained is that a preliminary fusion data set is obtained through feature extraction and strategy selection. The preliminary fusion process includes the following steps:

[0056] Step S11, data source preparation and feature extraction: Using a metadata analysis tool, extract the data source features of each data source in the multi-source heterogeneous toll data of the highway, including data structure, field type, data distribution, and collection frequency;

[0057] Step S12, real-time strategy selection and execution: Build a strategy library containing multi-source heterogeneous data fusion methods, match the extracted data source features with the methods in the strategy library to generate candidate fusion strategies; Select the optimal strategy and execute it according to the historical fusion effect and data quality indicators;

[0058] Step S13, data fusion: Execute the selected fusion strategy to integrate the multi-source heterogeneous data into a data set in a unified format.

[0059] In this implementation, what needs to be further explained is that the computing power center deployed at the edge of each toll station is recorded as an edge node; The method includes steps for allocating computing power to edge nodes, including the following steps:

[0060] Step S21, predict the traffic flow change trend in the next hour through machine learning, generate a traffic flow prediction distribution, and extract the computing power demand profile of each edge node based on the traffic flow prediction distribution. The profile is based on the data processing load per unit time;

[0061] The detailed implementation process includes: Real-time collection of traffic flow data from the gantries, lanes, and toll station terminals of highway toll stations, including the number of vehicle passages, average speed, and timestamp, etc. Integrate historical data (it is recommended to have at least 6 months) as the training set, including context variables such as holidays and weather conditions; Normalize the data, map the traffic flow values to the interval [0, 1] to avoid the influence of dimensional differences on model training;

[0062] Perform statistical processing on the predicted values, calculate the traffic flow mean and standard deviation every 5 minutes; Use kernel density estimation (KDE) to smooth the prediction curve and generate a continuous traffic flow distribution function; Store the results as a time series for subsequent computing power demand calculation;

[0063] Define the data processing load per unit time. According to the system benchmark test, assume that each passage record requires 0.1 unit of computing power (in units of CPU cycles or FLOPS); Calculate the computing power demand of each edge node. For each 5-minute time period, the demand is the product of the traffic flow prediction mean and the computing power per unit record; Organize the results by node and time to form a computing power demand profile;

[0064] Step S22: Analyze the non-linear fluctuation characteristics of the computing power demand profile using a chaotic dynamics model, identify the chaotic attractors of high-load nodes, and generate the overall computing power boundary of the cluster. Explanation: The chaotic attractors of high-load nodes refer to the set of trajectories that stably exist over time in the computing power demand fluctuations, exhibited by high-load nodes (such as toll stations processing a large number of vehicle passing records), reflecting the long-term behavior patterns of computing power. The overall computing power boundary of the cluster refers to the maximum computing power upper limit that the entire edge node cluster (i.e., the computing power centers of all toll stations) can provide within a given time period, considering constraints such as the computing power resources and network bandwidth of each node.

[0065] The detailed implementation process includes: Using the generated computing power demand profile as input, organizing it by edge node and time series; Using the chaospy library in Python or the chaotic analysis module in MATLAB to construct a chaotic model based on the Logistic map, normalizing the computing power demand to the interval [0, 1], setting control parameters to generate a chaotic sequence, and running 1000 iterations to ensure the sequence enters a stable chaotic state; Calculating the Lyapunov exponent, if it is positive (for example, 0.5), confirm that the demand fluctuations have chaotic characteristics; Using delay embedding technology to reconstruct the phase space; Statistically analyzing the peak computing power demands of all nodes, calculating the total available computing power of the cluster, and integrating it by time period to form a dynamic boundary curve.

[0066] Step S23: If the computing power demand exceeds the preset value, for example, exceeding the available computing power of the node by 20%, it is determined as a high-load type. Based on the chaotic control strategy, dynamically schedule the computing power, borrow computing power across nodes and correct the prediction deviation through online learning, and adaptively adjust the borrowing ratio according to the stability of the chaotic attractor.

[0067] Step S24: Real-time monitor the scheduling effect. If the consistency error is lower than 5%, feedback the scheduling strategy to the machine learning model to optimize the prediction for the next cycle.

[0068] In this implementation, it needs to be further explained that when there is information conflict in multi-source heterogeneous data and it is detected that there is a conflict in the key fields of multi-source data (such as the time stamp difference exceeds a reasonable threshold), then trigger the fusion strategy adjustment step based on dynamic trust degree and game theory optimization, which specifically includes:

[0069] In real-time policy selection and execution, based on the integrity, consistency, and timeliness indicators of the data at the acquisition end, calculate the trust degree scores of each data source in real-time, and use the Nash equilibrium model in game theory to dynamically adjust the fusion weights of each data source.

[0070] If the trust degree of a certain data source is lower than the preset threshold, reduce the corresponding weight through the zero-sum game strategy, and introduce an adversarial verification mechanism to verify the impact of the weight adjustment on the fusion consistency.

[0071] Combined with the computing power status of real-time feedback from edge nodes, dynamically update the payoff matrix of the game model to ensure the collaborative optimization of weight adjustment and computing power allocation;

[0072] Verify the consistency of the adjusted fusion result with the trustworthiness change log, and generate an analysis report including the game optimization effect.

[0073] Furthermore, the integrity index, consistency index, and timeliness index of the data at the acquisition end are respectively quantified as the dynamic integrity quality index, multi-source consistency quality index, and real-time timeliness quality index; the dynamic integrity quality index DCIQ is quantified based on the field completeness rate FCR, the percentage of missing records MRP, and the standard deviation of the missing field distribution; the multi-source consistency quality index MCIQ is quantified based on the matching rate of non-missing record fields, the conflict ratio of non-missing records, and the standard deviation of the conflict field distribution; the real-time timeliness quality index RTI is quantified based on the update frequency UF, the expected update frequency, the average delay time DL, and the standard deviation of the delay distribution; jointly analyze the dynamic integrity quality index, multi-source consistency quality index, and real-time timeliness quality index to obtain the trustworthiness score CQI.

[0074] In a possible embodiment, the integrity index is used to measure the missing and null values of key fields (such as timestamps, license plates, passing records, etc.) in the data at the acquisition end; the quantification method of the integrity index is: calculate the complete rate of required fields for each data packet, that is, the ratio of the actual filled value of the key field to the value that should be filled; use the sliding window technology to calculate the percentage of missing records within a certain time window;

[0075] In a possible embodiment, the consistency index is used to evaluate the matching degree of multi-source data in the data at the acquisition end in terms of logical relationship, data format, and semantics; the quantification method of the consistency index is: for the same vehicle and the same passing event, perform field matching in different data sources and calculate the matching rate of field values; count and analyze the proportion of records that conflict with each other or have inconsistent data formats to evaluate the data consistency;

[0076] In a possible embodiment, the timeliness index is used to measure the delay time of the data at the acquisition end from the acquisition end to the computing power center, as well as the real-time update frequency of the data. The real-time update frequency of the data refers to the number of times the data is updated within a specific time, reflecting the freshness and timeliness of the data; the quantification method of the delay time is: record timestamps at each link of data acquisition, transmission, and processing, and determine the delay time by calculating the time difference between each link; the quantification method of the real-time update frequency of the data is: by setting a time window, record the number of times the data is updated within the window, and calculate the ratio of the number of updates to the window duration to obtain the update frequency of the data at the acquisition end.

[0077] For ease of understanding, the specific method for quantifying the trust score provided by the embodiments of the present invention is as follows:

[0078] Through the formula the dynamic integrity quality index DCIQ is obtained; the calculation method of the field completion rate FCR is the ratio of the actual number of filled fields to the number of fields to be filled within the time window; the calculation method of the missing record percentage MRP is the ratio of the number of missing records to the total number of records within the time window; σ m issing represents the standard deviation of the missing field distribution;

[0079] Through the formula the multi-source consistency quality index MCIQ is obtained, and the non-missing record field matching rate FMR non-missing is calculated as the ratio of the number of matching fields to the total number of comparison fields among non-missing records; the non-missing record conflict ratio CRP non-missing is calculated as the ratio of the number of conflict records to the total number of non-missing records among non-missing records; the standard deviation of the conflict field distribution σ c onflict is calculated as the standard deviation of the number of conflict fields among non-missing records;

[0080] The expected update frequency UF e xp, the average delay time DL, and the standard deviation of the delay distribution σ delay are obtained. Through the formula the real-time timeliness quality index RTI is obtained. The calculation method of the update frequency UF is the ratio of the number of updates to the window duration within the time window;

[0081] By jointly analyzing the dynamic integrity quality index (DCIQ), the multi-source consistency quality index (MCIQ), and the real-time timeliness quality index (RTIQ), the trust score CQI is obtained;

[0082]

[0083] Explanation: The numerator is used to comprehensively consider the contributions of the three and emphasize the short-board effect; if any of the indicators (DCIQ, MCIQ, RTIQ) is too low, the CQI will decrease significantly; in the denominator, the constant 3 (the sum of the maximum values of the three) ensures a reasonable range of the denominator, and the square root term is used for non-linear punishment of low-value indicators.

[0084] Furthermore, the steps for optimizing the time window selection based on the traffic pattern specifically include: using historical traffic flow data, speed data, and real-time gantry traffic counts, training a traffic pattern classifier using the random forest algorithm to identify peak, trough, and congestion patterns; dynamically adjusting the time window length according to the traffic pattern, matching a short window length for the peak period. For example, the window length for the peak period is 1 to 5 minutes, the window length for the trough period is 30 to 60 minutes, and the window length for the congestion period is 10 to 20 minutes.

[0085] Summary: In the embodiments of the present invention, by collecting data of gantries, lanes, and toll stations in real time, performing preliminary fusion after cleaning, predicting traffic flow using machine learning, analyzing the fluctuation of computing power requirements in combination with chaotic dynamics, and dynamically scheduling the computing power of edge nodes, the problem of insufficient computing power in high-traffic scenarios is solved; at the same time, the trustworthiness scores of data sources are calculated based on integrity, consistency, and timeliness indicators, and game theory (Nash equilibrium and zero-sum game) is used to dynamically adjust the fusion weights to improve data consistency. After the fusion result passes the consistency verification, a preliminary data set is output and fed back to optimize the prediction model; traffic pattern classification is also introduced to dynamically adjust the time window to adapt to different traffic scenarios.

[0086] Embodiment 2. To solve the existing data transmission delay problem, it is applied to the process of transmitting the collected data initially fused at the edge-deployed computing center through the data link and network devices to the highway toll network data center deployed at the cloud edge. Refer to Figure 3 the transmission network optimization flowchart. The method further includes transmission network optimization steps, specifically including the following:

[0087] Step S31: Optimize bandwidth allocation and traffic scheduling through a heuristic optimization algorithm; the specific implementation process is as follows: In the data link between the edge node and the cloud, deploy a heuristic optimization algorithm to optimize bandwidth allocation; based on historical traffic flow data and real-time road conditions (such as holiday peaks, accident-prone periods), predict the traffic flow change trend in the next 5 - 30 minutes through machine learning, analyze multiple variables in parallel (such as vehicle speed, toll station distribution, network load), and generate a dynamic bandwidth scheduling strategy to solve the computational bottleneck of traditional deep learning in optimizing ultra-large-scale variables;

[0088] Step S32: Adaptive neural network topology optimization: In the network between the edge node and the cloud (the highway toll network data center deployed at the cloud edge), deploy a network topology simulated by a neural network to imitate the adaptive characteristics of the connection of human brain neurons; use a generative adversarial network to dynamically generate the optimal network path and respond to topology changes in real time;

[0089] Step S33: Holographic data compression and transmission: Perform holographic encoding on the collected data at the edge node, map multi-source heterogeneous data (such as images recorded by gantries, numerical values of toll data) to a high-dimensional holographic space, and generate compressed data packets;

[0090] Step S34: Combine redundant encoding with a multi-path transmission protocol, fragment the data packets and transmit them in parallel through multiple paths, and reconstruct them at the cloud.

[0091] Furthermore, in holographic data compression and transmission, an adaptive error correction mechanism is adopted to enhance the robustness of multi-path parallel transmission. The specific implementation process is as follows: When generating holographic encoded data packets at the edge node, the redundant encoding and error correction verification mechanism is used to bind the data packet and the error correction bit into an entangled pair. During the multi-path transmission process, the packet loss rate of each path is monitored in real time. If the packet loss rate of a certain path exceeds the preset value, such as 3%, the allocation ratio of the entangled pair is dynamically adjusted, and the error correction bit is preferentially transmitted to the path with a low packet loss rate. When reconstructing at the cloud, the lost data is recovered based on the measurement technology of the entangled state, and the error correction success rate can reach more than 98%. The adaptive error correction mechanism combined with the entanglement heuristic protocol breaks through the limitations of traditional error correction methods (such as Reed-Solomon codes) in a dynamic network environment and improves the transmission reliability of highway toll data in high-dynamic and multi-interference scenarios.

[0092] Furthermore, the security real-time optimization under the zero-trust architecture: During the data transmission process, the zero-trust network architecture is implemented, and each data stream needs to pass through dynamic authentication (such as digital signature based on vehicle ID) and integrity check (such as hash check). Combined with lightweight encryption, security is guaranteed without increasing latency. The security real-time optimization steps under the zero-trust architecture include: In dynamic authentication, a unique time-sensitive digital signature is generated based on the vehicle ID and the toll station location, and the signature is encrypted through a lightweight encryption algorithm (where the time sensitivity is accurately specified to milliseconds by attaching the current timestamp) to ensure that the signature is valid within a specific time window, such as 5 seconds. In integrity check, the hash value of the data stream is calculated in real time and compared with the pre-generated check value at the edge node. If an abnormality is found, the abnormal event is recorded and an alarm is triggered. For example, if the two are inconsistent, it is determined that the data may have been tampered with, and the time of the abnormal event, the data stream identifier, and the details of the difference are recorded, and an alarm is triggered to notify the computing power center.

[0093] In the embodiments of the present invention, it needs to be further explained that in the holographic data compression and transmission steps, it further includes:

[0094] Different holographic encoding strategies are adopted for the image data and numerical data of the data at the acquisition end. After the image data is mapped to a high-dimensional holographic space, the sparse representation technology is used to reduce redundancy and generate a compressed image feature matrix. The numerical data generates a compressed feature vector through vector quantization to retain the key toll information.

[0095] After holographic encoding is completed at the edge node, the compressed image feature matrix and numerical feature vector are integrated into a unified data packet, and the compression rate and encoding time of each data type are recorded.

[0096] When reorganizing in the cloud, a deep learning model is used to decode the compressed data packets to restore the original data structure. At the same time, the loss rate of the data packets is calculated through integrity verification. If the loss rate exceeds a preset threshold (such as 5%), the retransmission mechanism is triggered to obtain the lost data packets from the edge nodes again; the deep learning model is trained based on historical transmission data and adaptively adjusts the decoding parameters to cope with the data distribution changes in different road sections and time periods.

[0097] In summary, the embodiments of the present invention effectively solve the problems of delay, packet loss, and security in highway toll data transmission by optimizing bandwidth allocation and traffic prediction, adaptively adjusting the transmission path by a neural network, holographic data compression combined with an entanglement error correction mechanism, and strengthening security verification with a zero-trust architecture; the holographic compression technology maps multi-source heterogeneous data to a high-dimensional space to solve the problem of data transmission delay in the process of highway toll data transmission.

[0098] Embodiment 3, the difference between the embodiments of the present invention and Embodiments 1 and 2 is that the deep fusion includes the following steps:

[0099] Data alignment: Obtain the preliminarily fused acquisition-side data with timestamps transmitted by each edge node, and align them based on the location of each edge node (such as the longitude and latitude of the toll station) and the timestamp.

[0100] Outlier correction: Repair data errors through statistical rules or machine learning methods, including correcting license plate recognition deviations and filling in missing fields.

[0101] Verify the logical relationship between multi-source data. For records that fail the verification, corrective measures (such as calibration, interpolation, confidence selection) are preferentially taken. If the records cannot be corrected, the conflicting records are excluded, and the reasons for the anomalies are recorded.

[0102] Furthermore, the process of verifying the logical relationship between multi-source data at least includes verifying the rationality of the toll amount and the passing time, and excluding conflicting records.

[0103] Verify the matching of the vehicle identification (license plate number) and the passing record (the license plate number of the same vehicle should be consistent in the records of gantries, lanes, and toll stations).

[0104] Verify the sequential consistency of the passing time and the geographical location (verify whether the time sequence of vehicle passing conforms to the actual path according to the timestamp and the longitude and latitude of the toll station).

[0105] Verify the rationality of the passing speed and the time interval (the time interval between adjacent toll stations should be reasonably matched with the vehicle passing speed and the distance).

[0106] It is necessary to further explain in the embodiment of the present invention that, in the deep fusion process, an implicit time drift correction step based on path topology and multi-scale time drift compensation is included, specifically including:

[0107] Construct a highway path topology map, use the longitude and latitude coordinates of toll booths and gantries to generate sequential constraints for vehicle travel paths, and calculate theoretical travel time ranges based on historical speed distribution and road speed limits;

[0108] The distribution of timestamps recorded by each device is analyzed using short-term sliding windows and long-term sliding windows. The short-term window length ranges from 1 to 5 minutes, and the long-term window length ranges from 30 to 60 minutes. The mean shift and variance fluctuation characteristics of the time drift are extracted.

[0109] Based on the path topology constraints and the extracted time drift features, a Q-learning-based reinforcement learning algorithm is used to dynamically correct the timestamp deviation of each device, with the rationality of vehicle speed and data alignment consistency as the reward function;

[0110] Furthermore, each edge node is collaboratively optimized by sharing time drift characteristics and correction strategies to build a global time drift correction network. The correction step size is dynamically adjusted between edge nodes through a joint reward function based on reinforcement learning (combining local consistency and global consistency), prioritizing the correction of drift deviations in high-traffic nodes, thereby improving the time synchronization accuracy of the entire network and the consistency of fused data.

[0111] Explanation: The joint reward function is a function used in reinforcement learning to evaluate the correction effect. It combines local consistency (data alignment accuracy of a single node, such as timestamp deviation < 1 second) and global consistency (overall accuracy of data alignment of all nodes, such as license plate matching rate > 95%) to comprehensively optimize time correction. High-traffic nodes refer to edge nodes that process a large number of vehicle passage records within a specific time period. For example, during peak holiday periods, a toll station processes 200 records per minute, far higher than the average of 50.

[0112] The corrected timestamp is applied to the data alignment step, and the detected time drift anomaly is fed back to the equipment maintenance module through the network alarm system.

[0113] Summary: High-quality fusion is achieved through data alignment (based on location and timestamp), outlier correction (statistical rules or machine learning to correct license plate deviations and complete missing fields), and verification of the logical relationships of multi-source data (including toll amount reasonableness, license plate matching, passing order, and speed consistency). For records that fail the verification, calibration or interpolation is preferred first. If it cannot be corrected, it is excluded and the anomaly is recorded. In addition, the method introduces implicit time drift correction based on path topology and multi-scale time drift compensation, constructs a path topology graph, extracts drift features using short-term (1-5 minutes) and long-term (30-60 minutes) sliding windows, and uses Q-learning reinforcement learning to dynamically correct timestamp deviations. Further, through edge node collaborative optimization, sharing drift features and correction strategies, a global time drift correction network is constructed, and the correction step size is dynamically adjusted in combination with a joint reward function (fusing local and global consistency), giving priority to correcting high-traffic nodes to improve time synchronization accuracy.

[0114] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A real-time acquisition and fusion method for multi-source heterogeneous toll collection data on expressways, characterized in that, Including: Real-time collect the portal frame, lane passing record data and toll data of highway toll stations, and obtain the collected data at the collection end after cleaning. Then transmit it to the edge-deployed computing center through the data link and network devices for preliminary fusion to obtain the preliminary fusion data set; Denote the computing center edge-deployed at each toll station as an edge node. Predict the future traffic flow change trend through machine learning, generate the traffic flow prediction distribution, and extract the computing power demand profile of each edge node based on the traffic flow prediction distribution. Machine learning includes: collect historical and real-time traffic flow data, after normalization processing, use the time series prediction model for time series analysis, map the traffic flow data to a high-dimensional feature representation, and adjust the parameters through the optimizer to output the future traffic flow prediction value; calculate the computing power demand of each edge node based on the traffic flow prediction value to form the computing power demand profile; Use the chaotic dynamics model to analyze the non-linear fluctuation characteristics of the computing power demand profile, identify the chaotic attractor of high-load nodes, generate the overall computing power boundary of the cluster, and dynamically schedule the computing power based on the chaos control strategy. Among them, if the computing power demand exceeds the preset ratio of the available computing power of the node, borrow the computing power across nodes and correct the prediction deviation through online learning; When detecting conflicts in keyword fields of multi-source data, optimize and adjust the fusion strategy based on dynamic trust and game theory.

2. The real-time acquisition and fusion method of multi-source heterogeneous toll data for expressways according to claim 1, characterized in that The chaotic attractor of high-load nodes refers to the set of trajectories that stably exist over time shown by high-load nodes during the fluctuation of computing power demand, reflecting the long-term behavior pattern of computing power; The overall computing power boundary of the cluster refers to the maximum computing power upper limit that the entire edge node cluster can provide within a given time period, considering the computing power resources and network bandwidth constraints of each node.

3. A real-time acquisition and fusion method for multi-source heterogeneous toll data on expressways according to claim 1, characterized in that Obtain the preliminary fusion data set through feature extraction and strategy selection. The preliminary fusion process includes the following steps: Step S11, data source preparation and feature extraction: Use the metadata analysis tool to extract the data source features of each data source in the multi-source heterogeneous toll data of the highway, including data structure, field type, data distribution, and collection frequency; Step S12, real-time strategy selection and execution: Build a strategy library containing multi-source heterogeneous data fusion methods, match the extracted data source features with the methods in the strategy library to generate candidate fusion strategies; select the optimal strategy and execute it according to the historical fusion effect and data quality indicators; Step S13, data fusion: Execute the selected fusion strategy to integrate the multi-source heterogeneous data into a unified format data set.

4. A real-time acquisition and fusion method for multi-source heterogeneous toll collection data on expressways according to claim 1, characterized in that, Including: Transmit the preliminary fusion data set to the highway toll network data center deployed at the cloud edge through the data link and network devices, and perform deep fusion to output the fused highway toll data; the deep fusion includes location alignment, time alignment, and outlier correction.

5. A real-time acquisition and fusion method for multi-source heterogeneous toll collection data on expressways according to claim 3, characterized in that When there is an information conflict in multi-source heterogeneous data, trigger the fusion strategy adjustment step optimized based on dynamic trust and game theory, specifically including: In real-time strategy selection and execution, based on the integrity, consistency, and timeliness indicators of the collected data at the collection end, calculate the trust score of each data source in real-time, and dynamically adjust the fusion weight of each data source using the Nash equilibrium model in game theory; If the trustworthiness of a certain data source is lower than the preset threshold, the corresponding weight is reduced through a zero-sum game strategy, and an adversarial verification mechanism is introduced to verify the impact of weight adjustment on fusion consistency; Combined with the computing power status feedback by the edge nodes in real time, the payoff matrix of the game model is dynamically updated to ensure the collaborative optimization of weight adjustment and computing power allocation; The adjusted fusion result is verified for consistency with the trustworthiness change log, and an analysis report containing the game optimization effect is generated.

6. The real-time acquisition and fusion method of multi-source heterogeneous toll data for expressways according to claim 5, characterized in that, The method for obtaining the trustworthiness score is as follows: Quantitatively represent the dynamic integrity quality index DCIQ based on the field completeness rate FCR, the percentage of missing records MRP, and the standard deviation of the missing field distribution; Quantitatively represent the multi-source consistency quality index MCIQ based on the non-missing record field matching rate, the non-missing record conflict ratio, and the standard deviation of the conflict field distribution; Quantitatively represent the real-time timeliness quality index RTI based on the update frequency UF, the expected update frequency, the average delay time DL, and the standard deviation of the delay distribution; Jointly analyze the dynamic integrity quality index, the multi-source consistency quality index, and the real-time timeliness quality index to obtain the trustworthiness score CQI.

7. A real-time acquisition and fusion method for multi-source heterogeneous toll data on expressways according to any one of claims 1-4, characterized in that The method further includes a transmission network optimization step, which specifically includes: Step S31, optimize the bandwidth allocation and traffic scheduling through a heuristic optimization algorithm; Step S32, adaptive neural network topology optimization: in the network between the edge nodes and the cloud, deploy a network topology simulated by a neural network to imitate the adaptive characteristics of the connection of human brain neurons; use a generative adversarial network to dynamically generate the optimal network path and respond to topology changes in real time; Step S33, holographic data compression and transmission: perform holographic encoding on the data collected by the edge nodes, map the multi-source heterogeneous data to a high-dimensional holographic space, and generate compressed data packets; Step S34, combine redundant coding with a multi-path transmission protocol, fragment the data packets and transmit them in parallel through multiple paths, and recombine them in the cloud.

8. A real-time acquisition and fusion method for multi-source heterogeneous toll data on expressways according to claim 7, characterized in that When generating holographic encoded data packets at the edge nodes, use redundant coding and error correction verification mechanisms to bind the data packets and error correction bits into entangled pairs; during the multi-path transmission process, monitor the packet loss rate of each path in real time. If the packet loss rate of a certain path exceeds the preset value, dynamically adjust the allocation ratio of the entangled pairs, and preferentially transmit the error correction bits to the low packet loss path.

9. A real-time acquisition and fusion method for multi-source heterogeneous toll data on expressways according to claim 7, characterized in that, In the holographic data compression and transmission step: Adopt different holographic encoding strategies for the image data and numerical data of the data collected by the acquisition end. After the image data is mapped to a high-dimensional holographic space, use sparse representation technology to reduce redundancy and generate a compressed image feature matrix; the numerical data generates compressed feature vectors through vector quantization, and retains the key charging information; After completing the holographic encoding at the edge nodes, integrate the compressed image feature matrix and numerical feature vectors into a unified data packet, and record the compression rate and encoding time of each data type; When recombining in the cloud, use a deep learning model to decode the compressed data packet to restore the original data structure, and at the same time calculate the packet loss rate of the data packet through integrity verification. If the packet loss rate exceeds the preset threshold, trigger the retransmission mechanism and re-obtain the lost data packet from the edge nodes; The deep learning model is trained based on historical transmission data and adaptively adjusts the decoding parameters to cope with the data distribution changes in different road sections and time periods.

10. A real-time acquisition and fusion method for multi-source heterogeneous toll data on expressways according to claim 1, characterized in that, During the deep fusion process, it includes an implicit time drift correction step based on path topology and multi-scale time drift compensation, including: Construct a highway path topology map, generate sequential constraints of vehicle passing paths using the longitude and latitude coordinates of toll stations and gantries, and calculate the theoretical passing time range according to the historical speed distribution and road speed limits; Analyze the timestamp distributions recorded by each device through short-period sliding windows and long-period sliding windows respectively, where the short-period window length is 1 to 5 minutes and the long-period window length is 30 to 60 minutes, and extract the mean shift and variance fluctuation characteristics of time drift; Based on the path topology constraints and the extracted time drift characteristics, use a reinforcement learning algorithm based on Q-learning to dynamically correct the timestamp deviations of each device, with the rationality of vehicle passing speed and the consistency of data alignment as the reward function.

Citation Information

Patent Citations

  • Heterogeneous computing power platform design method, platform and resource scheduling method

    CN117421108A

  • Smart city traffic management method and system based on multi-source data fusion

    CN117912251A

  • Heterogeneous computing power fusion and dynamic optimization distribution method and system

    CN119847736A

Cited By

  • Cloud edge collaborative data processing method and system for remote under-pressure operation

    CN120856742A

  • Method and device for determining traffic volume data based on multi-source data, equipment and medium

    CN120998024A