A public transport travel service big data processing method and system
By generating spatiotemporal grid entities and a four-dimensional rights and interests rule decision tree, the problems of public transportation data fusion and resource scheduling are solved, achieving efficient, secure, and traceable data processing and route planning, and improving the system's dynamic response capability and compliance.
Patent Information
- Application Number
- CN202511082273.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-04
AI Technical Summary
Existing technologies in the public transportation sector face challenges in fusing multi-source heterogeneous data, failing to dynamically respond to environmental changes, lacking flexible arbitration mechanisms for ownership management, and failing to combine real-time confidence arbitration and circuit breaker control in route planning, resulting in instruction transmission delays and insufficient robustness of dynamic decision-making during resource competition.
By calibrating the coordinate deviation of heterogeneous data, a spatiotemporal grid entity is generated. A priority decision tree is constructed based on four-dimensional rights and interests rules, triggering a forced preemption strategy, executing a three-level dynamic rule chain, generating a confidence arbitration result, and verifying the feasibility of the path, thus forming a closed-loop governance process.
It achieves spatiotemporal consistency of multi-source data, ensures the priority of resource allocation for high-priority tasks, reduces the risk of path infeasibility, meets compliance requirements, and enhances the system's adaptability.
Smart Images

Figure CN120579797B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of public transportation service technology, and more specifically, to a method and system for big data processing of public transportation services. Background Technology
[0002] The current development of intelligent transportation faces the challenge of integrating multi-source heterogeneous data. Data generated in the public transportation sector, such as bus GPS, shared bicycle orders, and ride-hailing trajectories, form serious data silos due to differences in protocols and coordinate systems. Traditional static processing solutions are unable to meet the real-time scheduling needs of emergencies. Furthermore, data ownership conflicts among stakeholders such as enterprises, individuals, and operators lead to low efficiency in cross-stakeholder collaboration, significant delays in instruction transmission during resource competition, and insufficient robustness in dynamic decision-making. There is an urgent need to build a new technical system that takes into account real-time performance, ownership compliance, and closed-loop self-optimization.
[0003] Existing solutions rely on manual calibration of protocol deviations to address data silos, which is inefficient and unable to dynamically respond to environmental changes. Ownership management uses a fixed priority strategy, which lacks a flexible arbitration mechanism during resource contention, resulting in the obstruction of high-priority instructions. Path planning does not combine real-time confidence arbitration and circuit breaker control, resulting in a high proportion of failed instructions. Optimization of historical issues relies on manual rule adjustments, lacks parameter self-iteration capabilities, and makes it difficult to continuously improve system adaptability. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution:
[0005] A method and system for big data processing of public transportation travel services, comprising:
[0006] S1: Access heterogeneous public transportation data, calibrate coordinate deviations, and fuse real-time meteorological events to generate spatiotemporal grid entities;
[0007] S2: Based on spatiotemporal grid entities, a priority decision tree and rule matching logic are constructed according to four-dimensional rights and interests rules to generate scheduling vectors with ownership labels; when network resource contention is detected, a forced preemption strategy is triggered, and network resource arbitration logs are generated synchronously.
[0008] S3: Integrate scheduling vectors and network resource arbitration logs, define and execute a three-level dynamic rule chain, generate confidence arbitration results, perform path feasibility verification, and then generate a time-space constrained scheduling path;
[0009] S4: Integrate scheduling paths, confidence arbitration results, and network resource arbitration logs to generate a related dataset. Then, based on the four-dimensional rights and interests rules, formulate a sharing and distribution strategy to classify and process the related dataset.
[0010] S5: Perform a three-level quality check on the associated dataset after classification, trigger the anomaly correction mechanism and generate a data quality assessment report, and optimize in reverse through spatiotemporal analysis to form a closed-loop governance process.
[0011] Furthermore, the generation method of the spatiotemporal grid entity includes:
[0012] Detecting distance deviations between heterogeneous coordinate systems in heterogeneous public transportation data;
[0013] Obtain the road network connectivity density of road segments with distance deviations and extract the topological correlation; receive meteorological warning data, analyze the meteorological event types, and assign weight coefficients as the intensity of meteorological event impact;
[0014] By combining topological correlation with the intensity of meteorological event impact, a spatiotemporal benchmark alignment engine is constructed to calibrate distance deviation and generate a calibration coordinate set.
[0015] The calibration coordinate set is grid-encoded to obtain the road network data for each grid, and the grid-level topological correlation degree of each grid is calculated.
[0016] Integrate the grid-level topological correlation of all grids and the intensity of meteorological event impact to generate spatiotemporal grid entities.
[0017] Furthermore, the method for generating the scheduling vector includes:
[0018] Construct a priority decision tree based on the four-dimensional rights and interests rules, and define the rule matching logic;
[0019] Analyze spatiotemporal grid entities and match corresponding interest rules based on priority decision trees and rule matching logic;
[0020] Based on the rights and interests rules and topological correlation of each grid, the types and corresponding quantities of scheduling actions are dynamically generated, and each scheduling action is bound to multiple dimensions of rights and interests and encapsulated as a rights and interests signature.
[0021] Integrate scheduling actions, quantities, and ownership signatures to generate scheduling vectors with ownership labels.
[0022] Furthermore, the method of triggering a forced preemption strategy and synchronously generating a network resource arbitration log when resource contention is detected includes:
[0023] Define preemption trigger conditions and QoS level priority rules to detect and identify network resource contention events in real time;
[0024] Based on network resource contention events, predefined resource preemption strategies are dynamically matched and executed according to QoS level priority rules, and network resource arbitration logs are generated synchronously for each preemption operation.
[0025] Furthermore, the method for generating the confidence level arbitration result includes:
[0026] Define three categories of rules: integrity rules, logical rules, and compliance rules, along with verification and allocation logic, to form a three-level dynamic rule chain;
[0027] The scheduling vector and network resource arbitration log are parsed, and a three-level dynamic rule chain is executed step by step to verify the integrity of fields, logical rationality, and compliance. If the verification fails, the corresponding error code or confidence level adjustment is triggered.
[0028] The confidence level of the scheduling action is dynamically generated based on the rule verification results and integrated into a confidence arbitration result.
[0029] Furthermore, the method for performing path feasibility verification and then generating a time-space constrained scheduling path includes:
[0030] Based on the confidence level arbitration results, key constraints are extracted from heterogeneous public transportation data to form path verification logic;
[0031] A feasibility check is performed based on the path verification logic. If the feasibility check fails, a circuit breaker mechanism is triggered.
[0032] The system acquires historical traffic flow, weather conditions, and event urgency, and dynamically generates scheduling routes. If the circuit breaker mechanism is triggered, an alternative route is generated as the final scheduling route.
[0033] Furthermore, the method for formulating a sharing and distribution strategy to classify related datasets includes:
[0034] The confidence arbitration results, scheduling paths, and network resource arbitration logs are correlated and formatted to form a correlated dataset;
[0035] Based on the four-dimensional rights and interests rules, a sharing and distribution strategy is formulated to classify and process the related datasets;
[0036] Based on the classification results, tiered access permissions are provided to different stakeholders through API interfaces.
[0037] Furthermore, the data quality assessment report is generated in the following ways:
[0038] The associated dataset after classification is subjected to three levels of quality verification, including integrity verification, consistency verification, and compliance verification.
[0039] If the verification result does not meet expectations, an abnormal event is triggered, and the automatic correction mechanism is executed.
[0040] Integrate the results of the three-level quality verification, abnormal event records, and automatic correction records to generate a data quality assessment report.
[0041] Furthermore, the method of forming a closed-loop governance process through spatiotemporal analysis and reverse optimization includes:
[0042] Perform spatiotemporal analysis on data quality assessment reports to identify the root causes of data quality problems and generate actionable analysis results;
[0043] The analysis results are transformed into optimization actions, driving the optimization of the three-level dynamic rule chain and data source, forming a closed-loop governance process.
[0044] Based on the same inventive concept, a public transportation travel service big data processing system is also proposed, implemented based on the aforementioned public transportation travel service big data processing method, including:
[0045] Multi-source spatiotemporal fusion module: Integrates heterogeneous public transportation data, calibrates coordinate deviations, fuses real-time meteorological events, and generates spatiotemporal grid entities;
[0046] The four-dimensional rights and interests scheduling module: Based on the spatiotemporal grid entity, it constructs a priority decision tree and rule matching logic according to the four-dimensional rights and interests rules, and generates a scheduling vector with ownership labels; when network resource contention is detected, it triggers a forced preemption strategy and generates a network resource arbitration log simultaneously;
[0047] Rule Arbitration and Path Deduction Module: Integrates scheduling vectors and network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, performs path feasibility verification, and then generates a time-space constrained scheduling path;
[0048] Multi-type rights data encapsulation module: integrates scheduling path, confidence arbitration results and network resource arbitration logs to generate associated datasets, and then formulates a sharing and distribution strategy to classify and process the associated datasets according to the four-dimensional rights rules;
[0049] Evaluation and optimization module: Performs three-level quality checks on the classified and processed associated datasets, triggers anomaly correction mechanisms and generates data quality assessment reports, and performs reverse optimization through spatiotemporal analysis to form a closed-loop governance process.
[0050] The technical effects and advantages of this invention are as follows:
[0051] This invention generates a high-precision spatiotemporal grid entity by calibrating the coordinate deviations of heterogeneous data and fusing real-time meteorological events. This step solves the accuracy problem of traditional static coordinate transformation, and by dynamically adjusting the road network topology density and the intensity of meteorological influence, it ensures the spatiotemporal consistency of multi-source data, providing a reliable foundation for subsequent scheduling.
[0052] Secondly, a priority decision tree and rule matching logic are constructed based on the four-dimensional rights and interests rules to generate scheduling vectors with ownership labels, and a forced preemption strategy is triggered when resource contention is detected. This step ensures the priority of resource allocation for high-priority tasks (such as emergency scheduling) through a dynamic priority management mechanism, and generates network resource arbitration logs to achieve operation traceability, effectively resolving resource conflict issues.
[0053] Furthermore, by integrating scheduling vectors and arbitration logs, a three-level dynamic rule chain is executed to generate confidence-based arbitration results, and a spatiotemporally constrained scheduling path is generated through path feasibility verification. This step ensures the legality and safety of scheduling actions through step-by-step verification using three types of rules: completeness, logicality, and compliance. Combined with real-time traffic conditions and physical constraints, feasible or alternative paths are dynamically generated, significantly reducing the risk of path infeasibility.
[0054] Subsequently, the associated datasets are categorized using a sharing and distribution strategy, and hierarchical sharing of multiple data types is achieved based on four-dimensional rights and interests rules. This step provides differentiated access permissions to different rights holders through an API interface, satisfying compliance requirements such as government data notarization and enterprise data anonymization while protecting personal privacy, thus solving the compliance challenges of traditional data sharing.
[0055] Finally, a closed-loop governance process is formed through three-level quality verification and spatiotemporal analysis for reverse optimization. This step performs integrity, consistency, and compliance checks on the categorized data, triggers anomaly correction mechanisms, and generates a quality assessment report. Furthermore, spatiotemporal analysis identifies the root causes of data quality issues, drives rule chain optimization and data source improvement, and continuously enhances the system's adaptive capabilities.
[0056] Through the above steps, this invention constructs a complete process system from data fusion, resource scheduling, rule verification to shared governance, effectively solving the shortcomings of existing technologies in spatiotemporal alignment, rights scheduling, compliant sharing and path planning, and providing an efficient, safe and traceable systematic solution for public transportation big data processing. Attached Figure Description
[0057] Figure 1 A flowchart illustrating the big data processing method for this public transportation travel service.
[0058] Figure 2 This is a flowchart illustrating the rule arbitration and route deduction module in the big data processing method for public transportation travel services.
[0059] Figure 3 A flowchart illustrating the big data processing system for public transportation travel services. Detailed Implementation
[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0061] Example 1
[0062] Please see Figure 1 and Figure 2 As shown in this embodiment, a big data processing method for public transportation travel services includes:
[0063] S1: Access heterogeneous public transportation data, calibrate coordinate deviations, and fuse real-time meteorological events to generate spatiotemporal grid entities;
[0064] S2: Based on spatiotemporal grid entities, a priority decision tree and rule matching logic are constructed according to four-dimensional rights and interests rules to generate scheduling vectors with ownership labels; when network resource contention is detected, a forced preemption strategy is triggered, and network resource arbitration logs are generated synchronously.
[0065] S3: Integrate scheduling vectors and network resource arbitration logs, define and execute a three-level dynamic rule chain, generate confidence arbitration results, perform path feasibility verification, and then generate a time-space constrained scheduling path;
[0066] S4: Integrate scheduling paths, confidence arbitration results, and network resource arbitration logs to generate a related dataset. Then, based on the four-dimensional rights and interests rules, formulate a sharing and distribution strategy to classify and process the related dataset.
[0067] S5: Perform a three-level quality check on the associated dataset after classification, trigger the anomaly correction mechanism and generate a data quality assessment report, and optimize in reverse through spatiotemporal analysis to form a closed-loop governance process.
[0068] The methods for generating spatiotemporal grid entities include:
[0069] Use the MQTT / HTTP / 2 protocol to receive heterogeneous public transportation data streams, such as bus GPS data (parse the JT808 protocol header (0x7E identifier) and extract WGS84 coordinates), shared bicycle order data (parse JSON fields and extract GCJ02 coordinates), and ride-hailing trajectory data (parse the JT905 protocol and extract BD09 coordinates).
[0070] The protocol's legitimacy is verified by the protocol header identifier (such as the 0x7E start character of JT808), thereby filtering invalid data packets; JT808 is used for public transportation GPS data (0x7E start character), and JT905 is used for ride-hailing trajectory data (0x7F start character).
[0071] Detect the distance deviation between the heterogeneous coordinate systems of various public transportation devices in heterogeneous public transportation data;
[0072] Specifically, the Haversine algorithm is used to calculate the distance deviation between heterogeneous coordinate systems (such as WGS84 vs GCJ02 vs BD09); for example, if the GCJ02 coordinates of a shared bicycle order deviate from the WGS84 coordinates of the bus GPS data by 180 meters, it is marked as needing calibration.
[0073] It should be noted that the coordinate deviation here refers to the deviation of the same location in different coordinate systems. If two data sources (such as shared bicycle orders and public transport GPS) record the same physical location (e.g., an intersection), but use different coordinate systems (e.g., GCJ02 vs WGS84), their coordinate values will be different. This deviation is a mathematical difference in coordinate system transformation, rather than a change in the actual geographical location.
[0074] Example: The actual location of an intersection is (31.2304, 121.4737) in the GCJ02 coordinate system, but may be (31.2315, 121.4748) in the WGS84 coordinate system. There is a distance discrepancy between the two. The purpose of calibration is to eliminate this difference.
[0075] For heterogeneous coordinate systems that need calibration, obtain the road network connection density of the corresponding road segment, extract the topological correlation degree (the method of obtaining the topological correlation degree is consistent with the calculation method of the subsequent grid-level topological correlation degree), receive real-time meteorological warning data, analyze the meteorological event type, and assign weight coefficients as the influence intensity of the meteorological event.
[0076] For example, the types of meteorological events and the intensity of their impact (weighting coefficients): Red rainstorm warning (1.5), large-scale events (2.0), traffic accidents (0.8);
[0077] By combining topological correlation with the impact intensity of meteorological events, a spatiotemporal benchmark alignment engine is constructed to calibrate distance deviations. All coordinates (including coordinates that do not need calibration and calibrated coordinates) are integrated to generate a calibration coordinate set (including protocol conflict marker fields).
[0078] Specifically, the spatiotemporal reference alignment engine takes the ratio of topological correlation degree to the intensity of meteorological event impact as a weight value, and then uses this weight value to calibrate the corresponding heterogeneous coordinates to obtain the calibration coordinates;
[0079] Calibration coordinates = weight value × heterogeneous coordinate 1 + (1 - weight value) × heterogeneous coordinate 2;
[0080] Assuming the distance discrepancy exists between the WGS84 coordinates and the GCJ02 coordinates, then the calibration coordinates = weight value × WGS84 + (1 - weight value) × GCJ02;
[0081] The calibration coordinate set is grid-coded (generating Geohash 7-level codes (accuracy 100 meters), example: coordinates (31.2304°N, 121.4737°E) → Geohash 7-level code wx4er), the road network connection density is extracted, and the grid-level topological correlation degree is obtained;
[0082] The calculation method based on road network connectivity density topological correlation is as follows:
[0083] Basic indicators are extracted from the road network data of each grid, including the number of intersections (the total number of intersections within the grid area, reflecting the node density of the road network), the total road length (the total length of roads within the grid area, reflecting the coverage of the road network), the road width (the number of lanes or average width of each road within the grid (e.g., the width of a main road is 4 lanes and a branch road is 2 lanes)) and the area (geographical range (i.e., the area of the grid area)).
[0084] Then, the width of each road (such as the number of lanes) is converted into a width weight to calculate the average width within the grid. That is, the corresponding width weight is weighted and fused with the corresponding road length, and the weighted fusion result is then calculated as a ratio to the total road length to obtain the average width of the roads within the grid.
[0085] The number of intersections, total road length, and average width are standardized to eliminate dimensional differences. Then, each is assigned a weight (the assigned weight needs to be adjusted according to actual needs, such as through expert scoring or regression analysis; usually, the sum of the three weights is 1), and a weighted fusion is performed to obtain the topological correlation degree.
[0086] The grid-level topological correlation and meteorological event impact intensity of all grids are integrated to generate grid basic data containing meteorological weights and topological correlation. Then, the SHA256 algorithm is used to hash each grid ID and timestamp to generate a check code to ensure data integrity.
[0087] Integrate wx4er grid ID, meteorological event impact intensity, topological correlation degree, and check code to generate spatiotemporal grid entities.
[0088] The methods for generating scheduling vectors include:
[0089] Construct a priority decision tree based on the four-dimensional rights and interests rule (government > enterprise > individual > operator):
[0090] [Government Needs] →|Emergency Dispatch| (Highest Priority); [Enterprise Needs] →|Business Analysis| (Medium Priority); [Personal Needs] →|Trajectory Query| (Low Priority); [Operational Revenue Needs] →|Resource Optimization| (Lowest Priority);
[0091] The rule matching logic is defined as follows: when there is a priority conflict, the higher priority right automatically overrides the lower priority right. If there are multiple right requests for the same priority, they are handled according to the first-come, first-served principle.
[0092] Parse the fields in the spatiotemporal grid entity (wx4er grid ID: target scheduling area, event weight: impact of meteorological events (e.g., rainstorm weight 1.5), topological correlation: road network connection strength (e.g., 0.87)), and match the corresponding rights and interests rules according to the priority decision tree and rule matching logic;
[0093] For example, if the grid ID is wx4er and the impact intensity of the meteorological event is 1.5 (red alert for heavy rain), then the government's emergency dispatch priority will be triggered.
[0094] Based on the rights and interests rules and topological correlation of each grid, the type and corresponding quantity of scheduling actions are dynamically generated;
[0095] Specifically, the intensity of the impact of meteorological events drives the matching of dispatch action types, the number of corresponding actions is quantified through topological correlation, and then combined with the matching rights and interests rules and flow restriction mechanism, the dynamic generation of dispatch instructions and the optimized allocation of resources are realized, thereby ensuring that high-priority needs (such as government emergencies) are executed first, while taking into account the reasonable demands of enterprises, individuals and operators.
[0096] Based on the priority decision tree and grid status of the four-dimensional rights and interests rules, the scheduling action type is dynamically matched;
[0097] For example, action type classification and matching logic:
[0098] Scheduling action type, triggering condition, and priority:
[0099] Supply → Government emergency needs (such as rainstorm warnings) → Government (highest priority);
[0100] Analysis → Enterprise business analysis needs (such as road network traffic forecasting) → Enterprise (medium priority);
[0101] Query → Personal Tracking Query Request → Individual (Low Priority);
[0102] Optimization → Resource optimization needs of operators → Operators (lowest priority);
[0103] For example, if the grid ID is wx4er, the intensity of the meteorological event is 1.5 (red rainstorm warning), and the priority is the highest priority government emergency dispatch, then the supply action will be triggered (government priority).
[0104] If the topological correlation degree is 0.87 (dense road network), the enterprise may initiate analysis actions (such as analyzing traffic flow).
[0105] The number of corresponding dispatch actions is determined by the intensity of the meteorological event's impact and the degree of topological correlation, and is adjusted in conjunction with rights and interests rules and flow restriction mechanisms;
[0106] The calculation logic is as follows:
[0107] Government Supply Amount: Supply Amount = Event Weight × Basic Supply Coefficient;
[0108] Enterprise Analysis Quantity: Analysis Quantity = Topological Relationship Degree × Basic Analysis Coefficient;
[0109] Operational optimization amount: Optimization amount = Road network density × Basic optimization coefficient;
[0110] The adjustment logic of the flow restriction mechanism is as follows: if the calculation result exceeds the system resource limit (e.g., the government supply of 150 exceeds the inventory of 50), the flow restriction will be adjusted, including: adjusting to the inventory limit, dynamically correcting the quantity according to resource availability (e.g., adjusting to 50), preserving the priority of the rights and interests rules and forcibly covering low priority actions.
[0111] For example, the calculation logic for rights type, action type, and quantity:
[0112] Government → Supply → Supply Amount for Rainstorm Areas = Event Weight × 100;
[0113] Enterprise → Analytics → Data Analysis Request = Topological Relevance × 50;
[0114] Individual → Query → Track Query Count = 1 time / user;
[0115] Operator → Optimization → Resource Optimization Suggestion = Road Network Density × 20;
[0116] Furthermore, each scheduling action is bound to multiple dimensions of ownership and encapsulated into an ownership signature, that is, ownership tags are attached to each scheduling action;
[0117] For example, governments: attach blockchain licenses; enterprises: attach differential privacy identifiers; individuals: attach user identity hashes; operators: attach API key hashes;
[0118] Integrate scheduling actions and ownership signatures to generate scheduling vectors with ownership tags, including grid ID, action type, quantity, and a list of ownership signatures with multiple rights binding.
[0119] When resource contention is detected, the forced preemption strategy is triggered, and network resource arbitration logs are generated synchronously in the following ways:
[0120] Define preemption trigger conditions and QoS level priority rules to detect and identify network resource contention events in real time;
[0121] Specifically, it involves real-time collection of 5G channel metrics, such as bandwidth utilization, QoS (Quality of Service) level, and channel latency.
[0122] The preemption trigger condition is defined as follows: if the bandwidth utilization rate does not meet the preset threshold (e.g., <10Mbps), the resource preemption policy is triggered; if a high-priority scheduling command (e.g., government or enterprise) cannot be executed due to insufficient resources, the resource preemption policy is forcibly activated.
[0123] The QoS level priority rule is defined as follows: network resource allocation priority is distinguished according to QoS level, such as Level 0 (government emergency, absolute priority, preempting other network resources), Level 1 (enterprise needs, medium priority), and Level 2 (commercial advertising push, low priority, can be preempted).
[0124] Furthermore, based on real-time network resource status and QoS level priority rules, predefined resource preemption strategies are dynamically matched and executed, including interrupting low-priority network resources and allocating high-priority network resources.
[0125] Specifically, this involves releasing the bandwidth occupied by Level 2 (commercial advertising) (e.g., interrupting the streaming of advertising media) and dynamically allocating the released bandwidth to Level 0 (government emergency dispatch instructions).
[0126] Establish network resource isolation safeguards and use 5G slicing technology to isolate Level 0 network resources to ensure their exclusivity;
[0127] After preemption, the system continuously monitors the duration of Level 0 resource usage. If the Level 0 task is completed, the network resources are automatically released and Level 2 services are restored. If preemption fails (e.g., due to insufficient resources), a circuit breaker mechanism is triggered (e.g., service degradation).
[0128] Synchronously generate a network resource arbitration log for each preemption operation, including timestamp, event type (e.g., resource arbitration), action type (e.g., preempting Level 2 channel), target (e.g., guarantee scheduling instruction), resource status change (including duration, before resource status (e.g., "Level 2 network resource utilization": "85%", "bandwidth utilization": "9 Mbps"), after resource status (e.g., "Level 0 network resource utilization": "100%", "Level 2 utilization": "0%"), and ownership signature;
[0129] Then, the network resource arbitration log is hashed using the SHA256 hash algorithm and stored as a blockchain distributed ledger to ensure that the data content is immutable.
[0130] The methods for generating confidence level arbitration results include:
[0131] Define the verification and allocation logic for three categories of rules: integrity, logic, and compliance, forming a three-level dynamic rule chain;
[0132] The three-level dynamic rule chain includes:
[0133] Integrity rules: Verify that the fields in the scheduling vector and network resource arbitration log are complete (such as GPS trajectory, car rental time, number of empty cars, usage scope, etc.).
[0134] Logical rules: Verify the rationality of scheduling actions (such as the matching of car borrowing time with the number of available cars, and the spatiotemporal continuity of resource allocation).
[0135] Compliance rules: Ensure that scheduling actions comply with laws and regulations (such as the restrictions on the scope of data use under the Data Security Law);
[0136] The scheduling vector and network resource arbitration logs are parsed, and a three-level dynamic rule chain is executed step by step through the federated engine to verify the integrity of fields, logical rationality and compliance in turn; if the verification fails, the corresponding error code or confidence adjustment is triggered;
[0137] The logic of the three-level dynamic rule chain's step-by-step verification can be described by the following example:
[0138] Integrity rules: Field: "GPS track", validation condition: "not empty and continuous"; Field: "car rental time", validation condition: "format YYYY-MM-DD HH:MM:SS";
[0139] Integrity verification: Input GPS track, detect missing points → mark error code 1001 → generate anomaly description: GPS track is incomplete;
[0140] When the car rental time is entered, the system detects that the format does not conform to the specifications → error code 1002 is generated → time format error;
[0141] Logical rules: Condition: "Rental time > number of empty vehicles × 2 hours", scheduling action: "Trigger logical conflict"; Condition: "Network resource arbitration log timestamp is inconsistent with scheduling time", scheduling action: "Mark an exception";
[0142] Logical rationality verification: Input the car borrowing time and the number of empty cars (the number of empty cars after borrowing a car). If the detection finds that the car borrowing time is greater than the time corresponding to the number of empty cars, error code 2003 is generated, indicating a contradiction between the car borrowing time and the number of empty cars.
[0143] Input the timestamp of the network resource arbitration log. If it is found to be outside the scheduling time window, error code 2004 is generated: timestamp mismatch.
[0144] Compliance rules: Clause: "Data Security Law", verification condition: "Data usage scope does not exceed the authorized geographical area"; Clause: "Cybersecurity Law", verification condition: "No sensitive data sharing is involved";
[0145] Compliance Verification: Detection revealed that the scope of use exceeded the authorized geographical area under the Data Security Law → Error code 3007 was generated → Compliance violation:
[0146] The test found that the input data involved sensitive data and was not anonymized → error code 3008 was generated → data security violation;
[0147] The confidence level of the scheduling action is dynamically generated based on the rule verification result, and then the confidence arbitration result is generated, including the error code, confidence level and verification result (including arbitration time and exception description).
[0148] Specifically, the confidence level of the scheduling action is generated as follows: each rule verification is assigned a weight (the weight is determined by expert scoring, and the sum is 1 (i.e., 100%), such as completeness accounting for 30%, logic accounting for 40%, and compliance accounting for 30%). After the rule verification, the score is dynamically adjusted according to the verification results.
[0149] The logic of dynamically adjusting the score is defined as follows: when all rules pass, the total confidence score is 1; if a rule fails, the weight ratio of that rule is subtracted from the total score.
[0150] For example, suppose the integrity check weight is 30%: if an integrity rule fails, 30% of the score is deducted from the total score;
[0151] If multiple rules fail, the weights will be deducted cumulatively, but the total score will not be less than 0.0;
[0152] If the confidence score is lower than the preset confidence threshold (e.g., 0.6), the circuit breaker is triggered, which means that the scheduling action is prohibited from being executed and the log is recorded.
[0153] Methods for performing path feasibility verification and then generating time- and space-constrained scheduling paths include:
[0154] Based on the confidence level arbitration results, key constraints (including road height restrictions and road right-of-way, which can also be adjusted according to actual needs) are extracted from the heterogeneous public transportation data to form route verification logic, and then a feasibility verification is performed. If the feasibility verification fails, a circuit breaker mechanism is triggered (i.e., route generation is prohibited).
[0155] Among them, the path verification logic: taking the height limit detection verification logic as an example, the height limit detection mainly detects the height limit node and the vehicle height;
[0156] Height restriction nodes are road height restrictions that exist in the road (such as overpasses).
[0157] Therefore, the height restriction detection and verification logic is as follows: if the path contains a height restriction node and the vehicle height exceeds the height limit of the height restriction node, the circuit breaker mechanism is triggered to prevent path generation.
[0158] Other constraint checks can be added as needed, such as road right-of-way checks (to verify whether dispatched vehicles have the right to pass on specific roads (such as emergency lanes, traffic restriction policies)) and spatiotemporal continuity checks (to ensure that height restriction nodes meet the continuity requirements in time and space (such as the distance between adjacent height restriction nodes does not exceed the maximum driving speed limit)).
[0159] The system acquires historical traffic flow, weather conditions, and event urgency, and dynamically generates scheduling paths under spatiotemporal constraints. If a circuit breaker is triggered, an alternative path is generated as the final scheduling path.
[0160] Specifically, it involves acquiring historical traffic flow data (such as traffic flow data for each road segment in the past hour), weather conditions (such as rainfall and visibility), and event urgency (such as urgency scores for traffic accidents and road closures). These data are then used as input to predict future traffic flow and congestion levels using machine learning methods (such as Long Short-Term Memory (LSTM) models). For example, it can predict the probability of congestion for each road segment in the next 30 minutes (probability range: 0.0-1.0).
[0161] Then, based on the predicted road congestion, the optimization objective is defined as minimizing the total travel time, and the optimal scheduling path is generated through a route planning algorithm.
[0162] An example of the optimal scheduling path:
[0163] Path nodes: ["Start point 1", "Stop point 1", "Destination 1"], Estimated travel time: "45 minutes", Congestion probability: 0.2;
[0164] Among them, the path nodes are vehicle stops (such as bus stops and shared bicycle parking spots).
[0165] If a path triggers a circuit breaker due to height restrictions or other constraints, an alternative path algorithm (such as Dijkstra's shortest path algorithm) is automatically invoked to generate alternative paths. After generating alternative paths, real-time congestion probability optimization is applied to obtain a new optimal scheduling path.
[0166] Example circuit breaker result: Error code 4001, reason: "Height limit detection failed (overpass X height limit)", alternative path: ["Start point 1", "Detour 2", "End point 1"];
[0167] The final scheduling path and path attributes (such as travel time and congestion probability) are integrated as the basis for subsequent execution;
[0168] Methods for classifying related datasets using shared distribution strategies include:
[0169] The confidence level arbitration results, scheduling paths, and network resource arbitration logs are correlated and formatted uniformly. Specifically, the correlation is performed as follows:
[0170] The confidence arbitration result is associated with the scheduling path by scheduling ID to form a "path + confidence" combination; a unique ID is added to each scheduling path as the scheduling ID.
[0171] Associate the blockchain-based notarized hash of the network resource arbitration log with the path nodes of the scheduling path to ensure traceability of path resource allocation and supplement resource usage dynamics (such as "Level 2 channel at docking point 1 has been preempted").
[0172] The unified format is as follows: convert all data to JSON / XML format, and standardize field naming.
[0173] The confidence arbitration results, scheduling paths, and network resource arbitration logs, after being integrated, correlated, and formatted in a unified manner, are formed into a correlated dataset.
[0174] Based on the four-dimensional rights and interests rules (government, enterprise, individual, operator), a sharing and distribution strategy is formulated, and the related datasets are classified and processed.
[0175] The sharing and distribution strategy is as follows: government data (such as scheduling paths and resource arbitration logs) are stored on the blockchain using blockchain notarization technology to generate ownership audit packages and ensure data traceability.
[0176] Differential privacy technology is used to anonymize enterprise data (such as de-identified riding trajectories and shared bicycle route and stop / replacement data) and generate an anonymized sharing package to protect the privacy of commercial data.
[0177] Personal data (such as user identity and behavioral patterns) is anonymized and then distributed with user authorization.
[0178] Operator data (such as IoT device status and parking pillar vacancy rate) requires negotiation with equipment manufacturers regarding permissions to ensure operational compliance, legality of data use, and closed-loop management.
[0179] Furthermore, by providing tiered access permissions to different stakeholders through API interfaces, compliant data sharing and dynamic governance can be achieved.
[0180] Different API interfaces correspond to different rights holders:
[0181] Government data is provided through a dedicated government API interface, supporting audit traceability and real-time query.
[0182] Enterprise data provides data subscription services through internal enterprise API interfaces, supporting on-demand retrieval of specific fields (such as confidence scores and resource arbitration events).
[0183] Individuals and operators need to go through an approval process to obtain limited access rights;
[0184] It should be noted that the logic behind the sharing and distribution strategy is to classify and process data according to the different natures of the four-dimensional rights (including government data notarization, enterprise data anonymization, personal data anonymization, and operator data authorization).
[0185] Data quality assessment reports can be generated in the following ways:
[0186] The associated dataset after classification is subjected to three levels of quality verification, including integrity verification, consistency verification, and compliance verification.
[0187] If the verification result does not meet expectations, an abnormal event is triggered, and the automatic correction mechanism is executed.
[0188] Integrity verification includes field-level verification and data-level verification, with the aim of ensuring that there are no missing or abnormal data during transmission and storage.
[0189] Field-level validation involves checking whether there are any missing fields in the associated dataset after classification (e.g., the scheduling path is missing the "stop point 1" path node). If any fields are missing, they are marked as abnormal events.
[0190] Data-level verification involves comparing the data volume of the four types of data divided by the sharing and distribution strategy. If the difference rate exceeds a preset difference threshold (e.g., 5%), it is marked as an abnormal event.
[0191] The automatic correction mechanism involves automatically performing data completion tasks using ETL tools and recording completion logs.
[0192] For data that cannot be automatically completed, a manual completion process is triggered to notify the operations and maintenance personnel for handling;
[0193] Consistency checks include conflict event checks and timestamp checks, with the aim of ensuring the consistency of the four types of data and preventing data tampering or logical conflicts.
[0194] The conflict event verification involves comparing the consistency of fields in four types of data (such as path node sequences and resource arbitration events).
[0195] Example: If the ownership audit package records "starting point 1 → docking point 2 → ending point 3", but the de-identified shared package displays "starting point 1 → ending point 3", it is determined to be a conflict and marked as an abnormal event;
[0196] The timestamp verification is as follows: it verifies whether the timestamps of the four types of data match. If the difference exceeds the preset time difference threshold (e.g., 10 seconds), it is marked as an abnormal event.
[0197] The automatic correction mechanism is as follows: if a conflict exists, the original data is extracted from the data type with more complete data to cover the missing parts in other data types; if there is an abnormal timestamp, an alarm is triggered.
[0198] For conflict events that cannot be automatically resolved, a manual review process is triggered, and the data governance team is notified to handle them.
[0199] Compliance verification includes privacy compliance checks and legal compliance scans to ensure that shared data complies with laws and regulations (such as the Data Security Law) and privacy protection requirements;
[0200] Privacy compliance check: Verify whether sensitive fields have been correctly de-identified (e.g., user identity fields are irreversibly encrypted). If they have not been correctly de-identified, mark them as abnormal events.
[0201] Legal compliance scan: Automatically scans highly sensitive fields (such as personal identification and device unique ID) in shared data. If unsensitized fields are found, an alarm is triggered.
[0202] The automatic correction mechanism automatically adjusts the differential privacy parameters in the differential privacy technology (such as increasing the ε value to 0.05) and regenerates the de-identified shared package if any un-de-identified fields exist.
[0203] For incidents involving significant compliance risks (such as leaks of user privacy), a manual approval process is triggered, and the legal team intervenes to handle the matter.
[0204] Integrate the results of the three-level quality verification, abnormal event records, and automatic correction records to generate a data quality assessment report.
[0205] Methods for forming a closed-loop governance process through spatiotemporal analysis and reverse optimization include:
[0206] Perform spatiotemporal analysis on data quality assessment reports to identify the root causes of data quality problems and generate actionable analysis results;
[0207] Specifically, the spatiotemporal dimension analysis involves using spatiotemporal clustering algorithms (such as DBSCAN) to statistically analyze areas (such as core urban areas and transportation hubs) where scheduling routes are concentrated during peak traffic hours (such as 7:00-9:00), and these areas are recorded as high-frequency routes.
[0208] Simultaneously, by using spatiotemporal trend analysis algorithms (such as time series analysis), abnormal spatiotemporal points in the scheduling path are identified (such as a sudden increase in traffic flow on a certain road section at night, or a path breakage caused by missing intermediate nodes in a certain area), and these are recorded as abnormal areas. These abnormal spatiotemporal points can be identified as data quality issues.
[0209] This generates a heatmap of the spatiotemporal distribution of scheduling paths, which is marked with high-frequency paths and abnormal regions.
[0210] Synchronously generate a list of abnormal areas, including specific locations, time ranges, and problem types (such as missing fields or inconsistent timestamps).
[0211] Simultaneously, it integrates the spatiotemporal distribution heatmap of scheduling paths, the list of abnormal areas, and the statistics of abnormal data fields to generate a spatiotemporal distribution analysis report, which is submitted to operations personnel for data management and data analysis.
[0212] The analysis results are transformed into optimization actions, driving the optimization of the three-level dynamic rule chain and data source, forming a closed-loop governance process.
[0213] Specifically, it involves transforming the analysis results into specific tasks of three-level dynamic rule chain optimization and data source optimization;
[0214] Three-level dynamic rule chain optimization: The field missing rate of abnormal areas (such as a path node missing rate of 15%) is fed back to the process of three-level dynamic rule chain for step-by-step verification and generation of scheduling action confidence, thereby increasing the weight of integrity rules.
[0215] That is, simultaneously reduce the weight of compliance rules and integrity rules, and then add the reduced weight to the weight of integrity rules;
[0216] Data source optimization: Feedback abnormal areas in spatiotemporal analysis (such as a sudden increase in traffic flow on a certain road section at night) to the heterogeneous public transportation data receiving and acquisition module;
[0217] Adjust data acquisition strategies (such as increasing the frequency of sensor acquisition at night and fixing interface errors, thereby enhancing the reception of heterogeneous public transportation data).
[0218] Example 2
[0219] Please see Figure 3 As shown, parts not described in detail in this embodiment are described in Embodiment 1. A public transportation travel service big data processing system is provided, including:
[0220] Multi-source spatiotemporal fusion module: Integrates heterogeneous public transportation data, constructs a spatiotemporal benchmark alignment engine to calibrate coordinate deviations, fuses real-time meteorological events, and generates spatiotemporal grid entities;
[0221] The four-dimensional rights and interests scheduling module: Based on the spatiotemporal grid entity, it constructs a priority decision tree and rule matching logic according to the four-dimensional rights and interests rules, and generates a scheduling vector with ownership labels; when network resource contention is detected, it triggers a forced preemption strategy and generates a network resource arbitration log simultaneously;
[0222] Rule Arbitration and Path Deduction Module: Integrates scheduling vectors and network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, performs path feasibility verification, and then generates a time-space constrained scheduling path;
[0223] Multi-type rights data encapsulation module: integrates scheduling path, confidence arbitration results and network resource arbitration logs to generate associated datasets, and then formulates a sharing and distribution strategy to classify and process the associated datasets according to the four-dimensional rights rules;
[0224] Evaluation and optimization module: Performs three-level quality checks on the classified and processed associated datasets, triggers anomaly correction mechanisms and generates data quality assessment reports, and performs reverse optimization through spatiotemporal analysis to form a closed-loop governance process.
[0225] Example 3
[0226] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the operation mode of the public transportation travel service big data processing system described above.
[0227] Since the electronic device described in this embodiment is an electronic device used to implement the public transportation travel service big data processing method described in this application embodiment, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the public transportation travel service big data processing method described in this application embodiment. Therefore, how the electronic device implements the method in this application embodiment will not be described in detail here. Any electronic device used by those skilled in the art to implement the public transportation travel service big data processing method in this application embodiment falls within the scope of protection of this application.
[0228] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0229] The above description is merely a preferred embodiment of the present invention, and the scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for users of ordinary technical skills, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for processing big data related to public transportation travel services, characterized in that, include: S1: Access heterogeneous public transportation data, calibrate coordinate deviations, and fuse real-time meteorological events to generate spatiotemporal grid entities; The methods for generating the spatiotemporal grid entities include: Detecting distance deviations between heterogeneous coordinate systems in heterogeneous public transportation data; Obtain the road network connectivity density of road segments with distance deviations and extract the topological correlation; receive meteorological warning data, analyze the meteorological event types, and assign weight coefficients as the intensity of meteorological event impact; By combining topological correlation with the intensity of meteorological event impact, a spatiotemporal benchmark alignment engine is constructed to calibrate distance deviation and generate a calibration coordinate set. The calibration coordinate set is grid-encoded to obtain the road network data for each grid, and the grid-level topological correlation degree of each grid is calculated. Integrate the grid-level topological correlations and meteorological event impact intensities of all grids to generate spatiotemporal grid entities; S2: Based on spatiotemporal grid entities, a priority decision tree and rule matching logic are constructed according to four-dimensional rights and interests rules to generate scheduling vectors with ownership labels; when network resource contention is detected, a forced preemption strategy is triggered, and network resource arbitration logs are generated synchronously. S3: Integrate scheduling vectors and network resource arbitration logs, define and execute a three-level dynamic rule chain, generate confidence arbitration results, perform path feasibility verification, and then generate a time-space constrained scheduling path; S4: Integrate scheduling paths, confidence arbitration results, and network resource arbitration logs to generate a related dataset. Then, based on the four-dimensional rights and interests rules, formulate a sharing and distribution strategy to classify and process the related dataset. S5: Perform a three-level quality check on the associated dataset after classification, trigger the anomaly correction mechanism and generate a data quality assessment report, and optimize in reverse through spatiotemporal analysis to form a closed-loop governance process.
2. The method for big data processing of public transportation travel services according to claim 1, characterized in that, The scheduling vector is generated in the following ways: Construct a priority decision tree based on the four-dimensional rights and interests rules, and define the rule matching logic; Analyze spatiotemporal grid entities and match corresponding interest rules based on priority decision trees and rule matching logic; Based on the rights and interests rules and topological correlation of each grid, the types and corresponding quantities of scheduling actions are dynamically generated, and each scheduling action is bound to multiple dimensions of rights and interests and encapsulated as a rights and interests signature. Integrate scheduling actions, quantities, and ownership signatures to generate scheduling vectors with ownership labels.
3. The method for big data processing of public transportation travel services according to claim 2, characterized in that, The methods for triggering a forced preemption strategy and synchronously generating a network resource arbitration log when resource contention is detected include: Define preemption trigger conditions and QoS level priority rules to detect and identify network resource contention events in real time; Based on network resource contention events, predefined resource preemption strategies are dynamically matched and executed according to QoS level priority rules, and network resource arbitration logs are generated synchronously for each preemption operation.
4. The method for big data processing of public transportation travel services according to claim 3, characterized in that, The methods for generating the confidence level arbitration result include: Define three categories of rules: integrity rules, logical rules, and compliance rules, along with verification and allocation logic, to form a three-level dynamic rule chain; The scheduling vector and network resource arbitration log are parsed, and a three-level dynamic rule chain is executed step by step to verify the integrity of fields, logical rationality, and compliance. If the verification fails, the corresponding error code or confidence level adjustment is triggered. The confidence level of the scheduling action is dynamically generated based on the rule verification results and integrated into a confidence arbitration result.
5. The method for big data processing of public transportation travel services according to claim 4, characterized in that, The methods for performing path feasibility verification and then generating time-space constrained scheduling paths include: Based on the confidence level arbitration results, key constraints are extracted from heterogeneous public transportation data to form path verification logic; A feasibility check is performed based on the path verification logic. If the feasibility check fails, a circuit breaker mechanism is triggered. The system acquires historical traffic flow, weather conditions, and event urgency, and dynamically generates scheduling routes. If the circuit breaker mechanism is triggered, an alternative route is generated as the final scheduling route.
6. The method for big data processing of public transportation travel services according to claim 5, characterized in that, The methods for formulating a shared distribution strategy to classify related datasets include: The confidence arbitration results, scheduling paths, and network resource arbitration logs are correlated and formatted to form a correlated dataset; Based on the four-dimensional rights and interests rules, a sharing and distribution strategy is formulated to classify and process the related datasets; Based on the classification results, tiered access permissions are provided to different stakeholders through API interfaces.
7. The method for big data processing of public transportation travel services according to claim 6, characterized in that, The data quality assessment report is generated in the following ways: The associated dataset after classification is subjected to three levels of quality verification, including integrity verification, consistency verification, and compliance verification. If the verification result does not meet expectations, an abnormal event is triggered, and the automatic correction mechanism is executed. Integrate the results of the three-level quality verification, abnormal event records, and automatic correction records to generate a data quality assessment report.
8. The method for big data processing of public transportation travel services according to claim 7, characterized in that, The method of forming a closed-loop governance process through spatiotemporal analysis and reverse optimization includes: Perform spatiotemporal analysis on data quality assessment reports to identify the root causes of data quality problems and generate actionable analysis results; The analysis results are transformed into optimization actions, driving the optimization of the three-level dynamic rule chain and data source, forming a closed-loop governance process.
9. A public transportation travel service big data processing system, implemented based on the public transportation travel service big data processing method according to any one of claims 1 to 8, characterized in that, include: Multi-source spatiotemporal fusion module: Integrates heterogeneous public transportation data, calibrates coordinate deviations, fuses real-time meteorological events, and generates spatiotemporal grid entities; The four-dimensional rights and interests scheduling module: Based on the spatiotemporal grid entity, it constructs a priority decision tree and rule matching logic according to the four-dimensional rights and interests rules, and generates a scheduling vector with ownership labels; when network resource contention is detected, it triggers a forced preemption strategy and generates a network resource arbitration log simultaneously; Rule Arbitration and Path Deduction Module: Integrates scheduling vectors and network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, performs path feasibility verification, and then generates a time-space constrained scheduling path; Multi-type rights data encapsulation module: integrates scheduling path, confidence arbitration results and network resource arbitration logs to generate associated datasets, and then formulates a sharing and distribution strategy to classify and process the associated datasets according to the four-dimensional rights rules; Evaluation and optimization module: Performs three-level quality checks on the classified and processed associated datasets, triggers anomaly correction mechanisms and generates data quality assessment reports, and performs reverse optimization through spatiotemporal analysis to form a closed-loop governance process.
Citation Information
Patent Citations
Heterogeneous data processing optimization system based on intelligent edge computing
CN120336021A