Public transportation travel service big data processing method and system

By generating space-time grid entities and four-protecting rights and interests rules, the problems of data silos and resource competition in the public transportation system are solved, and compliance with real-time scheduling and data sharing is achieved, and the system's adaptability and rationality of path planning are improved.

CN120579797AActive Publication Date: 2025-09-02QUANZHOU BIG DATA OPERATION SERVICE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511082273.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-09-02
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

The existing public transportation system has problems such as inefficiency in data silos, equity management and dynamic decision-making, delay in resource competition and unreasonable path planning, making it difficult to achieve real-time scheduling and continuous optimization.

Method used

By calibrating the coordinate deviation of heterogeneous data, fusing real-time meteorological events to generate spatiotemporal grid entities, building scheduling vectors based on the four-dimensional rights and interests rules and triggering forced preemption strategies, implementing three-level dynamic rule chains to generate confidence arbitration results, conducting path feasibility checks and data quality evaluations, and forming a closed-loop governance process.

Benefits of technology

It realizes the spatial and temporal consistency of multi-source data, ensures resource allocation for high-priority tasks, reduces the risk of path infeasibility, ensures data sharing compliance and system adaptability, and provides an efficient and safe public transportation big data processing solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579797A_ABST
    Figure CN120579797A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of public transport travel services, and discloses a public transport travel service big data processing method and system, and the method comprises the steps: accessing heterogeneous public transport data, calibrating coordinate deviation, fusing a real-time meteorological event, and generating a space-time grid entity; constructing a priority decision tree and rule matching logic according to the four-dimensional right rule, generating a scheduling vector with an ownership label, and synchronously generating a network resource arbitration log; fusing the scheduling vector and a network resource arbitration log, defining and executing a three-level dynamic rule chain, generating a confidence arbitration result, performing path feasibility verification, and further generating a space-time constrained scheduling path; according to a four-dimensional right rule, a sharing distribution strategy is formulated to classify the associated data set; and performing three-level quality verification on the associated data set after classification processing, triggering an exception correction mechanism, generating a data quality evaluation report, and performing space-time analysis reverse optimization to form a closed-loop treatment process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of public transportation travel services, and more specifically, to a method and system for processing public transportation travel service big data. Background Art

[0002] The current development of smart transportation faces the dilemma of integrating multi-source heterogeneous data. Data such as bus GPS, shared bicycle orders, and online car-hailing trajectories generated in the public transportation field form serious data silos due to protocol differences and inconsistent coordinate systems. Traditional static processing solutions are difficult to cope with the real-time scheduling needs of emergencies. In addition, data ownership conflicts between stakeholders such as enterprises, individuals, and operators lead to low cross-subject collaboration efficiency, significant command transmission delays during resource competition, and insufficient dynamic decision-making robustness. There is an urgent need to build a new technical system that takes into account real-time performance, ownership compliance, and closed-loop self-optimization.

[0003] Existing solutions rely on manual calibration of protocol deviations to solve the data island problem, which is inefficient and unable to dynamically respond to environmental changes; ownership management adopts a fixed priority strategy and lacks a flexible arbitration mechanism during resource competition, resulting in the obstruction of high-priority instruction execution; path planning does not combine real-time confidence arbitration and circuit breaker control, and the proportion of failed instructions remains high; historical problem optimization relies on manual rule adjustment and lacks parameter self-iteration capabilities, making it difficult to continuously improve system adaptability. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solutions: A method and system for processing public transportation travel service big data, comprising: S1: Access heterogeneous public transportation data, calibrate coordinate deviations, integrate real-time meteorological events, and generate spatiotemporal grid entities; S2: Based on the spatiotemporal grid entity, a priority decision tree and rule matching logic are constructed according to the four-dimensional rights and interests rules to generate a scheduling vector with ownership tags. When network resource competition is detected, a forced preemption strategy is triggered and a network resource arbitration log is generated simultaneously. S3: Integrates scheduling vectors with network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, and performs path feasibility verification to generate scheduling paths with time and space constraints. S4: Fusion of scheduling paths, confidence arbitration results, and network resource arbitration logs to generate a related dataset. Based on the four-dimensional equity rules, a sharing and distribution strategy is formulated to classify and process the related dataset. S5: Perform three-level quality verification on the associated data sets after classification processing, trigger the anomaly correction mechanism and generate a data quality assessment report, and form a closed-loop governance process through reverse optimization through spatiotemporal analysis.

[0005] Furthermore, the generation method of the space-time grid entity includes: Detecting distance deviations in heterogeneous coordinate systems in heterogeneous public transportation data; Obtain the road network connection density of sections with distance deviation and extract the topological correlation; receive meteorological warning data, analyze the type of meteorological event, and assign a weight coefficient as the impact intensity of the meteorological event; Combining topological correlation with the impact intensity of meteorological events, a spatiotemporal benchmark alignment engine is constructed to calibrate distance deviations and generate a calibration coordinate set. Grid encode the calibration coordinate set, obtain the road network data of each grid, and calculate the grid-level topological correlation of each grid; The grid-level topological correlation of all grids and the impact intensity of meteorological events are integrated to generate a spatiotemporal grid entity.

[0006] Furthermore, the scheduling vector is generated in the following manner: Build a priority decision tree based on the four-dimensional equity rules and define the rule matching logic; Parse the spatiotemporal grid entity and match the corresponding equity rules based on the priority decision tree and rule matching logic; Based on the equity rules and topological correlation of each grid, the type and corresponding quantity of scheduling actions are dynamically generated, and multi-dimensional ownership binding is performed for each scheduling action and encapsulated into an ownership signature; Integrate scheduling actions, quantities and ownership signatures to generate scheduling vectors with ownership labels.

[0007] Furthermore, when resource contention is detected, the forced preemption policy is triggered and the network resource arbitration log is generated synchronously, including: Define preemption trigger conditions and QoS level priority rules to detect and identify network resource competition events in real time; Based on network resource competition events, it dynamically matches and executes predefined resource preemption strategies according to QoS level priority rules, and simultaneously generates network resource arbitration logs for each preemption operation.

[0008] Furthermore, the confidence arbitration result is generated in the following manner: Define three types of rules: integrity rules, logic rules, and compliance rules, as well as verification allocation logic, to form a three-level dynamic rule chain; Parse scheduling vectors and network resource arbitration logs, execute a three-level dynamic rule chain, and verify field integrity, logical rationality, and compliance. If verification fails, trigger a corresponding error code or confidence adjustment. The confidence level of the scheduling action is dynamically generated based on the rule verification results and integrated into the confidence arbitration result.

[0009] Furthermore, the method of performing path feasibility verification and then generating a scheduling path with time and space constraints includes: Based on the confidence arbitration results, key constraints in heterogeneous public transportation data are extracted to form path verification logic; Perform feasibility check based on the path verification logic. If the feasibility check fails, the circuit breaker mechanism is triggered. Obtain historical traffic volume, weather conditions, and event urgency, deduce and dynamically generate a dispatch path. If the fuse mechanism is triggered, an alternative path is generated as the final dispatch path.

[0010] Furthermore, the method of formulating a shared distribution strategy to classify the associated data sets includes: Correlate the confidence arbitration results, scheduling paths, and network resource arbitration logs and unify their formats to form a correlated dataset; Based on the four-dimensional rights and interests rules, a sharing and distribution strategy is formulated to classify and process the associated data sets; Based on the classification results, hierarchical access rights are provided to different stakeholders through the API interface.

[0011] Furthermore, the data quality assessment report is generated in the following manner: Perform three-level quality checks on the classified linked data sets, including integrity check, consistency check, and compliance check; If the verification result does not meet expectations, it is determined that an abnormal event has been triggered and an automatic correction mechanism is executed; Integrate the three-level quality verification results, abnormal event records and automatic correction records to generate a data quality assessment report.

[0012] Furthermore, the method of forming a closed-loop governance process through reverse optimization through spatiotemporal analysis includes: Conduct spatiotemporal analysis on data quality assessment reports to identify the root causes of data quality issues and generate actionable analysis results; Convert analysis results into optimization actions, drive three-level dynamic rule chain optimization and data source optimization, and form a closed-loop governance process.

[0013] Based on the same inventive concept, a public transportation travel service big data processing system is also proposed, which is implemented based on the above-mentioned public transportation travel service big data processing method and includes: Multi-source spatiotemporal fusion module: accesses heterogeneous public transportation data, calibrates coordinate deviations, integrates real-time meteorological events, and generates spatiotemporal grid entities; Four-dimensional equity scheduling module: Based on the spatiotemporal grid entity, it constructs a priority decision tree and rule matching logic according to four-dimensional equity rules to generate a scheduling vector with ownership tags. When network resource competition is detected, it triggers a forced preemption strategy and simultaneously generates a network resource arbitration log. Rule arbitration and path deduction module: This module integrates scheduling vectors and network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, and performs path feasibility verification to generate scheduling paths with time and space constraints. Multi-type equity data encapsulation module: This module integrates scheduling paths, confidence arbitration results, and network resource arbitration logs to generate related data sets. It then formulates a sharing and distribution strategy based on four-dimensional equity rules to classify and process the related data sets. Evaluation and Optimization Module: Performs three-level quality verification on the associated data sets after classification processing, triggers the anomaly correction mechanism and generates a data quality assessment report, and forms a closed-loop governance process through reverse optimization through spatiotemporal analysis.

[0014] The technical effects and advantages of the present invention are as follows: This method generates a high-precision spatiotemporal grid entity by calibrating coordinate deviations in heterogeneous data and integrating real-time meteorological events. This step addresses the accuracy issues of traditional static coordinate conversion. By dynamically adjusting the network topology density and meteorological influence intensity, it ensures the spatiotemporal consistency of multi-source data, providing a reliable foundation for subsequent scheduling.

[0015] Next, a priority decision tree and rule matching logic are constructed based on the four-dimensional rights and interests rules to generate a scheduling vector with ownership tags. When resource contention is detected, a forced preemption policy is triggered. This step ensures resource allocation priority for high-priority tasks (such as emergency dispatch) through a dynamic priority management mechanism. It also generates a network resource arbitration log for operation traceability, effectively resolving resource conflicts.

[0016] Next, the dispatch vector and arbitration log are integrated, and a three-level dynamic rule chain is executed to generate a confidence arbitration result. A dispatch path with spatiotemporal constraints is then generated through path feasibility verification. This step ensures the legality and safety of dispatch actions through a step-by-step verification of integrity, logic, and compliance rules. It also dynamically generates feasible or alternative paths based on real-time traffic conditions and physical constraints, significantly reducing the risk of path infeasibility.

[0017] Subsequently, the related datasets are categorized through a shared distribution strategy, enabling hierarchical sharing of multiple data types based on four-dimensional rights and interests rules. This step provides differentiated access rights to different stakeholders through APIs, meeting compliance requirements such as government data storage and enterprise data desensitization while protecting individual privacy and addressing compliance challenges associated with traditional data sharing.

[0018] Finally, a closed-loop governance process is formed through three-level quality verification and reverse optimization through spatiotemporal analysis. This step verifies the integrity, consistency, and compliance of the classified data, triggers anomaly correction mechanisms, and generates a quality assessment report. Spatiotemporal analysis further identifies the root causes of data quality issues, drives rule chain optimization and data source improvements, and continuously enhances the system's adaptability.

[0019] Through the above steps, the present invention constructs a full-process system from data fusion, resource scheduling, rule verification to shared governance, effectively solving the shortcomings of existing technologies in time and space alignment, rights and interests scheduling, compliance sharing and path planning, and providing an efficient, safe and traceable systematic solution for public transportation big data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flowchart of the public transportation service big data processing method.

[0021] Figure 2 This is a flow chart of the rule arbitration and path deduction modules in this public transportation service big data processing method.

[0022] Figure 3 This is a flow chart of the public transportation service big data processing system. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0024] Example 1

[0025] See also Figure 1 and Figure 2 As shown, the method for processing public transportation service big data described in this embodiment includes: S1: Access heterogeneous public transportation data, calibrate coordinate deviations, integrate real-time meteorological events, and generate spatiotemporal grid entities; S2: Based on the spatiotemporal grid entity, a priority decision tree and rule matching logic are constructed according to the four-dimensional rights and interests rules to generate a scheduling vector with ownership tags. When network resource competition is detected, a forced preemption strategy is triggered and a network resource arbitration log is generated simultaneously. S3: Integrates scheduling vectors with network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, and performs path feasibility verification to generate scheduling paths with time and space constraints. S4: Fusion of scheduling paths, confidence arbitration results, and network resource arbitration logs to generate a related dataset. Based on the four-dimensional equity rules, a sharing and distribution strategy is formulated to classify and process the related dataset. S5: Perform three-level quality verification on the associated data sets after classification processing, trigger the anomaly correction mechanism and generate a data quality assessment report, and form a closed-loop governance process through reverse optimization through spatiotemporal analysis.

[0026] The generation methods of space-time grid entities include: Use MQTT / HTTP / 2 to receive heterogeneous public transportation data streams, such as bus GPS data (parse the JT808 protocol header (0x7E identifier) ​​to extract WGS84 coordinates), shared bike order data (parse the JSON field to extract GCJ02 coordinates), and online ride-hailing trajectory data (parse the JT905 protocol to extract BD09 coordinates). Verify the legitimacy of the protocol through the protocol header identifier (such as the 0x7E start character of JT808) to filter out invalid data packets. JT808 is used for bus GPS data (0x7E start character), and JT905 is used for online car-hailing trajectory data (0x7F start character). Detect the distance deviation of the heterogeneous coordinate systems of each public transportation equipment in heterogeneous public transportation data; Specifically, the Haversine algorithm is used to calculate the distance deviation between heterogeneous coordinate systems (such as WGS84 vs. GCJ02 vs. BD09). For example, the GCJ02 coordinates of a shared bike order and the WGS84 coordinates of the bus GPS data deviate by 180 meters, marking it as requiring calibration. It should be noted that the coordinate deviation here refers to the deviation between the same location in different coordinate systems. If two data sources (such as shared bicycle orders and bus GPS) record the same physical location (for example, a certain intersection), but use different coordinate systems (such as GCJ02 vs. WGS84), their coordinate values ​​will deviate. This deviation is the mathematical difference in coordinate system conversion, not the actual change in geographic location.

[0027] For example, the actual location of a certain intersection is (31.2304, 121.4737) in the GCJ02 coordinate system, but it may be (31.2315, 121.4748) in the WGS84 coordinate system. There is a distance deviation between the two. The purpose of calibration is to eliminate this difference. For heterogeneous coordinate systems that require calibration, the road network connection density of the corresponding road section is obtained, the topological correlation is extracted (the topological correlation is obtained in the same way as the subsequent calculation of the grid-level topological correlation), real-time meteorological warning data is received, the meteorological event type is analyzed, and a weight coefficient is assigned as the impact intensity of the meteorological event; For example, the meteorological event type and meteorological event impact intensity (weight coefficient) are: red alert for rainstorm (1.5), large-scale event (2.0), and traffic accident (0.8); Combining topological correlation with the impact intensity of meteorological events, a spatiotemporal benchmark alignment engine is built to calibrate distance deviations. All coordinates (including those that do not require calibration and those that have been calibrated) are integrated to generate a calibrated coordinate set (including a protocol conflict marker field). Specifically, the spatiotemporal benchmark alignment engine takes the ratio of the topological correlation to the intensity of the meteorological event impact as the weight value, and then uses the weight value to calibrate the corresponding heterogeneous coordinates to obtain the calibrated coordinates; Calibration coordinate = weight value × heterogeneous coordinate 1 + (1-weight value) × heterogeneous coordinate 2; Assuming that the coordinates with distance deviation are WGS84 coordinates and GCJ02 coordinates, the calibration coordinates = weight value × WGS84 + (1-weight value) × GCJ02; Perform grid encoding on the calibrated coordinate set (generate Geohash level 7 encoding (100-meter accuracy), for example: coordinate (31.2304°N, 121.4737°E) → Geohash level 7 encoding wx4er), extract the road network connection density, and obtain the grid-level topological correlation; The calculation method of topological correlation based on road network connection density is: Basic indicators are extracted from the road network data of each grid, including the number of intersections (the total number of intersections within the grid area, reflecting the node density of the road network), the total length of roads (the total length of roads within the grid area, reflecting the coverage of the road network), the road width (the number of lanes or the average width of each road in the grid (e.g., the width of the main road is 4 lanes and the width of the branch road is 2 lanes)), and the regional area (geographical scope (i.e., the area of ​​the grid area)). The width of each road (such as the number of lanes) is then converted into a width weight to calculate the average width within the grid. That is, the corresponding width weight is weighted and fused with the corresponding road length. The weighted fusion result is then ratioed with the total road length to obtain the average width of the roads within the grid. The number of intersections, total road length, and average width are standardized to eliminate dimensional differences. Each of these is then assigned a weight (the weight needs to be adjusted based on actual needs, such as through expert scoring or regression analysis. Usually, the sum of the three weights is 1). These weighted factors are then combined to obtain the topological correlation. Integrate the grid-level topological correlation and meteorological event impact intensity of all grids to generate grid basic data containing meteorological weights and topological correlation. Then use the SHA256 algorithm to hash each grid ID and timestamp to generate a checksum to ensure data integrity. Integrate wx4er grid ID, meteorological event impact intensity, topological correlation, and check code to generate a space-time grid entity.

[0028] The scheduling vector is generated in the following ways: Construct a priority decision tree based on the four-dimensional rights and interests rule (government > enterprise > individual > operator): [Government Needs] →|Emergency Dispatch| (Highest Priority); [Enterprise Needs] →|Business Analysis| (Medium Priority); [Personal Needs] →|Trajectory Query| (Low Priority); [Operational Revenue Needs] →|Resource Optimization| (Lowest Priority); The rule matching logic is defined such that when priorities conflict, higher-priority benefits automatically override lower-priority benefits. If there are multiple benefit requests of the same priority, they are processed on a first-come, first-served basis. Parse the fields in the spatiotemporal grid entity (wx4er grid ID: target dispatch area, event weight: meteorological event impact (such as rainstorm weight 1.5), topological correlation: road network connection strength (such as 0.87)), and match the corresponding equity rules according to the priority decision tree and rule matching logic; For example, if the grid ID is wx4er and the meteorological event impact intensity is 1.5 (red alert for heavy rain), the government emergency dispatch priority is triggered; Based on the equity rules and topological relevance of each grid, the type and corresponding number of scheduling actions are dynamically generated; Specifically, the intensity of meteorological events drives the matching of dispatch action types, and the number of corresponding actions is quantified through topological correlation. Combined with matching equity rules and flow control mechanisms, dynamic generation of dispatch instructions and optimal resource allocation are achieved, ensuring that high-priority demands (such as government emergencies) are prioritized while also taking into account the reasonable demands of enterprises, individuals, and operators. Dynamically match the scheduling action type based on the priority decision tree of the four-dimensional equity rule and the grid status; Exemplary action type classification and matching logic: Scheduling action type, triggering conditions and priority: Supplies → Government emergency needs (such as heavy rain warnings) → Government (highest priority); Analysis → Enterprise business analysis needs (such as road network traffic forecasting) → Enterprise (medium priority); Query → Personal trajectory query demand → Personal (low priority); Optimization → Operator resource optimization requirements → Operator (lowest priority); For example, if the grid ID is wx4er, the meteorological event impact intensity is 1.5 (red alert for heavy rain), and the priority is the highest priority government emergency dispatch, then the resupply action is triggered (government priority); If the topological correlation is 0.87 (dense road network), the enterprise may initiate an analytical action (e.g., analyzing traffic flow); The number of corresponding dispatch actions is determined by the intensity of the meteorological event impact and the topological correlation, and is adjusted in combination with the equity rules and flow control mechanism; The calculation logic is: Government supply: supply = event weight × basic supply coefficient; Enterprise analysis volume: analysis volume = topological correlation × basic analysis coefficient; Operation optimization amount: optimization amount = road network density × basic optimization coefficient; The flow control mechanism adjusts the logic as follows: if the calculated result exceeds the system resource limit (e.g., government supply of 150 exceeds inventory by 50), the flow control is adjusted, including: adjusting to the inventory limit, dynamically adjusting the quantity based on resource availability (e.g., adjusting to 50), preserving the priority of the equity rule and forcibly overwriting lower priority actions; For example, equity type, action type, and quantity calculation logic: Government → Recharge → Recharge amount in rainstorm area = event weight × 100; Enterprise → Analysis → Data Analysis Request = Topological Relevance × 50; Individual → Query → Track query times = 1 / user; Operator → Optimization → Resource Optimization Suggestion = Road Network Density × 20; Then, multi-dimensional ownership binding is performed for each scheduling action and encapsulated into an ownership signature, that is, an ownership tag is attached to each scheduling action; For example, the government: binds the blockchain license; the enterprise: attaches the differential privacy identifier; the individual: attaches the user identity hash; the operator: attaches the API key hash; Integrate scheduling actions and ownership signatures to generate a scheduling vector with ownership tags, including grid ID, action type, quantity, and a list of ownership signatures for multi-stake binding.

[0029] When resource contention is detected, the forced preemption policy is triggered and network resource arbitration logs are generated simultaneously in the following ways: Define preemption trigger conditions and QoS level priority rules to detect and identify network resource competition events in real time; Specifically, 5G channel metrics such as bandwidth occupancy, QoS (Quality of Service) level, and channel latency are collected in real time; The preemption trigger condition is defined as: if the bandwidth usage rate does not meet the preset threshold (such as <10Mbps), the resource preemption strategy is triggered; if the high-level priority scheduling instructions (such as government and enterprise) cannot be executed due to insufficient resources, the resource preemption strategy is forcibly activated; QoS priority rules are defined as follows: network resource allocation priorities are differentiated according to QoS levels, such as Level 0 (government emergency, absolute priority, preempting other network resources), Level 1 (enterprise needs, medium priority), and Level 2 (commercial advertising push, low priority, preemptible). Then, based on the real-time network resource status and QoS level priority rules, it dynamically matches and executes predefined resource preemption strategies, including interrupting low-priority network resources and allocating high-priority network resources; Specifically, it releases the bandwidth occupied by Level 2 (commercial advertisements) (e.g., interrupting advertisement streaming) and dynamically allocates the released bandwidth to Level 0 (government emergency dispatch instructions); Develop network resource isolation guarantees and isolate Level 0 network resources through 5G slicing technology to ensure their exclusivity; After preemption, the system continuously monitors the duration of Level 0 resource usage. If the Level 0 task is completed, it automatically releases network resources and resumes Level 2 services. If preemption fails (e.g., insufficient resources), it triggers a circuit breaker mechanism (e.g., downgrading services). A network resource arbitration log is generated for each preemption operation, including timestamp, event type (e.g., resource arbitration), action type (e.g., preempting a Level 2 channel), target (e.g., guaranteeing scheduling instructions), resource status change (including duration, pre-resource status (e.g., "Level 2 network resource utilization": "85%", "Bandwidth utilization": "9 Mbps"), post-resource status (e.g., "Level 0 network resource utilization": "100%", "Level 2 utilization": "0%"), and ownership signature. The network resource arbitration log is then hashed using the SHA256 hash algorithm and stored as a blockchain distributed ledger to ensure that the data content cannot be tampered with.

[0030] The confidence arbitration results are generated in the following ways: Define the verification and allocation logic for three types of rules: integrity, logic, and compliance, forming a three-level dynamic rule chain; Among them, the three-level dynamic rule chain includes: Integrity rules: Verify that the fields in the dispatch vector and network resource arbitration log are complete (such as GPS trajectory, borrowing time, number of empty vehicles, usage scope, etc.); Logical rules: verify the rationality of scheduling actions (such as the matching of vehicle borrowing time and the number of empty vehicles, and the spatiotemporal continuity of resource allocation); Compliance rules: Ensure that scheduling actions comply with laws and regulations (such as the Data Security Law's restrictions on data usage); The scheduling vector and network resource arbitration logs are parsed, and a three-level dynamic rule chain is executed step by step through the federation engine to verify field integrity, logical rationality, and compliance. If the verification fails, the corresponding error code or confidence adjustment is triggered. The logic of the three-level dynamic rule chain's step-by-step verification can be described by the following example: Integrity rules: Field: "GPS track", verification condition: "not empty and continuous"; Field: "borrowing time", verification condition: "format is YYYY-MM-DD HH:MM:SS"; Integrity verification: Input the GPS track, detect missing points → mark error code 1001 → generate exception description: GPS track is incomplete; Enter the rental time, and the format is found to be non-compliant → error code 1002 is generated → time format error; Logical rules: Condition: "Borrowed vehicle time > Number of empty vehicles × 2 hours", Scheduling action: "Trigger logical conflict"; Condition: "Network resource arbitration log timestamp is inconsistent with scheduling time", Scheduling action: "Mark exception"; Logical rationality verification: Enter the borrowing time and the number of empty cars (the number of empty cars left after borrowing). If the borrowing time is greater than the corresponding length of the empty car count, error code 2003 will be generated, indicating a discrepancy between the borrowing time and the number of empty cars. Enter the timestamp of the network resource arbitration log. If the timestamp is not within the scheduled time window, error code 2004 will be generated. The timestamp is inconsistent. Compliance rules: Clause: "Data Security Law", verification condition: "Data usage does not exceed the authorized region"; Clause: "Cybersecurity Law", verification condition: "No sensitive data sharing involved"; Compliance Verification: Detection reveals that the scope of use exceeds the authorized area of ​​the Data Security Law → Error code 3007 is generated → Compliance violation: The input data contains sensitive data and is not desensitized. Error code 3008 is generated, which indicates a data security violation. Dynamically generate the confidence level of the scheduling action based on the rule verification results, and then generate the confidence arbitration result, including the error code, confidence level, and verification result (including arbitration time and exception description); Specifically, the confidence level of a scheduling action is generated by assigning a weight to each rule verification (determined by expert scoring, with the total weight being 1 (i.e., 100%)), e.g., 30% for completeness, 40% for logic, and 30% for compliance). After the rule is verified, the score is dynamically adjusted based on the verification results. The logic of dynamically adjusting the score is defined as follows: when all rules are passed, the total confidence score is set to 1. If a rule fails, the weight ratio of the rule is subtracted from the total score; For example, assume that the integrity check weight is 30%: if the integrity rule fails, 30% of the score will be deducted from the total score; If multiple rules fail, the weights are deducted cumulatively, but the total score is not less than 0.0; If the confidence score is lower than the preset confidence threshold (such as 0.6), a circuit breaker is triggered, which prohibits the scheduling action from being executed and logs are recorded.

[0031] Methods for performing path feasibility verification and generating a scheduling path with time and space constraints include: Based on the confidence arbitration results, key constraints in heterogeneous public transportation data are extracted (including road height restrictions and road access rights, which can also be adjusted according to actual needs). Path verification logic is then formed, and feasibility verification is then performed. If the feasibility verification fails, a circuit breaker mechanism is triggered (i.e., path generation is prohibited). Among them, the path verification logic: Taking the height limit detection verification logic as an example, the height limit detection mainly detects the height limit node and vehicle height; Height restriction nodes are road height restrictions that exist on the road (such as overpasses); Therefore, the height limit detection and verification logic is that if the path contains a height limit node and the vehicle height exceeds the height limit of the height limit node, the fuse mechanism is triggered and the path generation is prohibited; Other constraint checks can be added based on actual conditions, such as road right-of-way checks (used to verify whether dispatch vehicles have access to specific roads (such as emergency lanes, traffic restrictions)), and spatiotemporal continuity checks (used to ensure that height-restricted nodes are coherent in time and space (such as the distance between adjacent height-restricted nodes does not exceed the maximum speed limit)). Obtain historical traffic volume, weather conditions, and event urgency, deduce and dynamically generate dispatch routes under time and space constraints, and generate alternative routes as the final dispatch routes if a circuit breaker is triggered. Specifically, it obtains historical traffic volume (such as the traffic volume of each road section in the past hour), weather conditions (such as rainfall and visibility), and event urgency (such as the urgency score of traffic accidents and road closures due to construction) as input. Then, it uses machine learning methods (such as those based on long short-term memory networks (LSTM models)) to predict and output traffic flow and congestion status in the future time period. For example, it predicts the congestion probability of each road section in the next 30 minutes (probability range is 0.0-1.0); Then, based on the prediction results of road congestion, the optimization goal is defined as minimizing the total travel time, and the optimal scheduling path is generated through the route planning algorithm; Example, optimal dispatch path: Path nodes: ["Starting Point 1", "Stop Point 1", "End Point 1"], Estimated Travel Time: "45 minutes", Congestion Probability: 0.2; Among them, the path nodes are vehicle stops (such as bus stops and shared bicycle parking spots); If a path triggers a circuit breaker due to height restrictions or other constraints, an alternative path algorithm (such as the Dijkstra shortest path algorithm) is automatically invoked to generate alternative paths. This is then combined with real-time congestion probability optimization to arrive at a new optimal dispatch path. Example of a fuse result: Error code 4001, Reason: "Height limit detection failed (overpass X height limit)", Alternative paths: ["Starting point 1", "Detour 2", "End point 1"]; Integrate the final dispatch path and path attributes (such as travel time and congestion probability) as the basis for subsequent execution; Ways to develop a shared distribution strategy to categorize linked datasets include: The confidence arbitration results, scheduling paths, and network resource arbitration logs are associated and formatted in a unified manner. Specifically, the association is performed as follows: The confidence arbitration result is associated with the dispatch path according to the dispatch ID to form a "path + confidence" combination. A unique ID is added to each dispatch path as the dispatch ID. The blockchain-stored hash of the network resource arbitration log is associated with the path nodes of the scheduling path to ensure traceability of path resource allocation and to provide resource usage dynamics (e.g., "Level 2 channel at stop 1 is preempted"). The unified format is: convert all data into JSON / XML format and standardize field names; Integrate and unify the confidence arbitration results, scheduling paths, and network resource arbitration logs in a unified format to form a related dataset; Based on the four-dimensional rights and interests rules (government, enterprise, individual, operator), a sharing and distribution strategy is formulated to classify the associated data sets; The shared distribution strategy is to use blockchain technology to store government data (such as dispatch paths and resource arbitration logs) on-chain, generate ownership audit packages, and ensure data traceability. Differential privacy technology is used to desensitize enterprise data (such as desensitized riding trajectories and shared bicycle path stop and change point data) to generate desensitized sharing packages to protect the privacy of commercial data. Anonymize personal data (such as user identity and behavior trajectory) and then distribute it with user authorization; Operator data (such as IoT device status and vehicle column vacancy rate) must be negotiated with device manufacturers for permissions to ensure operational compliance, legality of data use, and closed-loop management. Furthermore, through the API interface, hierarchical access rights are provided to different stakeholders to achieve compliant sharing and dynamic governance of data; Among them, different equity entities correspond to different API interfaces: Government data provides data call services through government-specific API interfaces, supporting audit tracing and real-time query; Enterprise data provides data subscription services through the enterprise's internal API interface, supporting on-demand acquisition of specific fields (such as confidence scores and resource arbitration events); Individuals and operators need to obtain limited access through approval; It should be noted that the setting logic of the shared distribution strategy is to carry out hierarchical processing based on the different nature of the four-dimensional rights and interests (including government data notarization, enterprise data desensitization, personal data anonymization, and operator data authorization).

[0032] The data quality assessment report can be generated by: Perform three-level quality checks on the classified linked data sets, including integrity check, consistency check, and compliance check; If the verification result does not meet expectations, it is determined that an abnormal event has been triggered and an automatic correction mechanism is executed; Among them, integrity verification includes field-level verification and data-level verification, the purpose of which is to ensure that there are no missing or abnormal data during transmission and storage; Field-level verification is to check whether there are any missing fields in the associated dataset after classification processing (for example, the dispatch path lacks the "Stop 1" path node). If there are any missing fields, mark them as abnormal events. Data-level verification involves comparing the data volumes of the four types of data divided by the shared distribution strategy. If the difference rate exceeds a preset difference threshold (e.g., 5%), it is marked as an abnormal event. The automatic correction mechanism automatically executes data completion tasks through ETL tools and records completion logs; For data that cannot be automatically completed, a manual compound process is triggered and the operation and maintenance personnel are notified to handle it; Consistency checks include conflict event checks and timestamp checks, which aim to ensure the consistency of the four types of data and avoid data tampering or logical conflicts. Conflict event verification involves comparing the field consistency of four types of data (such as path node sequences and resource arbitration events); Example: If the ownership audit package records "starting point 1 → stop point 2 → end point 3", but the desensitized shared package shows "starting point 1 → end point 3", it is considered a conflict and marked as an abnormal event; Timestamp verification is to verify whether the timestamps of the four types of data match. If the difference exceeds the preset time difference threshold (such as 10 seconds), it is marked as an abnormal event; The automatic correction mechanism is that if there is a conflict, the original data from the data type with more complete fields will be extracted first, covering the missing parts in other data types; if there is a timestamp anomaly, an alarm will be triggered; For conflict events that cannot be automatically repaired, a manual review process is triggered and the data governance team is notified to handle the conflict. Compliance verification includes privacy compliance checks and legal compliance scans to ensure that shared data complies with laws and regulations (such as the Data Security Law) and privacy protection requirements; Privacy compliance check: Verify whether sensitive fields are correctly desensitized (for example, user identity fields are irreversibly encrypted). If not, mark it as an abnormal event. Legal compliance scanning: Automatically scans highly sensitive fields (such as personal identification and device unique IDs) in shared data and triggers an alarm if unmasked fields are found. The automatic correction mechanism is to automatically adjust the differential privacy parameters in the differential privacy technology (such as increasing the ε value to 0.05) if there are unmasked fields, and regenerate the masked shared package; For incidents involving significant compliance risks (such as leaking user privacy), a manual approval process is triggered and the legal team intervenes to handle the matter; Integrate the three-level quality verification results, abnormal event records and automatic correction records to generate a data quality assessment report.

[0033] Through reverse optimization through spatiotemporal analysis, the closed-loop governance process is formed in the following ways: Conduct spatiotemporal analysis on data quality assessment reports to identify the root causes of data quality issues and generate actionable analysis results; Specifically, the spatiotemporal dimension analysis is as follows: using a spatiotemporal clustering algorithm (such as DBSCAN) to count the areas (such as core urban areas and transportation hubs) where dispatching paths are concentrated during peak traffic hours (such as 7:00-9:00), and record them as high-frequency paths; At the same time, through spatiotemporal trend analysis algorithms (such as time series analysis), abnormal spatiotemporal points in the dispatch path (such as a sudden increase in traffic flow on a certain road section at night, or a missing intermediate node in a certain area leading to a broken path) are identified and recorded as abnormal areas. These abnormal spatiotemporal points can be judged as data quality issues. This generates a spatiotemporal distribution heat map of dispatch paths that marks high-frequency paths and abnormal areas; A list of abnormal areas is generated simultaneously, including the specific location, time range, and problem type (such as missing fields and inconsistent timestamps); At the same time, it integrates the spatiotemporal distribution heat map of the dispatch path, the abnormal area list and the abnormal data field statistics to generate a spatiotemporal distribution analysis report for submission to the operation personnel for data management and data analysis; Convert analysis results into optimization actions, drive three-level dynamic rule chain optimization and data source optimization, and form a closed-loop governance process; Specifically, it is the specific task of converting the analysis results into three-level dynamic rule chain optimization and data source optimization; Optimization of the three-level dynamic rule chain: The field missing rate in abnormal areas (e.g., a path node missing rate of 15%) is fed back into the three-level dynamic rule chain to verify each level and generate the confidence level of the scheduling action, thereby increasing the weight of the integrity rule. That is, reduce the weights of the compliance rule and the integrity rule at the same time, and add the reduced weights to the integrity rule weight; Data source optimization: Abnormal areas in spatiotemporal analysis (such as a sudden increase in traffic on a certain road section at night) are fed back to the heterogeneous public transportation data reception and collection module; Adjust data collection strategies (e.g., increase sensor collection frequency at night, fix interface errors, and thus enhance reception of heterogeneous public transportation data).

[0034] Example 2

[0035] See also Figure 3 As shown, for the parts not described in detail in this embodiment, please refer to the description of Example 1. A public transportation travel service big data processing system is provided, including: Multi-source spatiotemporal fusion module: This module accesses heterogeneous public transportation data, builds a spatiotemporal benchmark alignment engine to calibrate coordinate deviations, integrates real-time meteorological events, and generates spatiotemporal grid entities. Four-dimensional equity scheduling module: Based on the spatiotemporal grid entity, it constructs a priority decision tree and rule matching logic according to four-dimensional equity rules to generate a scheduling vector with ownership tags. When network resource competition is detected, it triggers a forced preemption strategy and simultaneously generates a network resource arbitration log. Rule arbitration and path deduction module: This module integrates scheduling vectors and network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, and performs path feasibility verification to generate scheduling paths with time and space constraints. Multi-type equity data encapsulation module: This module integrates scheduling paths, confidence arbitration results, and network resource arbitration logs to generate related data sets. It then formulates a sharing and distribution strategy based on four-dimensional equity rules to classify and process the related data sets. Evaluation and Optimization Module: Performs three-level quality verification on the associated data sets after classification processing, triggers the anomaly correction mechanism and generates a data quality assessment report, and forms a closed-loop governance process through reverse optimization through spatiotemporal analysis.

[0036] Example 3

[0037] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the operation mode of the public transportation service big data processing system provided above is realized.

[0038] Since the electronic device introduced in this embodiment is an electronic device used to implement a method for processing big data of public transportation services in the embodiment of this application, based on the method for processing big data of public transportation services introduced in the embodiment of this application, those skilled in the art can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application will not be described in detail here. As long as those skilled in the art implement the electronic device used in the method for processing big data of public transportation services in the embodiment of this application, they are within the scope of protection to be provided by this application.

[0039] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.

[0040] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for users of ordinary skill in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for processing public transportation service big data, characterized in that: include: S1: Access heterogeneous public transportation data, calibrate coordinate deviations, integrate real-time meteorological events, and generate spatiotemporal grid entities; S2: Based on the spatiotemporal grid entity, a priority decision tree and rule matching logic are constructed according to the four-dimensional rights and interests rules to generate a scheduling vector with ownership tags. When network resource competition is detected, a forced preemption strategy is triggered and a network resource arbitration log is generated simultaneously. S3: Integrates scheduling vectors with network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, and performs path feasibility verification to generate scheduling paths with time and space constraints. S4: Fusion of scheduling paths, confidence arbitration results, and network resource arbitration logs to generate a related dataset. Based on the four-dimensional equity rules, a sharing and distribution strategy is formulated to classify and process the related dataset. S5: Perform three-level quality verification on the associated data sets after classification processing, trigger the anomaly correction mechanism and generate a data quality assessment report, and form a closed-loop governance process through reverse optimization through spatiotemporal analysis.

2. A method for processing public transportation service big data according to claim 1, characterized in that: The generation method of the space-time grid entity includes: Detecting distance deviations in heterogeneous coordinate systems in heterogeneous public transportation data; Obtain the road network connection density of sections with distance deviation and extract the topological correlation; receive meteorological warning data, analyze the type of meteorological event, and assign a weight coefficient as the impact intensity of the meteorological event; Combining topological correlation with the impact intensity of meteorological events, a spatiotemporal benchmark alignment engine is constructed to calibrate distance deviations and generate a calibration coordinate set. Grid encode the calibration coordinate set, obtain the road network data of each grid, and calculate the grid-level topological correlation of each grid; The grid-level topological correlation of all grids and the impact intensity of meteorological events are integrated to generate a spatiotemporal grid entity.

3. A method for processing public transportation service big data according to claim 2, characterized in that: The generation method of the scheduling vector includes: Build a priority decision tree based on the four-dimensional equity rules and define the rule matching logic; Parse the spatiotemporal grid entity and match the corresponding equity rules based on the priority decision tree and rule matching logic; Based on the equity rules and topological correlation of each grid, the type and corresponding quantity of scheduling actions are dynamically generated, and multi-dimensional ownership binding is performed for each scheduling action and encapsulated into an ownership signature; Integrate scheduling actions, quantities and ownership signatures to generate scheduling vectors with ownership labels.

4. A method for processing public transportation service big data according to claim 3, characterized in that: When resource contention is detected, the forced preemption policy is triggered and the network resource arbitration log is generated synchronously. The method includes: Define preemption trigger conditions and QoS level priority rules to detect and identify network resource competition events in real time; Based on network resource competition events, it dynamically matches and executes predefined resource preemption strategies according to QoS level priority rules, and simultaneously generates network resource arbitration logs for each preemption operation.

5. A method for processing public transportation service big data according to claim 4, characterized in that: The confidence arbitration result is generated in the following manner: Define three types of rules: integrity rules, logic rules, and compliance rules, as well as verification allocation logic, to form a three-level dynamic rule chain; Parse scheduling vectors and network resource arbitration logs, execute a three-level dynamic rule chain, and verify field integrity, logical rationality, and compliance. If verification fails, trigger a corresponding error code or confidence adjustment. The confidence level of the scheduling action is dynamically generated based on the rule verification results and integrated into the confidence arbitration result.

6. A method for processing public transportation service big data according to claim 5, characterized in that: The method of performing path feasibility verification and then generating a scheduling path with time and space constraints includes: Based on the confidence arbitration results, key constraints in heterogeneous public transportation data are extracted to form path verification logic; Perform feasibility check based on the path verification logic. If the feasibility check fails, the circuit breaker mechanism is triggered. Obtain historical traffic volume, weather conditions, and event urgency, deduce and dynamically generate a dispatch path. If the fuse mechanism is triggered, an alternative path is generated as the final dispatch path.

7. A method for processing public transportation service big data according to claim 6, characterized in that: The method of formulating a shared distribution strategy to classify the associated data sets includes: Correlate the confidence arbitration results, scheduling paths, and network resource arbitration logs and unify their formats to form a correlated dataset; Based on the four-dimensional rights and interests rules, a sharing and distribution strategy is formulated to classify and process the associated data sets; Based on the classification results, hierarchical access rights are provided to different stakeholders through the API interface.

8. A method for processing public transportation service big data according to claim 7, characterized in that: The data quality assessment report is generated in the following manner: Perform three-level quality checks on the classified linked data sets, including integrity check, consistency check, and compliance check; If the verification result does not meet expectations, it is determined that an abnormal event has been triggered and an automatic correction mechanism is executed; Integrate the three-level quality verification results, abnormal event records and automatic correction records to generate a data quality assessment report.

9. A method for processing public transportation service big data according to claim 8, characterized in that: The methods of forming a closed-loop governance process through reverse optimization through spatiotemporal analysis include: Conduct spatiotemporal analysis on data quality assessment reports to identify the root causes of data quality issues and generate actionable analysis results; Convert analysis results into optimization actions, drive three-level dynamic rule chain optimization and data source optimization, and form a closed-loop governance process.

10. A public transportation travel service big data processing system, implemented based on a public transportation travel service big data processing method according to any one of claims 1 to 9, characterized in that: include: Multi-source spatiotemporal fusion module: accesses heterogeneous public transportation data, calibrates coordinate deviations, integrates real-time meteorological events, and generates spatiotemporal grid entities; Four-dimensional equity scheduling module: Based on the spatiotemporal grid entity, it constructs a priority decision tree and rule matching logic according to four-dimensional equity rules to generate a scheduling vector with ownership tags. When network resource competition is detected, it triggers a forced preemption strategy and simultaneously generates a network resource arbitration log. Rule arbitration and path deduction module: This module integrates scheduling vectors and network resource arbitration logs, defines and executes a three-level dynamic rule chain, generates confidence arbitration results, and performs path feasibility verification to generate scheduling paths with time and space constraints. Multi-type equity data encapsulation module: This module integrates scheduling paths, confidence arbitration results, and network resource arbitration logs to generate related data sets. It then formulates a sharing and distribution strategy based on four-dimensional equity rules to classify and process the related data sets. Evaluation and Optimization Module: Performs three-level quality verification on the associated data sets after classification processing, triggers the anomaly correction mechanism and generates a data quality assessment report, and forms a closed-loop governance process through reverse optimization through spatiotemporal analysis.

Citation Information

Patent Citations

  • Multi-source data fusion traffic meteorological risk prediction and route planning method and system

    CN118966956A

  • Multi-source heterogeneous data fusion sky-ground intelligent monitoring system

    CN119691657A

  • Traffic flow intelligent sensing scheduling method and system under special meteorological conditions

    CN120183177A

  • High-positioning-precision vehicle adaptive navigation method and vehicle-mounted navigator

    CN120252774A

  • Heterogeneous data processing optimization system based on intelligent edge computing

    CN120336021A