Traffic data quality scoring and automatic grading method based on view consistency detection

By employing multi-view consistency detection and a minimum modifiable set algorithm, the problem of cross-view conflict identification and automatic classification in traffic data quality management was solved, achieving a fully automated closed loop and improving the scientific nature and automation level of data quality management.

CN121350517BActive Publication Date: 2026-02-13GUANGZHOU JIAOXIN INVESTMENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511924279.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-02-13
Estimated Expiration
2045-12-19

AI Technical Summary

Technical Problem

Existing traffic data quality detection methods lack systematic and operable multi-view consistency analysis, cannot effectively identify cross-view conflicts, and have unclear correction strategies, which leads to the masking of data quality problems and makes it difficult to achieve real-time automated classification and processing.

Method used

By detecting data consistency in parallel across multiple views and determining priority correction fields based on the minimum modifiable set, automatic hierarchical classification and triggering of processing actions are achieved. Parallel modeling and cross-view consistency analysis are conducted using three types of views: location, event, and behavior, and the correction strategy is dynamically updated to form a fully automated closed loop.

Benefits of technology

It significantly improves the comprehensiveness and accuracy of conflict detection, avoids blind corrections, enhances the scientific and automated level of data quality management, ensures the availability of high-quality data, and prevents low-quality data from polluting the core system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121350517B_ABST
    Figure CN121350517B_ABST
Patent Text Reader

Abstract

The application discloses a traffic data quality scoring and automatic grading method based on view consistency detection, comprising: collecting and integrating multi-source traffic data, performing time alignment and weight initialization; dividing fields into three types of views of position, event and behavior, respectively calculating view field correlation and generating a conflict matrix between views; solving the minimum modifiable subset based on the conflict matrix and the weight, and prioritizing the conflict fields; calculating the consistency score in each view and merging into a comprehensive score, performing the first round of automatic grading and correction; sequentially performing multi-dimensional consistency detection on the corrected data, dynamically updating the correction priority and the minimum modifiable subset, and performing iterative correction until convergence; realizing automatic quality grading based on the weighted consistency score, and triggering corresponding automatic data management strategies. The application realizes the whole-process automatic evaluation and management of traffic data quality, and improves the data consistency and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of traffic data governance and intelligent analysis technology, and more specifically to a multi-view-based approach. Figure One Traffic data quality scoring and automatic grading method for consistency detection. Background Technology

[0002] In recent years, with the deepening of the construction of intelligent transportation systems, traffic data has become characterized by multiple sources, massive volume, and heterogeneity. Its quality directly affects the reliability of traffic analysis, management decisions, and intelligent applications. However, existing traffic data quality management methods generally have the following shortcomings:

[0003] I. Existing traffic data quality inspection methods are mostly based on a single perspective or a single indicator, such as only detecting location accuracy or event completeness, lacking a systematic and operational multi-view approach. Figure One Consistency analysis methods. Therefore, when logical contradictions arise between data in different views, traditional methods cannot effectively identify such cross-view conflicts, leading to the masking of data quality issues.

[0004] Second, when data inconsistencies are discovered, existing methods lack clear correction strategies, often relying on domain expert experience or simple heuristics to determine the fields that need correction, lacking scientific methods to determine the optimal fields for correction. This can lead to over-correction, affecting the original accuracy of the data, or it may miss critical erroneous fields, making the correction process iterative and inefficient.

[0005] Third, current data quality scoring relies heavily on manual statistics or threshold judgments, which cannot achieve real-time, automated classification and processing of data quality. There is a lack of automated triggering and linkage mechanisms between quality classification and subsequent processing actions, resulting in a broken data processing flow, slow response, and difficulty in meeting the high requirements of intelligent transportation systems for data timeliness and accuracy.

[0006] Therefore, how to provide a traffic data quality scoring and automatic grading method that can detect data consistency through multiple views in parallel, determine the priority correction field by combining the minimum modifiable set, and realize automatic grading and triggering processing actions is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] In view of the above problems, the present invention is proposed to provide a vision-based solution that overcomes or at least partially solves the above problems. Figure One A traffic data quality scoring and automatic grading method based on consistency detection. This method utilizes parallel multi-view data consistency detection, determines priority correction fields based on the minimum modifiable set, and implements automatic grading and triggering of processing actions to improve the scientific and automated level of traffic data quality management.

[0008] To achieve the above object, the present application adopts the following technical solutions:

[0009] The embodiment of the present application provides a traffic data quality scoring and automatic grading method based on consistency detection, comprising the following steps: Figure One

[0010] S1: Collect and integrate multi-source traffic data to form data records containing position, event and behavior information; perform time stamp alignment on the multi-source traffic data; and initialize initial weights of each field based on field statistical characteristics;

[0011] S2: Divide the data fields into three types of views, i.e., position view, event view and behavior view; and calculate the correlation between the fields in each view respectively;

[0012] S3: Generate a conflict matrix between the views based on the value deviation of the fields across the views; solve a first minimum modifiable subset for eliminating conflicts based on the conflict matrix and the initial weights, and perform a first correction priority ranking on the conflict fields according to the first minimum modifiable subset;

[0013] S4: Calculate the intra-view consistency score of each view based on the correlation between the fields in each view and the conflict matrix; and fuse the intra-view consistency scores of all the views to obtain an inter-view comprehensive consistency score of the current data record;

[0014] S5: Perform a first round of automatic quality grading on the data record according to the inter-view comprehensive consistency score, and trigger a first correction action on the data according to the grading result and the first minimum modifiable subset, wherein the correction sequence follows the first correction priority ranking; and re-calculate and update the conflict matrix after the correction;

[0015] S6: Perform time sequence consistency detection on the data after the first correction, and calculate a first consistency comprehensive score;

[0016] S7: Further perform cross-regional position consistency detection, event sequence consistency detection and behavior pattern consistency detection on the data after the first correction, and obtain a second consistency comprehensive score;

[0017] S8: Dynamically update the correction priority ranking and the minimum modifiable subset of the conflict fields based on the second consistency comprehensive score, to obtain a second correction priority ranking and a second minimum modifiable subset; and trigger a second correction action on the data according to the second minimum modifiable subset, wherein the correction sequence follows the second correction priority ranking;

[0018] ​S9: Determine whether the convergence conditions are met. The convergence conditions include: no conflict in the conflict matrix or the number of iterations reaches the preset upper limit. If not met, return to S7 and perform the next round of iteration with the data after secondary correction and the updated priority and subset. If met, calculate the final consistency score based on the weighted sum of the first consistency comprehensive score and the second consistency comprehensive score.

[0019] S10: Based on the final consistency score, perform final automatic quality classification on the data records, and trigger the corresponding automated data governance strategy based on the quality classification results.

[0020] Preferably, the multi-source traffic data includes at least one of vehicle satellite positioning trajectory data, vehicle sensor data, dispatch logs, passenger boarding and alighting records, ticketing information, traffic signal status, and road congestion status.

[0021] Preferably, the step S1 of initializing the initial weights of each field based on the statistical characteristics of the field includes: setting the initial weights according to the variance of the field in the historical data, wherein the variance value is inversely proportional to the initial weight value.

[0022] Preferably, in step S2:

[0023] The location view includes vehicle satellite positioning coordinates, positioning station information, and road segment markings;

[0024] The event view includes dispatch instructions, fault alarms, and operational event records;

[0025] The behavioral view includes passenger boarding and alighting behavior, ticketing information, and vehicle operation behavior.

[0026] Preferably, step S2, which calculates the correlation between fields within each view, includes: calculating the correlation for continuous fields using the Pearson correlation coefficient and calculating the correlation for discrete fields using the mutual information metric.

[0027] Preferably, in steps S3 and S8, the first minimum modifiable subset and the second minimum modifiable subset are both solved using a heuristic search algorithm. The heuristic search algorithm adds the candidate correction set in order of priority, from small to large initial weights of the fields and from large to small number of times they participate in the conflict in the conflict matrix.

[0028] Preferably, step S4, which involves fusing the intra-view consistency scores of all views, includes: fusing the intra-view consistency scores of all views using a weighted average method, where the weights are the normalized values ​​of the initial weights of the fields in each view.

[0029] Preferably, the time series consistency detection step in step S6 includes:

[0030] Perform cross-time consistency checks on time series data:

[0031]

[0032] in, For standardized timestamps, For the first Data records are in timestamps field values, For the number of time points, This represents the similarity function.

[0033] Preferably, in step S7:

[0034] Cross-regional location consistency detection includes: determining whether the geographical areas where a vehicle is located within adjacent time periods satisfy the road network topology connectivity;

[0035] The event sequence consistency detection includes: verifying whether the scheduling event and the actual vehicle behavior are logically consistent in time sequence;

[0036] The behavioral pattern consistency detection includes comparing the similarity between the current passenger boarding and alighting behavior and historical travel models or typical route patterns.

[0037] Preferably, the automated data governance strategy is triggered based on the final quality rating, including:

[0038] A structured report is generated based on the final automatic quality rating, historical data of field corrections, and statistical data of conflicting fields;

[0039] Based on the final quality rating, match one of the following automated data governance strategies: automatically correct and record, mark anomalies and review, or refuse to store data.

[0040] This invention provides a view-based embodiment. Figure One The traffic data quality scoring and automatic grading method based on consistency detection, and the beneficial effects of the above technical solution, include at least the following:

[0041] This invention utilizes parallel modeling and cross-viewing of three types of views: location, event, and behavior. Figure One Consistency analysis effectively uncovers data contradictions that are difficult to detect from a single perspective, significantly improving the comprehensiveness and accuracy of conflict detection.

[0042] This invention introduces a minimum modifiable subset (MUS) mechanism, which combines field weights and conflict participation to prioritize the correction of fields that have the greatest impact on overall consistency and have the lowest modification cost, thus avoiding blind correction.

[0043] The application fuses in-view scoring, time sequence consistency and cross-region / event / behavior multi-dimensional consistency, dynamically updates and corrects strategies, realizes the full-process automation closed loop of "detection-scoring-classification-correction-reassessment", and reduces manual intervention.

[0044] The application automatically triggers classification processing strategies according to final consistency scoring, improves the availability of high-quality data, and prevents low-quality data from polluting core systems. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0046] Figure 1 The present application is based on traffic data quality scoring and automatic classification method based on consistency detection. Figure One The flow chart of the present application is shown in the figure.

[0047] Figure 2 The present application is based on traffic data quality scoring and automatic classification method based on consistency detection. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0049] The present application is based on traffic data quality scoring and automatic classification method based on consistency detection. Figure One The present application is based on traffic data quality scoring and automatic classification method based on consistency detection. Figure 1 The present application is based on traffic data quality scoring and automatic classification method based on consistency detection.

[0050] S1: Collect and integrate multi-source traffic data to form data records containing position, event and behavior information; align the time stamps of the multi-source traffic data; and initialize the initial weights of each field based on field statistical characteristics;

[0051] S2: divide the data fields into three types of views, i.e., a location view, an event view, and a behavior view; and calculate the correlations between the fields in each view;

[0052] S3: generate a conflict matrix between the views based on the value deviations of the fields across the views; and based on the conflict matrix and initial weights, solve a first minimal modifiable subset for eliminating the conflicts, and perform a first correction priority ranking of the conflicting fields according to the first minimal modifiable subset;

[0053] S4: calculate a view-intra consistency score of each view based on the correlations between the fields in the view and the conflict matrix; and fuse the view-intra consistency scores of all the views to obtain a view-inter comprehensive consistency score of the current data record;

[0054] S5: perform a first round of automatic quality grading of the data record according to the view-inter comprehensive consistency score, and trigger a first correction action of the data according to the grading result and the first minimal modifiable subset, wherein the correction sequence follows the first correction priority ranking; and re-calculate the conflict matrix after the correction;

[0055] S6: perform time sequence consistency detection on the data after the first correction, and calculate a first consistency comprehensive score;

[0056] S7: further perform cross-regional location consistency detection, event sequence consistency detection, and behavior pattern consistency detection on the data after the first correction, and obtain a second consistency comprehensive score;

[0057] S8: based on the second consistency comprehensive score, dynamically update the correction priority ranking and the minimal modifiable subset of the conflicting fields to obtain a second correction priority ranking and a second minimal modifiable subset; and trigger a second correction action of the data according to the second minimal modifiable subset, wherein the correction sequence follows the second correction priority ranking;

[0058] S9: determine whether a convergence condition is met, wherein the convergence condition includes that there is no conflict in the conflict matrix or the number of iterations reaches a preset upper limit; if the convergence condition is not met, return to S7 to perform a next round of iteration on the data after the second correction and the updated priority ranking and subset; and if the convergence condition is met, calculate a final consistency score based on a weighted sum of the first consistency comprehensive score and the second consistency comprehensive score;

[0059] S10: perform a final automatic quality grading of the data record according to the final consistency score, and trigger a corresponding automatic data governance strategy based on the quality grading result.

[0060] In one embodiment, the multi-source traffic data includes at least one of vehicle satellite positioning trajectory data, vehicle-mounted sensor data, scheduling logs, passenger boarding and alighting records, ticket information, traffic signal states, and road congestion states.

[0061] In one embodiment, in step S1, the step of collecting and integrating multi-source traffic data is included, and the specific execution process is as follows:

[0062] The multi-source data in the urban traffic system is comprehensively collected, different types of data are integrated, the coverage completeness of the data in the three dimensions of space, time and behavior is ensured, and reliable original basis is provided for subsequent multi-view Figure One consistency analysis. In the integration process, each data record is uniquely identified and time-aligned, so that the data of different sources can be corresponded in the same record, thereby forming a complete data matrix that can be used for calculation and analysis. The collected data can be formally represented as:

[0063]

[0064] wherein, is the i-th traffic data record, is the j-th record, is the i-th field of the j-th record, including position, event or behavior information, represents the total number of data, represents the number of fields of each data. In one embodiment, in step S1, in order to facilitate multi-view consistency analysis and field-dependent modeling, the step of field vectorization coding and high-order feature construction of multi-source traffic data is also included, and the specific execution process is as follows:

[0065] Figure One The fields of each data record are vectorized and coded, and the original discrete fields or continuous fields are mapped to a unified vector space:

[0066]

[0067]

[0068] Meanwhile, high-order features are constructed for the potential relationship between field combinations, differences or ratios, so as to enhance the representation ability between fields:

[0069]

[0070] wherein, is the vector representation of the field , for the field that needs to be vectorized and coded, the in the subsequent steps are replaced by the vectorized field features . is the vector dimension, which is set according to the field type, the continuous field =1, the discrete field can be mapped to =5~10 by embedding,​​ For high-level feature constructor, it can be a nonlinear combination or derived field, is related to other fields.

[0071] In one embodiment, in step S1, since the time stamp formats of multi-source data are different, there are offsets, errors or missing, so the step of timestamp standardization and unified indexing of multi-source traffic data is further included, and the specific execution process is:

[0072] This step standardizes all time fields. Different source time stamps are unified into a standard format, and then the times of the same record in different data sources are aligned to generate a unified time index. For fields with missing or abnormal times, linear interpolation or nearest neighbor interpolation is used to complete them to ensure that each field in the subsequent multi-view Figure One consistency analysis can correspond in the time dimension. By minimizing the time deviation between fields from different sources, the accuracy of multi-view Figure One consistency analysis is improved.

[0073] The time alignment formula is represented as:

[0074]

[0075] wherein, , represents the field value of different sources in time , represents the field difference measure, which can be the Euclidean distance or the absolute difference value. By minimizing the time deviation between fields from different sources, the accuracy of multi-view Figure One consistency analysis is improved.

[0076] In one embodiment, the step of initializing the initial weight of each field based on the statistical characteristics of the field in step S1 includes setting the initial weight according to the variance of the field in the historical data, and the value of the variance is inversely proportional to the value of the initial weight. The specific execution process is:

[0077] The importance of each field in consistency evaluation is different. In this step, the field weight is initialized according to the field variance. The larger the variance, the richer the information and the higher the priority:

[0078]

[0079] wherein, represents the variance of the field , is a weight normalization process, so that some fields are not completely ignored.

[0080] In one embodiment, in step S2, in order to realize multi-viewFigure One The consistency analysis divides the data fields into three categories: location view, event view and behavior view, as shown in Figure Two Specifically,

[0081] The location view includes vehicle satellite positioning coordinates, positioning site information and road section identification;

[0082] The event view includes dispatch instructions, fault alarms and operation event records;

[0083] The behavior view includes passenger boarding and alighting behavior, ticket information and vehicle operation behavior.

[0084] After the division, each record corresponds to a subset of fields in each view, providing a basis for conflict detection within and between views. The view division formula is:

[0085]

[0086]

[0087] .

[0088] Among them, , , represent the three view field sets, which can be dynamically adjusted according to the conflict influence, and the field classification can be automatically identified by expert definition or data label. The view division ensures that different types of fields are evaluated separately in subsequent conflict analysis, improving accuracy. The view division result ( , , ) will be used in the steps of view internal field correlation calculation, inter-view conflict matrix generation and view internal consistency scoring. When the location view and the event view conflict, cross-view conflict detection is triggered; the behavior view exception affects the event view, and cross-view conflict is dynamically adjusted.

[0089] In one embodiment, the step of calculating the correlation between fields in each view in step S2 includes: for continuous fields, using Pearson correlation coefficient to calculate the correlation, and for discrete fields, using mutual information metric method to calculate the correlation. The specific execution process is:

[0090] In order to quantify the potential dependency between fields, the field correlation is calculated within each view. For continuous fields, Pearson correlation coefficient is used, and for discrete fields, mutual information is used. Correlation is used to judge the degree of mutual influence of fields in consistency conflict, and the subsequent minimal modification set algorithm will adjust the high-impact fields according to the correlation. The formula is:

[0091]

[0092] in, , indicating field and For discrete fields, the correlation can be determined using mutual information. .

[0093] In one embodiment, in step S3, inconsistencies may exist between cross-view fields, requiring the generation of a conflict matrix. This is used to indicate conflicts between fields. For each pair of cross-view fields, if their values ​​deviate by more than a threshold, they are marked as conflicting. The formula is as follows:

[0094]

[0095] Among them, the conflict matrix The cross-view conflict flag indicates that the view... With View fields and A state of conflict; This indicates an adaptive threshold that can be adjusted adaptively based on the field's standard deviation. , , They are respectively The standard deviation of the field; S5 partitioned view fields in , , S5 partitioned view fields , .

[0096] In one embodiment, to optimize conflict correction, step S3 is based on the conflict matrix. and initial field weights The minimum modifiable subset is found through heuristic search. This means that conflicts can be eliminated by modifying only the fewest fields while ensuring consistency between views.

[0097] The formula for solving MUS is:

[0098]

[0099] in, This represents the subset of fields that were selected for modification. Represents the importance weight of the field. This indicates that the corrected conflict matrix eliminates all conflicts. The MUS algorithm can be solved using heuristic search or integer programming.

[0100] According to the MUS calculation result, the conflict fields are prioritized according to the importance weight and the view weight:

[0101]

[0102] wherein, , , represent the view weight factor, the field weight factor, and the conflict weight factor, and the initial value is set to = = =1 / 3 and can be updated through adaptive learning subsequently. represents the field conflict number statistics of other fields , which are used to calculate the priority of the field in the minimum modifiable set. represents the number of conflicts in which the field participates. The sorting result is used for subsequent automatic correction action priority selection of high-impact fields.

[0103] In one embodiment, the step of fusing the view internal consistency scores of all views in step S4 includes: fusing the view internal consistency scores of all views in a weighted average manner, and the weight is the normalized value of the initial weight of each field in the view. The specific execution process is:

[0104] In the view, the inter-field consistency score of each record is calculated, which is used to measure the integrity and correctness of the data in the same view. The score formula is:

[0105]

[0106] wherein, represents the consistency score of the i th record in the view , is the field correlation, is the conflict matrix, and the denominator is used for normalization to ensure that the score is in [0, 1].

[0107] The multi-view internal scores are integrated into a comprehensive consistency score:

[0108]

[0109] represents the total consistency score of the i th record, is the view weight, and the view weight The initial value is determined by domain experts and can be dynamically updated. The comprehensive score is used for subsequent automatic grading, with a score range of [0, 1], which facilitates setting grading thresholds.

[0110] In one embodiment, step S5 is based on the total consistency score Automatic grading of data. Set the threshold value:

[0111]

[0112] wherein, , , represent the grade threshold, initially set , 75, Grading can trigger different processing actions, such as direct storage, automatic correction, high-risk marking, or manual review, and the weight and scoring formula ensure that grading is scientific and controllable.

[0113] According to the minimum modifiable subset (MUS) and the priority sorting field, data with a grade lower than B triggers automatic correction. Formula:

[0114]

[0115] wherein, represents the corrected field value, represents the difference between the original field and the optimal estimate, is the correction ratio, , which is set to 0.7 by default and is used to control the correction amplitude, and the correction order follows the priority sorting. The correction ratio is associated with the business scenario, for example = 0.7 is suitable for high-precision positioning data, = 0.5 is suitable for passenger behavior data, and the adjustment rule is increased:

[0116]

[0117] Recalculate the conflict matrix for the corrected field:

[0118]

[0119] wherein, represents the corrected conflict matrix, if there is still a conflict, return to S3 to update the conflict matrix, and repeat steps S3-S5 iterations, with iterations to ensure convergence. This step ensures that data consistency is maximized.

[0120] In this embodiment, step S5 also includes a step of detecting and removing outliers from the field after consistency correction:

[0121] Outliers can be flagged or removed to improve data quality and reliability. Outliers can be identified using the standard deviation method or box plot method.

[0122]

[0123] , Represents the mean and standard deviation of the field. The sensitivity coefficient is adjustable; the default value is 3, corresponding to 3 in statistics. in principle. This indicates that a field is missing or invalid in the current record (i.e., an outlier marker). In similarity and consistency calculations, encountering... It will skip or assign a missing penalty value, without being confused with the value 0.

[0124] In this embodiment, step S5 further includes the step of readjusting the field weights based on the correction and anomaly removal results:

[0125]

[0126] in, The weighted smoothing factor is set to 0.7~0.9. This represents the error rate or anomaly rate of the field. Dynamic weights can improve the efficiency of subsequent conflict correction and scoring accuracy.

[0127] In one embodiment, the time series consistency detection step in step S6 includes:

[0128] Perform cross-time consistency checks on time-series data such as vehicle trajectories and passenger behavior:

[0129]

[0130] in, To standardize timestamps and ensure a consistent timeline across multiple data sources, It is a time series function, reflecting the continuous state of a field as it evolves over time. In this formula, Representing the Data records in timestamps The field values ​​are field characteristics after S2 vectorization. The data type depends on the view type, and the data type setting rules are as follows:

[0131] Location view: latitude and longitude coordinates, velocity, and direction angle;

[0132] Event View: Dispatch event codes, fault status

[0133] Behavioral view: Passenger flow, operational behavior coding.

[0134] For the number of time points, represents the similarity function, which is selected according to the field type:

[0135] Continuous fields (such as position, velocity): Euclidean distance (normalized) or cosine similarity.

[0136]

[0137] Discrete fields (such as event encoding): Jaccard similarity or Hamming distance.

[0138]

[0139] represents the time series consistency score, The higher the value, the stronger the time dimension consistency. Low score triggers anomaly value detection or field weight adjustment.

[0140] Combine time consistency and view Figure One consistency score to get comprehensive score:

[0141]

[0142] wherein, is the final score, is the multi-view comprehensive score, is the time series consistency score, is the weight coefficient, , used to balance view and time consistency, set . The comprehensive score is used for final automatic grading and correction strategy decision, ensuring comprehensive and scientific data quality evaluation.

[0143] In this embodiment, the data quality grading can be triggered according to the final score. The grading trigger strategy is:

[0144] According to the final score , the data is graded :

[0145]

[0146] and triggers different actions :

[0147]

[0148] Different levels trigger different automated processing, improving data governance efficiency.

[0149] In one embodiment, in step S7:

[0150] Cross-regional location consistency detection includes determining whether the geographical areas where a vehicle is located within adjacent time periods satisfy the road network topology connectivity requirement. The specific execution process is as follows:

[0151] Consistency checks are performed on vehicle location information across different roads or areas to ensure the rationality of cross-regional trajectories. A location continuity score is calculated for each record.

[0152]

[0153] in, No. Record number in The location vector of each region To record the number of cross-regional transactions, This represents the standard deviation of the location distance, used for normalization.

[0154] Event sequence consistency checking includes verifying whether the scheduling events and the actual vehicle behavior are logically consistent in time sequence. The specific execution process is as follows:

[0155] Perform chronological and logical consistency checks on the scheduling logs and fault events recorded in the event view. Calculate the event sequence consistency score:

[0156]

[0157] in, Representing the Record in time Event encoding, This is an indicator function that returns 1 if the condition is true, and 0 otherwise. This represents the number of time points in the event.

[0158] Behavioral pattern consistency detection includes comparing the similarity between current passenger boarding and alighting behavior and historical travel models or typical route patterns. The specific execution process is as follows:

[0159] Pattern consistency analysis was performed on passenger behavior and vehicle operation behavior. Behavior vectors. Based on higher-order features Construction, i.e. And calculate behavioral similarity:

[0160]

[0161] in, Representing the behavioral feature vectors of records For reference, the number of behavioral samples, This represents cosine similarity or Euclidean distance normalization.

[0162] In one embodiment, the specific implementation process of step S8 is as follows:

[0163] Combining the region position, event and behavior consistency score of the output , , Dynamically adjusting the priority of the cross-view conflict field:

[0164]

[0165] wherein, represents the weighted multi-view Figure One consistency score, = 0.5, = 0.3, = 0.2, which is used to adjust the original priority and the proportion of the view Figure One consistency impact. The dynamic priority is used for automatic correction in the next step, which can improve the correction efficiency and reduce the iteration number.

[0166] Based on the adjusted priority, the MUS subset is recalculated:

[0167]

[0168] wherein, represents the dynamically adjusted minimum modifiable field set, is the dynamic cross-view conflict priority. The dynamic MUS ensures optimal conflict correction and minimum field modification.

[0169] In one embodiment, step S9 is iterated until all conflicts are eliminated or the iteration number reaches the upper limit:

[0170]

[0171] wherein, is the corrected conflict matrix, which represents the cross-view conflict state of the field and , and the iteration number upper limit is times, which is used as an engineering bottom-up strategy to ensure algorithm convergence. The iteration number upper limit includes the full process loop of S7-S8.

[0172] is the full conflict statistical value. If it is 0, it means that all field conflicts have been eliminated. If ture or the iteration number reaches 5 times, the final consistency score is calculated; if False, a new round of cross-region detection is performed.

[0173] The convergence is determined by the corrected conflict matrix , and if =0 or If convergence is determined, the iteration is terminated after T=5 times, and after convergence, the final score and ranking are entered.

[0174] In this embodiment, the comprehensive view and time score, and the cross-region, event, and behavior score are combined to obtain a final consistency score:

[0175]

[0176] For The weighted comprehensive result of the cross-region / event / behavior score is normalized by the weight sum for final ranking determination.

[0177] In this embodiment, the steps of normalizing and fusing the quality of multi-source data are also included:

[0178] The quality scores of data from different sources are normalized and fused according to the reliability of the sources:

[0179]

[0180] wherein, is the score of the th data source , is the data source reliability weight, which is dynamically updated based on the historical correction success rate, and the fused score is used for global consistency judgment.

[0181] In one embodiment, the step S10 of automatically ranking according to the final consistency score includes:

[0182]

[0183] The ranking triggers corresponding processing actions, such as direct storage, recording of exceptions, automatic correction and manual review or rejection, to ensure data governance automation and controllability.

[0184] In one embodiment, the automatic data governance strategy is triggered according to the final quality ranking, including:

[0185] According to the final automatic quality ranking, field correction historical data, and conflict field statistical data, a structured report is generated;

[0186] According to the level of the final quality ranking, one of the following automatic data governance strategies is matched: automatic correction and recording, marking of exceptions and review, and rejection of data storage.

[0187] In this embodiment, the steps of generating a data quality report include:

[0188] Generate structured report according to final classification and field correction and conflict statistics:

[0189] Quality overview: comprehensive score (A / B / C / D ratio)

[0190] Multi-dimensional analysis: conflict heat map (according to view / field statistics ), time consistency trend ( );

[0191] Correction traceability: automatic correction record (field value before and after correction, gain ), abnormal field list;

[0192] Governance suggestions: dynamic weight adjustment , conflict pattern optimization ( ).

[0193] The report output is in JSON / PDF format, containing interactive visualization such as charts and version backtracking interface.

[0194] Among them, the steps of multi-dimensional view conflict statistical analysis include:

[0195] Count the number of conflicts of each field in different views:

[0196]

[0197] Among them, represents the total number of field conflicts, which is used to identify high-risk fields and optimize field weights.

[0198] In this embodiment, the steps of data correction record and version management include:

[0199] Record each correction operation to form data version control, and use historical records for traceability to facilitate analysis and optimization of correction effect.

[0200]

[0201] Among them, is the version iteration sequence, and is the total number of iterations, with a default upper limit of 5 times. is the th data record in version + 1 state, indicates the version iteration number. represents the value of the corrected field , which comes from the output of the first correction or the second automatic correction. ​

[0202] In this embodiment, the abnormal field is automatically labeled. For the fields that still have abnormalities or high conflict probability after correction, automatic labeling is performed for manual review:

[0203]

[0204] wherein, is the abnormal labeling threshold, is set to 0.65. 1 indicates that the field needs manual review, combined with automatic correction, to improve the overall data governance efficiency.

[0205] In one embodiment, it also includes dynamically adjusting the field weight, view weight and threshold based on historical correction records, conflict frequency and grading effect, and adaptively updating to ensure long-term stability and accuracy of the system:

[0206]

[0207] wherein, is the weight of the view in the T th iteration, represents the weight of the field f in the T th iteration, is the learning rate (default setting is 0.1). Based on the field error rate and conflict participation rate. is the adjustment amount calculated based on the historical conflict decline rate and correction effect, representing the average correction gain of the view :

[0208]

[0209] Set the upper limit of adaptive iteration, increase the rules in system learning, and dynamically optimize the iteration depth combined with the correction gain :

[0210] .

[0211] After completing automatic correction and iteration convergence, the correction effect of each record is evaluated. By comparing the consistency scores before and after correction, the correction gain is calculated:

[0212]

[0213] wherein, is the correction gain, is the final consistency score after correction, is the original uncorrected score, indicates that the correction is effective. ​

[0214] In one embodiment, it also includes cross-view Figure One Steps in learning consistent patterns:

[0215] By learning consistency patterns across fields using historical data and corrected records, a predictive model is built.

[0216]

[0217] in, Indicates and Other related field sets, For the prediction function, regression, decision trees, or neural networks can be used to provide reference values ​​for automatic correction in the next round, improving accuracy.

[0218] The steps of the above embodiments are integrated into a closed-loop optimization system, which ensures the system's adaptability and long-term stability. Each iteration adjusts the gain accordingly. Dynamically update weights, thresholds, and priorities:

[0219]

[0220] in, This includes field weights, view weights, and MUS priority. The range of adjustment for the learning rate control. Calculated based on the effects of this round of corrections.

[0221] A multi-level data report is generated based on the final score, classification, correction records, and anomaly field statistics.

[0222]

[0223] The report can be output as charts or visualizations, making it easier for management to monitor overall data quality and providing decision-making references and optimization suggestions.

[0224] In this embodiment, the tiered triggering of the data governance strategy ensures data quality levels and automates corresponding governance actions. Based on the final tiering results... Different strategies are triggered, including:

[0225]

[0226] in, Auto-Correct + Record : Invokes auto-correction, automatically corrects and records the changes to the repository;

[0227] Flag + Review Referencing anomaly annotation rules, marking anomalies and manually reviewing them;

[0228] Reject + Alert : Trigger an alarm mechanism, reject data, and record it in the conflict statistics report.

[0229] It should be noted that:

[0230] A / B Level governance: due to high data quality, only light correction is needed to support real-time business, such as vehicle scheduling;

[0231] C Level governance: medium-quality data needs manual review to avoid introducing new errors by automatic correction;

[0232] D Level governance: low-quality data is usually accompanied by high conflicts, such as contradictions between position and event views, and rejection can prevent pollution of the core library.

[0233] In one embodiment, it also includes: according to the historical data correction effect, conflict frequency, hierarchical distribution for long-term optimization, form adaptive model:

[0234]

[0235] Among them, Threshold, weight and priority parameters, Learning rate. Adaptive optimization ensures that the system is stable and accurate in the long run, and can continuously improve data quality scores and hierarchical accuracy.

[0236] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0237] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for traffic data quality scoring and automatic classification based on view consistency detection, characterized in that, Includes the following steps: S1: Collect and integrate multi-source traffic data to form data records containing location, event, and behavior information; align the multi-source traffic data with timestamps; and initialize the initial weights of each field based on the statistical characteristics of the fields. S2: Divide the data fields into three types of views: location view, event view, and behavior view; Calculate the correlation between fields within each view; S3: Generate a conflict matrix between views based on cross-view field value deviations; Based on the conflict matrix and initial weights, the first minimum modifiable subset for eliminating conflicts is solved, and the conflict fields are sorted according to the first priority of correction. S4: Based on the correlation of fields within each view and the conflict matrix, calculate the intra-view consistency score for each view; merge the intra-view consistency scores of all views to obtain the overall inter-view consistency score for the current data record. S5: Perform the first round of automatic quality classification on the data records based on the overall consistency score between views, and trigger a correction action on the data based on the classification results and the first minimum modifiable subset, with the correction order following the first correction priority. The conflict matrix was recalculated and updated after correction. S6: Perform time series consistency checks on the corrected data and calculate the first consistency comprehensive score; S7: Further cross-regional location consistency detection, event sequence consistency detection, and behavior pattern consistency detection are performed on the corrected data to obtain a second consistency comprehensive score; S8: Based on the second consistency comprehensive score, dynamically update the correction priority ranking and minimum modifiable subset of the conflicting fields to obtain the second correction priority ranking and the second minimum modifiable subset; The secondary correction action is triggered based on the second minimum modifiable subset of data, and the correction order follows the second correction priority order. S9: Determine whether the convergence conditions are met. The convergence conditions include: no conflict in the conflict matrix or the number of iterations reaches the preset upper limit. If not met, return to S7 and perform the next round of iteration with the data after secondary correction and the updated priority and subset. If met, calculate the final consistency score based on the weighted sum of the first consistency comprehensive score and the second consistency comprehensive score. S10: Based on the final consistency score, perform final automatic quality classification on the data records, and trigger the corresponding automated data governance strategy based on the quality classification results; In steps S3 and S8, the first minimum modifiable subset and the second minimum modifiable subset are solved using a heuristic search algorithm. The heuristic search algorithm adds the candidate correction set in order of priority: from the initial weight of the field to the largest, and from the number of times the conflict occurs in the conflict matrix to the smallest.

2. The traffic data quality scoring and automatic grading method based on view consistency detection according to claim 1, characterized in that, The multi-source traffic data includes at least one of the following: vehicle satellite positioning trajectory data, vehicle sensor data, dispatch logs, passenger boarding and alighting records, ticketing information, traffic signal status, and road congestion status.

3. The traffic data quality scoring and automatic grading method based on view consistency detection according to claim 1, characterized in that, The step S1, which initializes the initial weights of each field based on the statistical characteristics of the field, includes setting the initial weights according to the variance of the field in the historical data, wherein the variance value is inversely proportional to the initial weight value.

4. The traffic data quality scoring and automatic grading method based on view consistency detection according to claim 1, characterized in that, In step S2: The location view includes vehicle satellite positioning coordinates, positioning station information, and road segment markings; The event view includes dispatch instructions, fault alarms, and operational event records; The behavioral view includes passenger boarding and alighting behavior, ticketing information, and vehicle operation behavior.

5. The traffic data quality scoring and automatic grading method based on view consistency detection according to claim 1, characterized in that, Step S2, which calculates the correlation between fields within each view, includes: using the Pearson correlation coefficient to calculate the correlation for continuous fields, and using the mutual information metric to calculate the correlation for discrete fields.

6. The traffic data quality scoring and automatic grading method based on view consistency detection according to claim 1, characterized in that, Step S4, which involves fusing the intra-view consistency scores of all views, includes: using a weighted average method to fusing the intra-view consistency scores of all views, with the weights being the normalized values ​​of the initial weights of the fields in each view.

7. The traffic data quality scoring and automatic grading method based on view consistency detection according to claim 1, characterized in that, The time series consistency detection steps in step S6 include: Perform cross-time consistency checks on time series data: in, For standardized timestamps, For the first Data records are in timestamps field values, For the number of time points, This represents the similarity function.

8. The traffic data quality scoring and automatic grading method based on view consistency detection according to claim 1, characterized in that, In step S7: Cross-regional location consistency detection includes: determining whether the geographical areas where a vehicle is located within adjacent time periods satisfy the road network topology connectivity; The event sequence consistency detection includes: verifying whether the scheduling event and the actual vehicle behavior are logically consistent in time sequence; The behavioral pattern consistency detection includes comparing the similarity between the current passenger boarding and alighting behavior and historical travel models or typical route patterns.

9. The traffic data quality scoring and automatic grading method based on view consistency detection according to claim 1, characterized in that, The automated data governance strategy is triggered based on the final quality rating, including: A structured report is generated based on the final automatic quality rating, historical data of field corrections, and statistical data of conflicting fields; Based on the final quality rating, match one of the following automated data governance strategies: automatically correct and record, mark anomalies and review, or refuse to store data.

Citation Information

Patent Citations

  • Highway multi-source traffic data grading and classifying processing system

    CN121122025A