Real-time quality monitoring and abnormal data labeling method and system for road inspection data
By cleaning and labeling data in multiple dimensions, the problems of disordered format and strong concealment of outliers in highway inspection data have been solved, achieving data standardization and accurate labeling, and improving the efficiency of data quality monitoring and anomaly handling.
Patent Information
- Application Number
- CN202511525870.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2025-11-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for highway inspection data preprocessing and quality assessment suffer from problems such as chaotic data formats, uneven distribution of missing data segments, and strong concealment of outliers. This results in a lack of structure and standardization in data quality monitoring and anomaly labeling, making it difficult to meet the requirements for the integrity and standardization of highway inspection data.
Through multi-dimensional data cleaning, construction of a data status discrimination matrix, quantitative evaluation, and structured annotation, including data format unification, missing value filling, outlier removal, construction of a quality rule base based on business logic, and mapping of historical anomaly levels, the data is standardized, complete, and accurately annotated.
It significantly improves the efficiency of quality monitoring and the accuracy of anomaly labeling for highway inspection data, generates high-quality standard datasets, ensures the accuracy of data quality assessment and the structured nature of anomaly labeling, and meets the needs of real-time monitoring and efficient anomaly handling for highway inspection data.
Smart Images

Figure CN120995361A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for real-time monitoring of the quality of highway inspection data and annotation of abnormal data. Background Technology
[0002] Existing technologies have significant shortcomings in the preprocessing of highway inspection data. They do not conduct systematic multi-dimensional cleaning of the raw inspection data, but only perform simple format adjustments or fill in missing values. They have not developed specific processing procedures for problems such as chaotic data formats, uneven distribution of missing data segments, and strong concealment of outliers. As a result, the processed data still contains residual format differences, missing key information, and invalid outliers. It is impossible to form a unified standard analysis dataset, which leads to poor reliability of the basic data provided for subsequent data quality monitoring. This directly affects the accuracy of quality judgment results and makes it difficult to meet the requirements of completeness and standardization of highway inspection data.
[0003] Existing technologies have significant shortcomings in the quality assessment and anomaly labeling stages of highway inspection data. During quality assessment, a business logic-based quality rule base is not constructed; data quality is judged solely based on a single dimension or empirical threshold. Quantitative analysis of data integrity and accuracy is not performed, nor is a comprehensive data quality score calculated using nonlinear fusion algorithms, making it impossible to accurately quantify data quality levels. In the anomaly labeling stage, mapping rules are not established based on the anomaly levels of historical inspection data. Identified anomalies are simply marked without associating anomaly markers with specific fields in standard data or conducting semantic consistency checks. This results in a lack of structured and standardized anomaly labeling, making it impossible to accurately locate the position and type of anomalies. Subsequent data repair and analysis are inefficient, failing to meet the needs of real-time monitoring and efficient anomaly handling for highway inspection data. Summary of the Invention
[0004] This invention provides a method and system for real-time monitoring of highway inspection data quality and annotation of abnormal data, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, this invention provides a method for real-time quality monitoring and anomaly labeling of highway inspection data, comprising:
[0006] S1. Perform multi-dimensional cleaning on the highway inspection data to obtain the standard data of the highway inspection data;
[0007] S2. Based on preset data quality dimensions, comprehensively judge the real-time status of the standard data to obtain the integrity status and accuracy status of the standard data;
[0008] S3. Based on the quality rule base of the highway inspection data, quantitatively evaluate the integrity status and the accuracy status to obtain the quality evaluation result of the highway inspection data.
[0009] S4. Compare the quality assessment results with the abnormal data segments in the historical dynamic thresholds in multiple dimensions to obtain the abnormal identifiers of the highway inspection data.
[0010] S5. Based on the anomaly level of historical highway inspection data, map the anomaly identifier to the standard data to obtain structured anomaly labeling data of the highway inspection data.
[0011] In a preferred embodiment, the step of performing multi-dimensional cleaning on the highway inspection data to obtain standard data of the highway inspection data includes:
[0012] The data format of highway inspection data is standardized to obtain the standardized data of highway inspection data.
[0013] Identify missing values in the standardized data to obtain missing data identifiers for the standardized data;
[0014] Fill in the missing data segments in the missing identifier data to obtain the complete data of the highway inspection data;
[0015] By removing outliers from the complete data, standard data for the highway inspection data is obtained.
[0016] In a preferred embodiment, the step of comprehensively judging the real-time status of the standard data based on a preset data quality dimension to obtain the integrity status and accuracy status of the standard data includes:
[0017] Multimodal evolution is performed on the preset data quality dimensions to construct the data state discrimination matrix of the standard data;
[0018] Based on the integrity dimension in the data state discrimination matrix, the completeness of fields and the continuity of sequences in the standard data are checked one by one to obtain the initial integrity judgment state of the standard data.
[0019] Based on the accuracy dimension in the data state discrimination matrix, the value range compliance and logical consistency of the standard data are checked at multiple levels to obtain the initial accuracy judgment state of the standard data.
[0020] The initial integrity judgment state and the initial accuracy judgment state are input into the data state discrimination matrix to obtain the integrity state and accuracy state of the standard data.
[0021] In a preferred embodiment, the step of performing multimodal evolution on a preset data quality dimension to construct a data state discrimination matrix for the standard data includes:
[0022] Modal analysis is performed on the preset data quality dimensions to obtain the quality modal factors of the standard data;
[0023] Based on the business logic of the highway inspection data, establish dynamic association rules for the quality modal factors;
[0024] Based on the dynamic association rules, the quality modality factors are reconstructed to obtain the dimensional association topology graph of the standard data;
[0025] Based on the association relationships in the dimensional association topology graph, the influence degree of the quality modality factor is subjected to multimodal evolution to obtain the data state discrimination matrix of the standard data.
[0026] In a preferred embodiment, the quality rule base based on the highway inspection data is used to quantitatively evaluate the integrity status and the accuracy status to obtain the quality evaluation result of the highway inspection data, including:
[0027] The quality rule base of the highway inspection data is analyzed to obtain the integrity quantification rules and accuracy quantification rules of the highway inspection data.
[0028] Based on the integrity quantification rules, the missing data fields of the integrity status are analyzed to obtain the analysis results of the integrity status;
[0029] Based on the accuracy quantification rules, the accuracy status is logically checked to obtain the check result of the accuracy status;
[0030] The analysis results and the verification results are logically integrated to obtain the quality assessment results of the highway inspection data.
[0031] In a preferred embodiment, the logical integration of the analysis results and the verification results to obtain the quality assessment result of the highway inspection data includes:
[0032] Extract the integrity score parameters from the analysis results;
[0033] Analyze the accuracy scoring parameters in the verification results;
[0034] Based on the quality rule base of the highway inspection data, the consistency parameters of the standard data are obtained;
[0035] The integrity scoring parameter and the consistency parameter are fused nonlinearly in multiple dimensions to obtain the comprehensive quality score of the highway inspection data. The calculation formula for the comprehensive quality score is as follows:
[0036] ;
[0037] In the formula, For the overall quality score, As an integrity factor, For accuracy factor, As a consistency factor, The quantified value of the integrity scoring parameter. For the quantification of accuracy scoring parameters, This is the quantized value of the consistency parameter. It is an exponential function. It is an inverse trigonometric function. It is a natural exponential function;
[0038] The comprehensive quality score is mapped to a preset quality level range to obtain the quality assessment result of the highway inspection data.
[0039] In a preferred embodiment, the step of performing a multi-dimensional comparison between the quality assessment result and the abnormal data segments in the historical dynamic threshold to obtain the anomaly identifier of the highway inspection data includes:
[0040] The integrity and accuracy comparison benchmarks of abnormal data segments are extracted from historical dynamic thresholds.
[0041] Based on the integrity comparison benchmark, the integrity component in the quality assessment result is verified for consistency to obtain the integrity compliance of the quality assessment result;
[0042] The accuracy component in the quality assessment result is compared with the accuracy comparison benchmark by a deviation analysis to obtain the accuracy compliance of the quality assessment result.
[0043] The integrity compliance and accuracy compliance are logically combined and judged to obtain the anomaly identifier of the highway inspection data.
[0044] In a preferred embodiment, the step of mapping the anomaly identifier to the standard data based on the anomaly level of historical highway inspection data to obtain structured anomaly labeling data of the highway inspection data includes:
[0045] Based on the anomaly levels of the historical highway inspection data, an anomaly level mapping rule for the highway inspection data is constructed.
[0046] Based on the anomaly level mapping rule, the anomaly identifier is parsed in a structured manner to obtain a unified annotation symbol for the anomaly identifier;
[0047] The unified annotation symbols are associated with the corresponding data fields of the standard data to obtain the structured anomaly annotation data of the highway inspection data.
[0048] In a preferred embodiment, associating the unified annotation symbol with the corresponding data field of the standard data to obtain the structured anomaly annotation data of the highway inspection data includes:
[0049] Parse the data field structure of the standard data to obtain the field type description and field context information of the standard data;
[0050] Based on the annotation category of the unified annotation symbol, pattern matching is performed on the field type description to obtain a candidate field set for the standard data;
[0051] Based on the field context information, the data fields in the candidate field set are sorted by priority to obtain the target data fields of the standard data;
[0052] The unified annotation symbols are injected into the preset metadata location of the target data field to obtain the preliminary structured anomaly annotation data of the standard data;
[0053] The semantic consistency of the preliminary structured anomaly annotation data is verified to obtain the structured anomaly annotation data of the highway inspection data.
[0054] To address the aforementioned problems, this invention also provides a real-time monitoring and anomaly labeling system for highway inspection data, the system comprising:
[0055] The highway inspection data standardization and cleaning module is used to clean highway inspection data in multiple dimensions to obtain standard data of the highway inspection data.
[0056] The real-time data quality assessment module is used to comprehensively assess the real-time status of the standard data based on preset data quality dimensions, and to obtain the integrity status and accuracy status of the standard data.
[0057] The data quality quantitative assessment module is used to quantitatively assess the integrity status and the accuracy status based on the quality rule base of the highway inspection data, and obtain the quality assessment result of the highway inspection data.
[0058] The abnormal data dynamic identification module is used to compare the quality assessment results with the abnormal data segments in the historical dynamic threshold in multiple dimensions to obtain the abnormal identification of the highway inspection data.
[0059] The structured anomaly labeling and mapping module is used to map the anomaly identifiers to the standard data based on the anomaly levels of historical highway inspection data, thereby obtaining structured anomaly labeling data of the highway inspection data.
[0060] Compared with the prior art, the present invention has the following beneficial effects:
[0061] 1. This invention lays a high-quality foundation for highway inspection data quality monitoring through multi-dimensional data cleaning and precise quality assessment. Highway inspection data undergoes sequential format standardization, missing value imputation, and outlier removal to generate standardized, complete, and accurate data, effectively eliminating problems such as chaotic data formats, missing information, and abnormal interference. A data status assessment matrix is constructed based on preset data quality dimensions. From the integrity dimension, it checks the completeness of fields and the continuity of sequences; from the accuracy dimension, it verifies the compliance of value ranges and logical consistency. It accurately outputs the integrity and accuracy status of the data, comprehensively and objectively reflecting the real-time data quality and providing a clear basis for subsequent quantitative evaluation.
[0062] 2. This invention significantly improves the efficiency of highway inspection data quality monitoring and the accuracy of anomaly labeling by leveraging scientific quantitative assessment and structured annotation. Based on the completeness and accuracy quantification rules of the quality rule base, field analysis and logical verification of data quality status are performed. A comprehensive quality score is calculated using a nonlinear multidimensional fusion formula and mapped to a quality level, achieving precise quantification of data quality. The quality assessment results are compared with anomaly data segments in historical dynamic thresholds from multiple dimensions to generate accurate anomaly identifiers. Mapping rules are then constructed based on historical anomaly levels, parsing the anomaly identifiers into unified annotation symbols and associating them with corresponding fields in standard data. After semantic consistency verification, structured anomaly-annotated data is obtained. This process achieves standardization and precision in data quality monitoring and anomaly labeling, greatly improving the usability of highway inspection data and the efficiency of subsequent analysis. Attached Figure Description
[0063] Figure 1 This is a flowchart illustrating a method for real-time quality monitoring and abnormal data annotation of highway inspection data according to an embodiment of the present invention.
[0064] Figure 2 This is a functional module diagram of a real-time monitoring and anomaly labeling system for highway inspection data provided in an embodiment of the present invention;
[0065] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0066] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0067] This application provides a method for real-time monitoring of highway inspection data quality and anomaly labeling. The execution subject of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for real-time monitoring of highway inspection data quality and anomaly labeling can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0068] Reference Figure 1 The diagram shown is a flowchart illustrating a method for real-time quality monitoring and anomaly labeling of highway inspection data according to an embodiment of the present invention. In this embodiment, the method includes:
[0069] S1. Perform multi-dimensional cleaning on the highway inspection data to obtain the standard data of the highway inspection data;
[0070] In this embodiment of the invention, the step of performing multi-dimensional cleaning on the highway inspection data to obtain standard data of the highway inspection data includes:
[0071] The data format of highway inspection data is standardized to obtain the standardized data of highway inspection data.
[0072] Identify missing values in the standardized data to obtain missing data identifiers for the standardized data;
[0073] Fill in the missing data segments in the missing identifier data to obtain the complete data of the highway inspection data;
[0074] By removing outliers from the complete data, standard data for the highway inspection data is obtained.
[0075] Specifically, when standardizing the data format of highway inspection data, all highway inspection data to be processed is first collected, and the preset unified data format requirements are clarified. These requirements include the field names of data storage, the field arrangement order, the encoding format of data records, and the data storage type. Then, the format of each piece of highway inspection data is adjusted one by one. Field names that do not meet the preset unified format requirements are replaced with standard field names, the fields are rearranged according to the preset order, data with different encoding formats are converted to the preset encoding format, and data with different storage types are adjusted to the preset storage type. After the format adjustment of all highway inspection data is completed, all data that meets the unified format requirements is the standardized data of highway inspection data.
[0076] Furthermore, when identifying missing values in the standardized data, first determine the range of normal data content and the data existence format that each field in the standardized data should contain. Then, check each data record in the standardized data field by field traversal to see if each field in each data record has empty data content, unidentifiable data content, or data content that does not conform to the normal existence format of the field. If a field in a data record has any of the above situations, mark the field as missing and record the field name, data record number, and location information of the missing state. Compile and summarize all the marked missing state information to form the missing identification data of the standardized data, which contains all missing state markers and corresponding detailed information.
[0077] Furthermore, when filling in the missing data segments in the missing data identifier data, first extract the detailed information corresponding to all missing data segments recorded in the missing data identifier data, including the missing field name, the data record number to which it belongs, and the missing location. For each missing data segment, fill is performed according to the type of its field and the data pattern of that field in other complete data records. If the missing field is a highway mileage location field, and the field shows a continuous increasing or decreasing pattern in adjacent data records, then the interval between adjacent values is calculated based on the values of that field in two adjacent complete data records, and the value corresponding to the missing data segment is determined and filled according to this interval pattern. If the missing field is a highway defect type field, and the field has a repeated type in other data records of the same road segment, then the defect type with the highest frequency in that road segment is selected as the filling content for the missing data segment. After completing the filling of all missing data segments, the highway inspection data that no longer contains missing data is the complete highway inspection data.
[0078] Furthermore, when removing outliers from the complete data, the normal data range for each field in the complete data is first determined based on the actual highway inspection scenarios and historical normal inspection data corresponding to the highway inspection data. This normal data range must cover all reasonable data values that the field may have in actual inspections. Then, the field values of each data record in the complete data are checked one by one. The value of each field in each data record is compared with the normal data range of that field. If the value of a field in a data record exceeds the normal data range of that field, the value is determined to be an outlier. The entire data record containing the outlier is removed from the complete data. After completing the outlier checking and removal operations for all data records, all the remaining data that does not contain outliers is the standard data of the highway inspection data.
[0079] In summary, standardizing data formats eliminates format chaos, prevents format differences from affecting subsequent processing, and ensures a smooth workflow; identifying missing data marks the missing data, accurately locates the missing location and type, provides a clear basis for subsequent data filling, and avoids blind filling.
[0080] In summary, filling in missing data segments yields complete data; supplementing missing content according to patterns eliminates information gaps and meets the completeness requirements of quality assessment; removing outlier standard data filters out invalid interference and outputs high-quality data, ensuring objective quality assessment and accurate anomaly labeling.
[0081] S2. Based on preset data quality dimensions, comprehensively judge the real-time status of the standard data to obtain the integrity status and accuracy status of the standard data;
[0082] In this embodiment of the invention, the step of comprehensively judging the real-time status of the standard data based on a preset data quality dimension to obtain the integrity status and accuracy status of the standard data includes:
[0083] Multimodal evolution is performed on the preset data quality dimensions to construct the data state discrimination matrix of the standard data;
[0084] Based on the integrity dimension in the data state discrimination matrix, the completeness of fields and the continuity of sequences in the standard data are checked one by one to obtain the initial integrity judgment state of the standard data.
[0085] Based on the accuracy dimension in the data state discrimination matrix, the value range compliance and logical consistency of the standard data are checked at multiple levels to obtain the initial accuracy judgment state of the standard data.
[0086] The initial integrity judgment state and the initial accuracy judgment state are input into the data state discrimination matrix to obtain the integrity state and accuracy state of the standard data.
[0087] The step of performing multimodal evolution on preset data quality dimensions and constructing the data state discrimination matrix of the standard data includes:
[0088] Modal analysis is performed on the preset data quality dimensions to obtain the quality modal factors of the standard data;
[0089] Based on the business logic of the highway inspection data, establish dynamic association rules for the quality modal factors;
[0090] Based on the dynamic association rules, the quality modality factors are reconstructed to obtain the dimensional association topology graph of the standard data;
[0091] Based on the association relationships in the dimensional association topology graph, the influence degree of the quality modality factor is subjected to multimodal evolution to obtain the data state discrimination matrix of the standard data.
[0092] Specifically, when constructing a data state discrimination matrix for standard data by performing multimodal evolution on preset data quality dimensions, it is first clarified that the preset data quality dimensions include completeness and accuracy. The completeness dimension encompasses two sub-items: field completeness and sequence continuity. The accuracy dimension encompasses two sub-items: value range compliance and logical consistency. Then, for each dimension and sub-item, its corresponding discrimination criteria are determined. For example, the discrimination criterion for field completeness under the completeness dimension is that each data record must contain all preset mandatory fields; the discrimination criterion for sequence continuity is that the time or spatial sequence of the data record must be continuous without breaks. The discrimination criterion for value range compliance under the accuracy dimension is that the value of each field must be within a preset reasonable range; the discrimination criterion for logical consistency is that the numerical relationship between related fields must conform to the actual business logic. These dimensions, sub-items, and corresponding discrimination criteria are organized in matrix form. The rows of the matrix represent data quality dimensions and sub-items, the columns represent discrimination criteria, and each cell contains the specific discrimination requirements for the corresponding sub-item, thereby constructing the data state discrimination matrix for standard data.
[0093] Furthermore, when checking the completeness of fields and the continuity of sequences in the standard data one by one based on the integrity dimension in the data state discrimination matrix to obtain the initial integrity judgment state of the standard data, the discrimination criteria for field completeness and sequence continuity under the integrity dimension are first extracted from the data state discrimination matrix. For the field completeness check, each data record in the standard data is traversed one by one to check whether each record contains all the pre-set required fields in the integrity dimension. If a record is missing any required field, it is marked as incomplete. For the sequence continuity check, the sequence type of the standard data is first determined. If it is a time series, the data records are arranged in chronological order; if it is a spatial series, the data records are arranged in spatial order. Then, the continuity of sequences between two adjacent records is checked.
[0094] For example, the time interval between adjacent records in a time series must meet a preset interval requirement, and the positional interval between adjacent records in a spatial series must meet a preset distance requirement. If the sequence interval between two adjacent records exceeds the preset requirement, the positional sequence is marked as broken.
[0095] Furthermore, after completing the field completeness and sequence continuity checks of all data records, the number of records with incomplete fields and the number of sequence breaks are counted. If the proportion of records with incomplete fields is lower than a preset threshold and the number of sequence breaks is zero, the initial integrity status of the standard data is determined to be complete; if the proportion of records with incomplete fields exceeds the preset threshold or the number of sequence breaks is not zero, the initial integrity status is determined to be incomplete, thus obtaining the initial integrity status of the standard data.
[0096] Furthermore, when performing multi-level verification of the value range compliance and logical consistency in the standard data based on the accuracy dimension in the data state discrimination matrix to obtain the initial accuracy judgment state of the standard data, the discrimination criteria for value range compliance and logical consistency under the accuracy dimension are first extracted from the data state discrimination matrix.
[0097] Furthermore, for multi-level verification of value range compliance, the first level of verification checks the value of each field in the standard data one by one to determine whether it is within the reasonable range preset in the discrimination matrix. If the value exceeds the range, it is marked as a value range abnormality. The second level of verification performs a second verification on the value marked as a value range abnormality to confirm whether the abnormality is caused by data entry error. If it is confirmed to be an entry error, it is corrected and re-checked. If it cannot be corrected, the value range abnormality mark is retained.
[0098] Furthermore, for multi-level verification of logical consistency, the first level of verification selects logically related field combinations in the standard data, such as the road condition level and damage severity fields in highway inspection data, and checks whether the numerical relationship of the fields within the combination conforms to the preset business logic in the discrimination matrix. If it does not conform, it is marked as a logical anomaly. The second level of verification reproduces the business logic of the field combinations marked as logical anomalies to verify whether the anomaly is caused by a special business scenario. If it does not belong to a special scenario, the logical anomaly mark is retained. After completing the multi-level verification of value range compliance and logical consistency, the number of records with value range anomalies and logical anomalies is counted. If the proportion of anomaly records is lower than a preset threshold, the initial accuracy status of the standard data is determined to be accurate; if the proportion of anomaly records exceeds the preset threshold, the initial accuracy status is determined to be inaccurate. This yields the initial accuracy status of the standard data.
[0099] Furthermore, when inputting the initial integrity and accuracy judgment states into the data state discrimination matrix to obtain the integrity and accuracy states of the standard data, a new column is first added to the data state discrimination matrix to fill in the mapping relationship between the initial judgment state and the final state. It is clarified that when the initial integrity judgment state is complete and there are no other supplementary anomalies, the final integrity state is judged as complete; when the initial integrity judgment state is incomplete, the specific reasons for the incompleteness are further confirmed by combining the discrimination criteria of the integrity dimension in the matrix. If only a few non-critical fields are missing and can be supplemented in a reasonable way, the final integrity state can be adjusted to basically complete; otherwise, the incomplete judgment is maintained.
[0100] Furthermore, regarding the accuracy status, when the initial accuracy assessment is accurate and there are no other hidden anomalies, the final accuracy status is determined to be accurate. When the initial accuracy assessment is inaccurate, the cause of the anomaly is analyzed based on the discrimination criteria of the accuracy dimension in the matrix. If it is only a slight numerical deviation in a few fields and does not affect the overall data usability, the final accuracy status can be adjusted to basically accurate; otherwise, the inaccurate assessment is maintained. The previously obtained initial integrity and accuracy assessments are substituted into the mapping relationship in the matrix for determination, thereby obtaining the integrity and accuracy status of the standard data.
[0101] Specifically, when performing modal analysis on preset data quality dimensions to obtain the quality modal factors of standard data, it is first clarified that the preset dimensions include completeness and accuracy dimensions. The completeness dimension is broken down into features such as field coverage and data record continuity, while the accuracy dimension is broken down into features such as numerical range compliance and logical matching between fields. These core features are defined as quality modal factors. For example, field coverage corresponds to the actual number of required fields in each data record, and data record continuity corresponds to its degree of continuity in time or spatial sequence. Thus, the quality modal factors of standard data are obtained.
[0102] Furthermore, when establishing dynamic association rules for quality modal factors based on the business logic of highway inspection data, the core business logic is first sorted out, such as the correspondence between the location of the defect and the mileage marker, and the matching between the type of defect and the degree of defect. Then, the association relationships between quality modal factors are analyzed. For example, the absence of the existence factor of the mileage marker field will lead to the inability to determine the matching degree factor between the location of the defect and the mileage marker. Different defect type factors correspond to different ranges of defect degree values. Based on these associations, rules are formulated to clarify the impact of factor status on the judgment criteria of other associated factors, thereby establishing dynamic association rules for quality modal factors.
[0103] Furthermore, when reconstructing the graph of quality modal factors based on dynamic association rules to obtain the dimensional association topology of standard data, all quality modal factors are treated as graph nodes and labeled with discrimination content. Connections are drawn between the associated nodes according to dynamic association rules, and the thickness of the connection is used to distinguish the degree of association. Association rules are labeled on the connection lines. For example, the degree intervals corresponding to different disease types are labeled on the connection lines between disease types and disease severity numerical ranges. The factor association relationship is presented through nodes, connections, and rule labels, and the graph reconstruction is completed to obtain the dimensional association topology.
[0104] Furthermore, when performing multimodal evolution based on the influence of the correlation relationships in the dimensional correlation topology graph on the quality modal factors to obtain the data state discrimination matrix of the standard data, factor correlation relationships and rules are extracted from the topology graph. The influence of factors on other correlated factors is analyzed. For example, the absence of the mileage marker field will completely block the discrimination of the matching degree between the disease location and the mileage marker, and incorrect disease type will interfere with the discrimination of the numerical range of disease severity. The quality modal factors are used as matrix rows, and the data state discrimination dimensions are used as columns. The discrimination criteria, correlation influence and degree of the corresponding dimension of the factor are filled in the cells, and the correlation rule index is supplemented, thereby constructing the data state discrimination matrix of the standard data.
[0105] In summary, multimodal evolution pre-defines dimensions and constructs a data state discrimination matrix, clarifies the discrimination criteria for each dimension, avoids judgment confusion, and ensures an orderly discrimination process; it checks fields and sequences based on the matrix integrity dimension, accurately captures integrity defects, and solves the problem of one-sided integrity judgment.
[0106] In summary, by using the matrix accuracy dimension to perform multi-level calibration of the value range and logic, hidden defects are identified and the accuracy level of the data is fully reflected; the initial judgment state input matrix is used to obtain the final state, correct the initial judgment deviation, ensure that the results are consistent with reality, and lay the foundation for subsequent quantitative evaluation.
[0107] In summary, modal analysis pre-defines quality modal factors across dimensions, breaking down completeness and accuracy dimensions into specific discriminable features. This avoids the problem of abstract dimensions leading to a lack of clear discriminative objects, and provides a clear factor basis for building association rules.
[0108] In summary, by establishing dynamic factor association rules based on business logic and clarifying factor associations in conjunction with inspection scenarios, we can solve the problem of chaotic associations caused by neglecting business logic and ensure that factor associations meet actual application needs.
[0109] In summary, the dimensional correlation topology graph reconstructed from the graph transforms abstract factor correlations into intuitive graphs, solving the problem of difficult-to-understand correlations and providing a visual basis for subsequent evolution; the data state discrimination matrix derived from the topology graph clarifies the factor discrimination criteria and influences, solves the defect of arbitrary judgment due to the lack of unified discrimination standards, and ensures that the subsequent state discrimination results are standardized and accurate.
[0110] S3. Based on the quality rule base of the highway inspection data, quantitatively evaluate the integrity status and the accuracy status to obtain the quality evaluation result of the highway inspection data.
[0111] In this embodiment of the invention, the quality rule base based on the highway inspection data is used to quantitatively evaluate the integrity status and the accuracy status to obtain the quality evaluation result of the highway inspection data, including:
[0112] The quality rule base of the highway inspection data is analyzed to obtain the integrity quantification rules and accuracy quantification rules of the highway inspection data.
[0113] Based on the integrity quantification rules, the missing data fields of the integrity status are analyzed to obtain the analysis results of the integrity status;
[0114] Based on the accuracy quantification rules, the accuracy status is logically checked to obtain the check result of the accuracy status;
[0115] The analysis results and the verification results are logically integrated to obtain the quality assessment results of the highway inspection data.
[0116] The logical integration of the analysis results and the verification results yields the quality assessment results of the highway inspection data, including:
[0117] Extract the integrity score parameters from the analysis results;
[0118] Analyze the accuracy scoring parameters in the verification results;
[0119] Based on the quality rule base of the highway inspection data, the consistency parameters of the standard data are obtained;
[0120] The integrity scoring parameter, accuracy scoring parameter, and consistency parameter are fused nonlinearly in a multidimensional manner to obtain the comprehensive quality score of the highway inspection data. The formula for calculating the comprehensive quality score is as follows:
[0121] ;
[0122] In the formula, For the overall quality score, As an integrity factor, For accuracy factor, As a consistency factor, The quantified value of the integrity scoring parameter. For the quantification of accuracy scoring parameters, This is the quantized value of the consistency parameter. It is an exponential function. It is an inverse trigonometric function. It is a natural exponential function;
[0123] The comprehensive quality score is mapped to a preset quality level range to obtain the quality assessment result of the highway inspection data.
[0124] Specifically, the quality rule base for highway inspection data is established based on the business needs and logic of highway inspection. First, core rule dimensions, including completeness, accuracy, and consistency, are determined. Then, each dimension is refined and quantitative standards are set. Next, dynamic association rules between dimensions are built, threshold parameters are adjusted, weight factors are assigned, and a comprehensive quality scoring formula is established. Finally, this is implemented in data cleaning, quality assessment, and other stages, and is dynamically iterated and optimized based on application problems, new data, and needs. Specifically, when parsing the quality rule base for highway inspection data to obtain quantitative rules for the completeness and accuracy of highway inspection data, a pre-built quality rule base for highway inspection data is first obtained. This rule base contains various quantitative standards for data completeness and accuracy. All rules related to completeness are selected from the rule base and organized into completeness quantification rules, such as the number of required fields for each data record, the types of fields that can be missing and the maximum missing ratio, and the time or spatial interval requirements for the continuous and unbroken data sequence. At the same time, all rules related to accuracy are selected from the rule base and organized into accuracy quantification rules, such as the range requirements for the values of each field, the conditions that the logical relationships between different fields must meet, and the matching standards between data values and actual inspection scenarios. Thus, the completeness quantification rules and accuracy quantification rules for highway inspection data are obtained.
[0125] Furthermore, when performing missing data field analysis on the integrity status based on integrity quantification rules to obtain the analysis results, the current judgment result of the integrity status is first clarified. Then, according to the field requirements in the integrity quantification rules, missing fields are checked for in each record of the standard data one by one. The types of missing fields and the number of corresponding missing records in all records are counted, and the proportion of each missing field to the total number of records is calculated. At the same time, the data sequence is checked for breaks, and the number of breaks and the corresponding sequence intervals are counted. The statistically obtained missing field types, missing proportions, and number of sequence breaks are compared with the allowable standards in the integrity quantification rules. For example, if the missing proportion of a certain field exceeds the maximum allowed missing proportion of the rule, or the number of sequence breaks exceeds the allowed range of the rule, then the integrity status is marked as non-compliant; if all statistical results meet the rule requirements, then the integrity status is marked as compliant. These comparison results and markings are compiled and summarized to form the analysis results of the integrity status.
[0126] Furthermore, when performing logical verification of data values to obtain the accuracy status verification results based on the accuracy quantification rules, the values of each field in the standard data are checked one by one according to the numerical range requirements in the accuracy quantification rules, and the number of records with values exceeding the range and their corresponding fields are counted. Then, according to the logical relationship requirements in the rules, related field combinations are checked, such as highway mileage markers and disease location fields, disease type and disease severity fields, and the number of records that do not conform to logical relationships is counted. Simultaneously, combined with the scenario matching standards in the rules, the data values are verified to ensure they match the actual environment, road conditions, and other scenarios of the inspected road sections, and the number of mismatched records is counted. The number of records with values exceeding the range, logically inconsistent records, and scenario mismatched records are compared with the allowable standards in the accuracy quantification rules. If the number of a certain type of abnormal record exceeds the rule's allowable range, the accuracy status is marked as non-compliant; if all abnormal record numbers meet the rule requirements, the accuracy status is marked as compliant. These comparison results and markings are compiled and summarized to form the accuracy status verification results.
[0127] Furthermore, when logically integrating the analysis and verification results to obtain the quality assessment results of highway inspection data, the compliance status and non-compliance details of the integrity status in the analysis results are extracted separately, and the compliance status and non-compliance details of the accuracy status in the verification results are also extracted. The compliance status of integrity and accuracy are comprehensively judged. If both meet the corresponding quantitative rule requirements, the highway inspection data quality is initially determined to be up to standard; if either one has a non-compliance, the data quality is initially determined to be substandard. Subsequently, all non-compliance details are integrated, clarifying the quality dimension to which the non-compliance belongs, the corresponding anomaly type, the number of anomaly records, and their proportion. Simultaneously, based on the priority settings in the quality rule base, key non-compliance affecting data usability is marked. The comprehensive judgment results, non-compliance details, and key marked content are compiled into a structured report. The report must clearly present the overall data quality situation and specific problems, thereby obtaining the quality assessment results of the highway inspection data.
[0128] Specifically, when extracting integrity scoring parameters from the analysis results, the integrity status analysis results document is first opened. This document records various specific analysis data obtained based on integrity quantification rules, including the proportion of missing required fields, the number of data sequence breaks, and the effective filling rate of non-required fields—quantitative information related to integrity. Based on the definition of integrity scoring parameters in the highway inspection data quality rule base, core information directly used to calculate the integrity score is selected from this analysis data. For example, the specific values corresponding to the proportion of missing required fields, the statistical results corresponding to the number of data sequence breaks, and the specific proportions corresponding to the effective filling rate of non-required fields are determined as integrity scoring parameters. This information is extracted and recorded one by one, ensuring that the value of each parameter is completely consistent with the original data in the analysis results. This yields the integrity scoring parameters in the analysis results.
[0129] Furthermore, when analyzing the accuracy scoring parameters in the verification results, the verification result document for accuracy status is consulted. This document contains various accuracy-related data generated based on accuracy quantification rules, such as the percentage of records with values outside the range, the number of records with logical inconsistencies in fields, and the frequency of data mismatch with the scenario. Based on the definition of accuracy scoring parameters in the highway inspection data quality rule library, key quantitative data for calculating the accuracy score is selected from the verification result document. For example, the specific values of the percentage of records with values outside the range, the statistical results of the number of records with logical inconsistencies in fields, and the specific values of the frequency of data mismatch with the scenario are determined as accuracy scoring parameters. These data are extracted and confirmed one by one to ensure that the value of each parameter accurately reflects the actual situation in the verification results. Thus, the accuracy scoring parameters in the verification results are obtained through analysis.
[0130] Furthermore, when obtaining the consistency parameters of standard data from the quality rule base of highway inspection data, the quality rule base is first retrieved. Rule entries related to standard data consistency are searched within the rule base. These entries clarify the acquisition method and judgment criteria for consistency parameters. For example, the rules stipulate that consistency parameters must include the data duplication rate at different inspection times for the same road segment, the data consistency of different acquisition devices in the same inspection task, and the matching rate of data format with the preset standard format. According to the rule requirements, corresponding information is extracted from the standard data for calculation. For example, the proportion of identical data content recorded at different inspection times for the same road segment is calculated to obtain the data duplication rate; the proportion of consistent data in the same field recorded by different acquisition devices in the same inspection task is compared to obtain the data consistency; and the proportion of records in the standard data whose format conforms to the preset standard format is checked to obtain the format matching rate. These calculated values are determined as the consistency parameters of the standard data, ensuring that the parameters meet the definition requirements in the quality rule base.
[0131] Furthermore, when obtaining the comprehensive quality score of highway inspection data through nonlinear multidimensional fusion of integrity, accuracy, and consistency parameters, the influence weight of each parameter in the comprehensive quality score is first clarified. This weight is pre-set by the highway inspection data quality rule base. For example, the rules stipulate that the integrity parameter has the highest weight, followed by the accuracy parameter, and the consistency parameter has the lowest weight. According to the weight, the value of each parameter is first converted into a corresponding score. For example, the lower the proportion of missing required fields in the integrity parameter, the higher the corresponding converted score; the lower the proportion of records with out-of-range values, the higher the converted score of the accuracy parameter; the higher the data duplication rate, consistency, and format matching rate, the higher the converted score of the consistency parameter. Subsequently, according to the nonlinear fusion requirements, the converted scores of each parameter are superimposed. During the superposition process, if the score of a certain parameter is lower than a preset threshold, the overall superposition result will be appropriately reduced; if the score of a certain parameter is higher than a preset excellent threshold, the overall superposition result will be appropriately increased. In this way, the parameters of the three dimensions are fused into a unified value, which is the comprehensive quality score of the highway inspection data, ensuring that the score can comprehensively reflect the actual situation of the three parameters.
[0132] Furthermore, when mapping the comprehensive quality score to preset quality level intervals to obtain the quality assessment results of highway inspection data, the preset quality level interval division standard is first obtained. This standard is set by the highway inspection data quality rule base. For example, the comprehensive quality score is divided into multiple consecutive intervals, each interval corresponding to a quality level. The interval with the highest score corresponds to the excellent level, the next highest to the good level, the middle interval to the acceptable level, the lower interval to the level requiring improvement, and the lowest interval to the unacceptable level. The calculated comprehensive quality score is compared with these preset intervals to determine the specific interval to which the score belongs. For example, if the comprehensive quality score is in the highest interval, the corresponding quality level is excellent; if it is in the interval requiring improvement, the corresponding quality level is requiring improvement. The comprehensive quality score and its corresponding quality level are then compiled together to form a structured result containing the score value and the level determination. This result is the quality assessment result of the highway inspection data, ensuring that the assessment result clearly reflects the data quality level.
[0133] Specifically, the integrity factor comes from the highway inspection data quality rule base. This rule base pre-sets the weight ratio of the integrity dimension in the comprehensive quality score. The specific value of the factor is determined according to the degree of influence of the integrity dimension on the overall quality of highway inspection data, so as to ensure that the factor can accurately reflect the importance of the integrity score parameter in the comprehensive score.
[0134] Furthermore, the accuracy factor originates from the highway inspection data quality rule base. Based on the impact of the accuracy dimension on the application value of highway inspection data, the rule base pre-sets the weight ratio of the accuracy dimension in the comprehensive quality score. This weight ratio is the specific value of the accuracy factor, ensuring that the factor can reflect the degree of contribution of the accuracy scoring parameter to the comprehensive score.
[0135] Furthermore, the consistency factor originates from the highway inspection data quality rule base. Based on the close correlation between the consistency dimension and data integrity and accuracy, as well as its impact on the overall data quality, the rule base pre-sets the weight ratio of the consistency dimension in the comprehensive quality score. This weight ratio is the specific value of the consistency factor, ensuring that the factor can reasonably reflect the role of the consistency parameter in the comprehensive score.
[0136] Furthermore, the quantitative value of the integrity scoring parameter comes from the analysis of the integrity component in the quality assessment results. By converting information such as the proportion of missing fields, the number of data sequence breaks, and the inclusion rate of required fields contained in the integrity component into specific values according to the quantitative method set in the highway inspection data quality rule base, this value is the quantitative value of the integrity scoring parameter, ensuring that it can accurately represent the actual level of data integrity.
[0137] Furthermore, the quantitative value of the accuracy scoring parameter comes from the analysis of the accuracy component in the quality assessment results. Information such as the proportion of records with values out of range, the number of records with logical inconsistencies in fields, and the frequency of data mismatch with the scenario included in the accuracy component are converted into specific values according to the quantitative standards in the highway inspection data quality rule library. This value is the quantitative value of the accuracy scoring parameter, ensuring that it can truly reflect the actual situation of data accuracy.
[0138] Furthermore, the quantitative value of the consistency parameter comes from the analysis of the consistency of standard data. According to the consistency calculation method set in the highway inspection data quality rule base, information such as the data duplication rate of different inspection times for the same road segment, the data consistency of different collection devices for the same inspection task, and the matching rate of data format with the preset standard format are converted into specific values. These values are the quantitative values of the consistency parameter, ensuring that the actual level of data consistency can be objectively reflected.
[0139] Furthermore, logarithmic processing is applied to the quantified values of the integrity scoring parameters to ensure that their contribution to the overall quality score increases reasonably as the integrity parameters improve, avoiding large fluctuations in the score due to small changes in the parameters. Arctangent processing is applied to the quantified values of the accuracy scoring parameters to limit the range of influence of the accuracy parameters on the overall score, preventing excessive increases or decreases in the overall score when the accuracy parameters are abnormally high. Exponential decay processing is applied to the quantified values of the consistency parameters to ensure that the consistency parameters have a more significant effect on improving the overall score at lower levels, and that the improvement gradually slows down after reaching higher levels, which conforms to the actual impact of data consistency on overall quality.
[0140] Furthermore, by assigning different weights to the three dimension parameters through completeness factors, accuracy factors, and consistency factors, the comprehensive quality score can reasonably reflect the contribution of each dimension parameter according to the actual impact of each dimension on the quality of highway inspection data, and finally obtain a comprehensive score that can comprehensively and objectively reflect the overall quality of highway inspection data.
[0141] Furthermore, as the quantified value of the integrity scoring parameter increases, its logarithmic result will also increase. With a fixed integrity factor, this part will contribute more to the overall quality score. Therefore, the overall quality score will show an upward trend as the quantified value of the integrity scoring parameter increases, but the rate of increase will gradually slow down because the growth rate of the logarithmic function will decrease as the input value increases, thus avoiding the continuous increase of the integrity parameter leading to an infinitely rapid increase in the score.
[0142] Furthermore, as the quantified value of the accuracy scoring parameter increases, its result after arctangent processing will also increase. However, the range of values for the arctangent function is fixed. No matter how large the quantified value of the accuracy parameter is, the processing result will remain stable within a specific range. Therefore, with the accuracy factor fixed, the contribution of this part to the overall quality score will increase as the quantified value of the accuracy parameter increases, eventually stabilizing. There will be no situation where the score gets out of control due to the infinite increase of the accuracy parameter.
[0143] Furthermore, as the quantified value of the consistency parameter increases, its result after exponential decay will also increase. When the quantified value of the consistency parameter is small, the processing result grows faster, and its contribution to the overall quality score is significantly improved. When the quantified value of the consistency parameter reaches a high level, the growth rate of the processing result will gradually slow down and eventually tend to a fixed value. Therefore, the overall quality score will increase with the increase of the quantified value of the consistency parameter, and the rate of increase will gradually slow down until it tends to stabilize, which is consistent with the actual situation that the impact of data consistency on overall quality weakens after it has been improved to a certain extent.
[0144] In summary, analyzing the completeness and accuracy quantification rules of the quality rule base clarifies evaluation standards and avoids evaluation confusion caused by the lack of unified rules. Based on the completeness quantification rules, missing fields are analyzed to accurately identify completeness issues, ensuring that the analysis results accurately reflect the actual level of data completeness. Accuracy quantification rules are used to verify data logic, identifying numerical and logical anomalies and ensuring that the verification results comprehensively reflect the data's accuracy. Integrating the analysis and verification results yields a quality assessment, comprehensively judging data quality, addressing the problem of single-minded judgments, and laying the foundation for subsequent comparisons.
[0145] In summary, extracting completeness and accuracy scoring parameters clarifies the core data for evaluation, providing a foundation for comprehensive scoring. Consistency parameters are derived from the quality rule base to supplement quality dimensions beyond completeness and accuracy, avoiding biased evaluation. A non-linear, multi-dimensional fusion calculation is used to arrive at a comprehensive quality score, rationally allocating the weights of each parameter through a formula to accurately quantify data quality. The score is mapped to a quality level range, transforming abstract scores into intuitive levels, providing a clear basis for subsequent anomaly comparisons.
[0146] S4. Compare the quality assessment results with the abnormal data segments in the historical dynamic thresholds in multiple dimensions to obtain the abnormal identifiers of the highway inspection data.
[0147] In this embodiment of the invention, the step of performing a multi-dimensional comparison between the quality assessment result and the abnormal data segments in the historical dynamic threshold to obtain the abnormality identifier of the highway inspection data includes:
[0148] The integrity and accuracy comparison benchmarks of abnormal data segments are extracted from historical dynamic thresholds.
[0149] Based on the integrity comparison benchmark, the integrity component in the quality assessment result is verified for consistency to obtain the integrity compliance of the quality assessment result;
[0150] The accuracy component in the quality assessment result is compared with the accuracy comparison benchmark by a deviation analysis to obtain the accuracy compliance of the quality assessment result.
[0151] The integrity compliance and accuracy compliance are logically combined and judged to obtain the anomaly identifier of the highway inspection data.
[0152] Specifically, when extracting the integrity and accuracy comparison benchmarks for abnormal data segments from historical dynamic thresholds, the historical dynamic thresholds are first retrieved. These thresholds are formed based on the quality characteristics of historical highway inspection data and contain various benchmark information used to determine whether the data is abnormal. Benchmark content related to data integrity is then selected, including the upper limit of the proportion of missing fields under normal circumstances in historical data, the maximum allowed number of breaks in the data sequence, and the minimum inclusion rate of required fields. This content is then compiled into the integrity comparison benchmark. Simultaneously, benchmark content related to data accuracy is selected, including the normal fluctuation range of values for each field in historical data, the allowable deviation of logical relationships between fields, and the minimum matching rate between the data and the actual scenario. This content is then compiled into the accuracy comparison benchmark, ensuring that both benchmarks can be directly used for subsequent comparative analysis.
[0153] Furthermore, when verifying the consistency of the integrity components in the quality assessment results based on the integrity comparison benchmark to obtain the integrity compliance of the quality assessment results, the integrity components are first extracted from the quality assessment results. These components include specific information such as the field missing rate, the number of data sequence breaks, and the mandatory field inclusion rate of the current highway inspection data. This information is then compared one by one with the corresponding content in the integrity comparison benchmark. For example, the current field missing rate is compared with the upper limit of the missing rate in the benchmark; the current number of data sequence breaks is compared with the maximum number of breaks in the benchmark; and the current mandatory field inclusion rate is compared with the minimum inclusion rate in the benchmark. The proportion of items that meet the benchmark requirements out of the total number of comparison items is calculated. This proportion represents the integrity compliance of the quality assessment results. If a comparison item fully meets the benchmark requirements, it is considered compliant; if it exceeds the benchmark limit, it is considered non-compliant, ensuring that the compliance accurately reflects the degree of consistency between the integrity components and the benchmark.
[0154] Furthermore, when performing deviation analysis between the accuracy component of the quality assessment results and the accuracy comparison benchmark to obtain the accuracy compliance rate, the accuracy component, which includes the fluctuation of numerical values of each field, logical matching between fields, and data matching with the scenario, is first extracted from the quality assessment results. Then, this information is checked against the accuracy comparison benchmark one by one to check whether the numerical fluctuations are within the normal range, whether the logical relationship deviations are compliant, and whether the data matching with the scenario meets the minimum requirements. Finally, the percentage of checks that meet the benchmark is calculated; this percentage is the accuracy compliance rate. Meeting the benchmark is counted as compliance, and exceeding the range is counted as non-compliance, ensuring that the compliance rate truly reflects the deviation situation.
[0155] Furthermore, when logically synthesizing the integrity and accuracy compliance to obtain anomaly indicators for highway inspection data, the logical rules for judgment are first determined: when both integrity and accuracy compliance reach a preset pass rate, the current highway inspection data is judged to be anomaly-free; when either compliance fails to reach the pass rate, the current data is judged to be anomaly-free. The calculated integrity and accuracy compliance are compared with the preset pass rates. If both reach or exceed the pass rate, an anomaly-free indicator is generated, indicating that the data integrity and accuracy meet the historical dynamic threshold requirements. If one or both compliance rates fail to reach the pass rate, an anomaly indicator is generated, clearly indicating the type of non-compliant compliance, the specific non-compliant comparison or inspection item, and the deviation from the historical dynamic threshold. This ensures that the anomaly indicator clearly reflects the specific location and cause of the data anomaly, thus obtaining the anomaly indicators for the highway inspection data.
[0156] In summary, analyzing historical thresholds provides a benchmark for comparing completeness and accuracy, offering historical references for anomaly detection and preventing arbitrary judgments due to a lack of standards. The completeness compliance verified against the completeness benchmark accurately measures whether data completeness meets standards and reflects anomalies at the completeness level.
[0157] In summary, by comparing the accuracy benchmark with the deviation analysis to obtain the accuracy compliance, accuracy deviations are clearly identified and anomalies at the accuracy level are revealed. Anomaly markers, which comprehensively determine compliance, clarify the anomaly types and causes, resolve the ambiguity of anomaly labeling in the background technology, and lay the foundation for subsequent mapping.
[0158] S5. Based on the anomaly level of historical highway inspection data, map the anomaly identifier to the standard data to obtain structured anomaly labeling data of the highway inspection data.
[0159] In this embodiment of the invention, the step of mapping the anomaly identifier to the standard data based on the anomaly level of historical highway inspection data to obtain structured anomaly labeling data of the highway inspection data includes:
[0160] Based on the anomaly levels of the historical highway inspection data, an anomaly level mapping rule for the highway inspection data is constructed.
[0161] Based on the anomaly level mapping rule, the anomaly identifier is parsed in a structured manner to obtain a unified annotation symbol for the anomaly identifier;
[0162] The unified annotation symbols are associated with the corresponding data fields of the standard data to obtain the structured anomaly annotation data of the highway inspection data.
[0163] The step of associating the unified annotation symbols with the corresponding data fields of the standard data to obtain the structured anomaly annotation data of the highway inspection data includes:
[0164] Parse the data field structure of the standard data to obtain the field type description and field context information of the standard data;
[0165] Based on the annotation category of the unified annotation symbol, pattern matching is performed on the field type description to obtain a candidate field set for the standard data;
[0166] Based on the field context information, the data fields in the candidate field set are sorted by priority to obtain the target data fields of the standard data;
[0167] The unified annotation symbols are injected into the preset metadata location of the target data field to obtain the preliminary structured anomaly annotation data of the standard data;
[0168] The semantic consistency of the preliminary structured anomaly annotation data is verified to obtain the structured anomaly annotation data of the highway inspection data.
[0169] Specifically, when constructing anomaly level mapping rules for highway inspection data based on the anomaly levels of historical highway inspection data, the anomaly level information already marked in the historical data is collected. These levels are divided into categories such as severe anomalies, minor anomalies, and tiny anomalies according to their degree of impact. The anomaly types corresponding to each level are sorted out. For example, severe anomalies include large-scale missing fields, and minor anomalies include missing individual non-critical fields. A unique mapping relationship is set for each level and its corresponding type. The level corresponding to the type to which the anomaly identifier belongs and the key information that must be included are clearly defined and organized into a structured rule document, which is the anomaly level mapping rule.
[0170] Furthermore, when performing structured parsing of anomaly identifiers based on anomaly level mapping rules to obtain unified annotation symbols for anomaly identifiers, the anomaly identifier document is opened, and information such as the quality dimension to which the anomaly belongs, its specific manifestations, and the range of data involved is extracted. The corresponding anomaly type and level are determined by referring to the mapping rules, and the unified annotation symbol provisions corresponding to the level and type are queried in the rules. The symbols are extracted according to the provisions and ensured to be consistent with the rules, thereby obtaining unified annotation symbols.
[0171] Furthermore, when linking the unified annotation symbols to the corresponding data fields of the standard data to obtain the structured anomaly annotation data of the highway inspection data, locate the fields in the standard data related to the anomaly identifier, add annotation columns to these fields and fill in the corresponding unified annotation symbols to ensure that the annotation of each anomaly data record is accurate. Add a structured description at the end of the standard data to indicate the correspondence between the symbols and the anomaly level and type, forming a complete document containing the original data, annotation columns and description, which is the structured anomaly annotation data.
[0172] Specifically, when parsing the data field structure of standard data to obtain the field type description and field context information of standard data, open the file storing the standard data, view all data fields, determine the type description of each field one by one, such as location, disease, time, etc., record the field name and type attribute, and at the same time analyze the relationship between each field and other fields and its role in the data, organize it into field context information, and ensure that the type and context of each field are clearly recorded.
[0173] Furthermore, when performing pattern matching on field type descriptions based on the annotation categories of unified annotation symbols to obtain a candidate field set for standard data, the annotation categories of unified annotation symbols are clearly defined, such as location, disease, or time anomalies. Fields matching the category are selected from the field type descriptions, and a temporary field list is formed. It is checked to ensure that no irrelevant field types are included, thereby obtaining a candidate field set.
[0174] Furthermore, when prioritizing the data fields in the candidate field set based on the field context information to obtain the target data field of the standard data, the context information of the candidate fields is viewed, the field most directly associated with the anomaly identifier is given the highest priority, the core application fields are given an appropriate increase in priority, and the fields are sorted according to the degree of association and importance, and the field ranked first is selected as the target data field.
[0175] Furthermore, when injecting uniform annotation symbols into the preset metadata location of the target data field to obtain preliminary structured anomaly annotation data of the standard data, the metadata area of the target data field is located, and uniform annotation symbols are filled in the preset anomaly annotation location according to the format to ensure that the original data is not affected. The integrity of the data structure and the accuracy of the annotation are checked, thereby obtaining preliminary structured anomaly annotation data.
[0176] Furthermore, when performing semantic consistency verification on the preliminary structured anomaly annotation data to obtain structured anomaly annotation data for highway inspection data, the annotation category is checked to ensure it matches the field type and the anomaly situation matches the field context, according to the standard that the annotation symbol matches the actual anomaly semantics of the target field. Once all annotation items meet the requirements, they are determined to be structured anomaly annotation data. If there are any discrepancies, they are corrected and re-verified.
[0177] In summary, mapping rules are established based on historical anomaly levels, ensuring that anomaly identifiers correspond to clear level standards and avoiding unfounded labeling. Parsing anomaly identifiers according to these rules yields standardized labeling symbols and a unified anomaly marking format, resolving labeling inconsistencies. Data conforming to the association standards and corresponding fields accurately pinpoints anomaly locations, generating structured labeled data that facilitates subsequent data repair and analysis.
[0178] In summary, parsing the type description and context of field structures clarifies field attributes and relationships, providing a basis for subsequent matching. Candidate fields matched by label category are filtered to identify related fields, avoiding label mismatches. Target fields are selected by sorting according to context, locking in core related fields to ensure targeted labeling. Symbol injection yields preliminary labeled data, which is then semantically validated to ensure accurate and compliant labeling, providing reliable labeling results for subsequent data processing.
[0179] like Figure 2 The diagram shown is a functional module diagram of a real-time monitoring and anomaly labeling system for highway inspection data provided in an embodiment of the present invention.
[0180] The real-time monitoring and anomaly labeling system 100 for highway inspection data quality described in this invention can be installed in an electronic device. Depending on the functions implemented, the system 100 may include a highway inspection data standardization and cleaning module 101, a real-time data quality discrimination module 102, a data quality quantitative assessment module 103, an anomaly dynamic identification module 104, and a structured anomaly labeling and mapping module 105. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.
[0181] In this embodiment, the functions of each module / unit are as follows:
[0182] The highway inspection data standardization and cleaning module 101 is used to perform multi-dimensional cleaning on the highway inspection data to obtain the standard data of the highway inspection data.
[0183] The real-time data quality discrimination module 102 is used to comprehensively discern the real-time status of the standard data based on a preset data quality dimension, and obtain the integrity status and accuracy status of the standard data.
[0184] The data quality quantitative assessment module 103 is used to quantitatively assess the integrity status and the accuracy status based on the quality rule base of the highway inspection data, and obtain the quality assessment result of the highway inspection data.
[0185] The abnormal data dynamic identification module 104 is used to compare the quality assessment result with the abnormal data segment in the historical dynamic threshold in multiple dimensions to obtain the abnormal identification of the highway inspection data.
[0186] The structured anomaly labeling and mapping module 105 is used to map the anomaly identifier to the standard data according to the anomaly level of historical highway inspection data, so as to obtain the structured anomaly labeling data of the highway inspection data.
[0187] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0188] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0189] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0190] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0191] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for real-time quality monitoring and anomaly labeling of highway inspection data, characterized in that, The method includes: S1. Perform multi-dimensional cleaning on the highway inspection data to obtain the standard data of the highway inspection data; S2. Based on preset data quality dimensions, comprehensively judge the real-time status of the standard data to obtain the integrity status and accuracy status of the standard data; S3. Based on the quality rule base of the highway inspection data, quantitatively evaluate the integrity status and the accuracy status to obtain the quality evaluation result of the highway inspection data. S4. Compare the quality assessment results with the abnormal data segments in the historical dynamic thresholds in multiple dimensions to obtain the abnormal identifiers of the highway inspection data. S5. Based on the anomaly level of historical highway inspection data, map the anomaly identifier to the standard data to obtain structured anomaly labeling data of the highway inspection data.
2. The method for real-time quality monitoring and abnormal data annotation of highway inspection data as described in claim 1, characterized in that, The process of cleaning the highway inspection data in multiple dimensions to obtain standard data for the highway inspection data includes: The data format of highway inspection data is standardized to obtain the standardized data of highway inspection data. Identify missing values in the standardized data to obtain missing data identifiers for the standardized data; Fill in the missing data segments in the missing identifier data to obtain the complete data of the highway inspection data; By removing outliers from the complete data, standard data for the highway inspection data is obtained.
3. The method for real-time quality monitoring and abnormal data annotation of highway inspection data as described in claim 1, characterized in that, The process of comprehensively judging the real-time status of the standard data based on preset data quality dimensions to obtain the integrity and accuracy status of the standard data includes: Multimodal evolution is performed on the preset data quality dimensions to construct the data state discrimination matrix of the standard data; Based on the integrity dimension in the data state discrimination matrix, the completeness of fields and the continuity of sequences in the standard data are checked one by one to obtain the initial integrity judgment state of the standard data. Based on the accuracy dimension in the data state discrimination matrix, the value range compliance and logical consistency of the standard data are checked at multiple levels to obtain the initial accuracy judgment state of the standard data. The initial integrity judgment state and the initial accuracy judgment state are input into the data state discrimination matrix to obtain the integrity state and accuracy state of the standard data.
4. The method for real-time quality monitoring and abnormal data annotation of highway inspection data as described in claim 3, characterized in that, The step of performing multimodal evolution on preset data quality dimensions and constructing the data state discrimination matrix of the standard data includes: Modal analysis is performed on the preset data quality dimensions to obtain the quality modal factors of the standard data; Based on the business logic of the highway inspection data, establish dynamic association rules for the quality modal factors; Based on the dynamic association rules, the quality modality factors are reconstructed to obtain the dimensional association topology graph of the standard data; Based on the association relationships in the dimensional association topology graph, the influence degree of the quality modality factor is subjected to multimodal evolution to obtain the data state discrimination matrix of the standard data.
5. The method for real-time quality monitoring and abnormal data annotation of highway inspection data as described in claim 1, characterized in that, The quality rule base based on the highway inspection data quantifies and evaluates the integrity status and the accuracy status to obtain the quality evaluation result of the highway inspection data, including: The quality rule base of the highway inspection data is analyzed to obtain the integrity quantification rules and accuracy quantification rules of the highway inspection data. Based on the integrity quantification rules, the missing data fields of the integrity status are analyzed to obtain the analysis results of the integrity status; Based on the accuracy quantification rules, the accuracy status is logically checked to obtain the check result of the accuracy status; The analysis results and the verification results are logically integrated to obtain the quality assessment results of the highway inspection data.
6. The method for real-time quality monitoring and abnormal data annotation of highway inspection data as described in claim 5, characterized in that, The logical integration of the analysis results and the verification results yields the quality assessment results of the highway inspection data, including: Extract the integrity score parameters from the analysis results; Analyze the accuracy scoring parameters in the verification results; Based on the quality rule base of the highway inspection data, the consistency parameters of the standard data are obtained; The integrity scoring parameter, accuracy scoring parameter, and consistency parameter are fused nonlinearly in a multidimensional manner to obtain the comprehensive quality score of the highway inspection data. The formula for calculating the comprehensive quality score is as follows: ; In the formula, For the overall quality score, As an integrity factor, For accuracy factor, As a consistency factor, The quantified value of the integrity scoring parameter. For the quantification of accuracy scoring parameters, This is the quantized value of the consistency parameter. It is an exponential function. It is an inverse trigonometric function. It is a natural exponential function; The comprehensive quality score is mapped to a preset quality level range to obtain the quality assessment result of the highway inspection data.
7. The method for real-time quality monitoring and abnormal data annotation of highway inspection data as described in claim 1, characterized in that, The step of comparing the quality assessment results with abnormal data segments in historical dynamic thresholds in multiple dimensions to obtain anomaly identifiers for the highway inspection data includes: The integrity and accuracy comparison benchmarks of abnormal data segments are extracted from historical dynamic thresholds. Based on the integrity comparison benchmark, the integrity component in the quality assessment result is verified for consistency to obtain the integrity compliance of the quality assessment result; The accuracy component in the quality assessment result is compared with the accuracy comparison benchmark by a deviation analysis to obtain the accuracy compliance of the quality assessment result. The integrity compliance and accuracy compliance are logically combined and judged to obtain the anomaly identifier of the highway inspection data.
8. The method for real-time quality monitoring and abnormal data annotation of highway inspection data as described in claim 1, characterized in that, The step of mapping the anomaly identifiers to the standard data based on the anomaly levels of historical highway inspection data to obtain structured anomaly labeling data for the highway inspection data includes: Based on the anomaly levels of the historical highway inspection data, an anomaly level mapping rule for the highway inspection data is constructed. Based on the anomaly level mapping rule, the anomaly identifier is parsed in a structured manner to obtain a unified annotation symbol for the anomaly identifier; The unified annotation symbols are associated with the corresponding data fields of the standard data to obtain the structured anomaly annotation data of the highway inspection data.
9. A method for real-time quality monitoring and abnormal data annotation of highway inspection data as described in claim 8, characterized in that, The step of associating the unified annotation symbols with the corresponding data fields of the standard data to obtain the structured anomaly annotation data of the highway inspection data includes: Parse the data field structure of the standard data to obtain the field type description and field context information of the standard data; Based on the annotation category of the unified annotation symbol, pattern matching is performed on the field type description to obtain a candidate field set for the standard data; Based on the field context information, the data fields in the candidate field set are sorted by priority to obtain the target data fields of the standard data; The unified annotation symbols are injected into the preset metadata location of the target data field to obtain the preliminary structured anomaly annotation data of the standard data; The semantic consistency of the preliminary structured anomaly annotation data is verified to obtain the structured anomaly annotation data of the highway inspection data.
10. A real-time monitoring system for the quality of highway inspection data and an anomaly data annotation system, characterized in that, The system includes: The highway inspection data standardization and cleaning module is used to clean highway inspection data in multiple dimensions to obtain standard data of the highway inspection data. The real-time data quality assessment module is used to comprehensively assess the real-time status of the standard data based on preset data quality dimensions, and to obtain the integrity status and accuracy status of the standard data. The data quality quantitative assessment module is used to quantitatively assess the integrity status and the accuracy status based on the quality rule base of the highway inspection data, and obtain the quality assessment result of the highway inspection data. The abnormal data dynamic identification module is used to compare the quality assessment results with the abnormal data segments in the historical dynamic threshold in multiple dimensions to obtain the abnormal identification of the highway inspection data. The structured anomaly labeling and mapping module is used to map the anomaly identifiers to the standard data based on the anomaly levels of historical highway inspection data, thereby obtaining structured anomaly labeling data of the highway inspection data.
Citation Information
Cited By
Road marking state real-time monitoring and evaluation method and system based on video recognition
CN121767946A