Traffic engineering material quality data analysis method and system with complex data source

By extracting data items from standardized testing reports, cleaning and processing the data, dividing time intervals for feature calculation, and adopting an integrated evaluation model, the accuracy problem of testing data analysis for transportation engineering materials was solved, and multi-dimensional quality evaluation and data sharing were achieved.

CN115809236BActive Publication Date: 2026-03-31HUBEI TRAFFIC INVESTMENT INTELLIGENT TESTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and accurately analyze material testing data for transportation engineering, resulting in difficulties in achieving precise material quality evaluation.

Method used

By extracting data items based on standardized test reports, cleaning up null, duplicate, and invalid values, merging and labeling data types, dividing time intervals for feature calculation, and using an integrated evaluation model for evaluation, a multi-dimensional quality analysis is formed.

Benefits of technology

It enables rapid and accurate analysis of material testing data in transportation engineering, simplifies the data relationships of complex data sources, improves the accuracy and interpretability of material quality evaluation, and facilitates data sharing and access control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115809236B_ABST
    Figure CN115809236B_ABST
Patent Text Reader

Abstract

The application discloses a traffic engineering material quality data analysis method and system with a complex data source, and relates to the field of data analysis.The method comprises the following steps: based on a normalized detection report, data items related to material quality analysis are extracted, and the attributes of the extracted data items are defined; the defined data items are cleaned of null values, repeated values, synonymous values and invalid values, and are subjected to data type merging and label processing, so that the original data are converted into a calculable data set; an evaluation range is determined according to a quality data analysis target, a time interval is determined according to the time range of the target, the data in the time interval are subjected to feature calculation, and a feature set is formed; the features are classified, an integrated evaluation result is obtained based on an integrated evaluation model, and the data set, the feature set and the integrated evaluation result are output.The application can realize fast and accurate analysis of traffic engineering material detection data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis, specifically to a method and system for analyzing the quality data of transportation engineering materials with complex data sources. Background Technology

[0002] With the rapid development of the transportation industry, the quality management of transportation engineering projects is becoming increasingly important, and material quality is a crucial factor in ensuring project quality. Testing and inspection professionals conduct tests according to a series of standards and specifications, and qualified testing institutions issue test reports to verify the quality indicators of materials, ensuring that the materials used in transportation engineering meet the required quality standards. Data analysis of material testing quality contributes to the intelligent management of transportation engineering quality.

[0003] In terms of materials quality data analysis, material producers, i.e., material manufacturers, need to control product quality when the product leaves the factory. They must test the strength value of the product and calculate the mean and variance of the strength value over a period of time to ensure that the product meets quality requirements. Materials researchers, whose research objective is to study the testing methods of a certain type of material, usually use conventional first-moment and second-moment statistical methods to evaluate material quality. However, the above statistical analysis methods are limited by their respective business scope, research scope, and time range, resulting in limited data scale.

[0004] On the other hand, users of materials and relevant units involved in the construction of transportation engineering projects have conducted research on business standardization and informatization around their business functions, with the main goal of improving business efficiency. They have collected and stored a large amount of material testing data. However, due to the complexity of the transportation testing business process, the wide range of testing parameters covered, and the complexity of factors affecting testing such as people, machines, materials, methods, and environment, the data relationships, data types, and data content in the testing business system are very complex. It is difficult to establish a test and inspection dataset that meets the quality management objectives and it is difficult to quickly obtain accurate and interpretable material quality analysis and evaluation results.

[0005] Therefore, how to quickly and accurately analyze the testing data of transportation engineering materials is an urgent problem to be solved. Summary of the Invention

[0006] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a method and system for analyzing the quality data of transportation engineering materials with complex data sources, which can realize rapid and accurate analysis of the test data of transportation engineering materials.

[0007] To achieve the above objectives, this invention provides a method for analyzing the quality data of transportation engineering materials with complex data sources, specifically including the following steps:

[0008] Based on standardized testing reports, data items related to material quality analysis are extracted, and the attributes of the extracted data items are defined.

[0009] Clean up null, duplicate, synonym, and invalid values ​​in the defined data items, and perform data type merging and tagging to transform the original data into a computable dataset;

[0010] The evaluation scope is determined based on the quality data analysis objectives, and the time interval is determined according to the time range of the objectives. Feature calculations are performed on the data within the time interval to form a feature set.

[0011] The features are classified, and the integrated evaluation results are obtained based on the integrated evaluation model. The dataset, feature set and integrated evaluation results are output.

[0012] Based on the above technical solution, the step of extracting material quality analysis-related data items from standardized testing reports and defining the attributes of the extracted data items includes the following steps:

[0013] Extract the data items related to material quality analysis sequentially from the raw data documents of standardized testing reports;

[0014] Define the data type and length of each data item;

[0015] The data items include the name of the testing unit, report number, project name, project location / purpose, testing basis, judgment basis, supplier, test date, testing parameters, test value, test result, and report date.

[0016] Based on the above technical solutions,

[0017] The null value is a value that is ignored and not filled in during the detection process. When the detection value is null, it is treated as invalid data.

[0018] The duplicate value is determined by the detection date and time and the report number of the detection report. When merging data, detections with the same date and the same report number belong to the same detection.

[0019] The invalid values ​​are cases where text descriptions are used instead of numerical records, and cases where input errors occur. Invalid values ​​are considered invalid data and are cleared.

[0020] The synonym values ​​appear in the project name, the purpose of the project section, the name of the testing unit, and the name of the manufacturer. For the project name, they are grouped according to the road section number and the construction or operation stage. For the purpose of the project section, they are grouped according to the three types: roadbed, pavement, and others. For the name of the testing unit and the name of the manufacturer, the abbreviations or simplified names are converted to the full names.

[0021] Based on the above technical solutions,

[0022] The data item labeling process specifically involves using codes to identify and distinguish the names of the testing unit, manufacturers, and projects.

[0023] The labeling of the testing unit's name is used for data analysis of the testing source;

[0024] The manufacturer's name is used for the manufacturer's quality data analysis;

[0025] The project name is used for material quality data analysis of the project.

[0026] Based on the above technical solution, the specific steps of determining the evaluation scope according to the quality data analysis target and determining the time interval according to the target time range include:

[0027] Based on the test result data table and the evaluation range determined according to the quality data analysis objectives, select the test values ​​of the corresponding objectives within the corresponding test time range;

[0028] Based on the time range of the dataset, the time range of the project duration, and the time range of the material usage period, the number of intervals to be divided is determined, and the detection time range is divided to obtain the time intervals.

[0029] Based on the above technical solutions,

[0030] The test result data table is specifically represented as follows:

[0031]

[0032] Among them, M N×T This represents the test result data table, m i,t Let N represent the t-th measurement value of the i-th target, N represent the number of targets, and T represent the number of days in the date and time range of the target detection.

[0033] The detection time range is divided into time intervals, and the corresponding formula is:

[0034] T0 = ​​TK

[0035] Where T0 represents the number of days in the time interval, and K represents the number of intervals to be divided.

[0036] Based on the above technical solution, the specific steps for performing feature calculations on data within the time interval to form a feature set include:

[0037] For the data within each time interval, amplitude-frequency characteristic analysis was performed to obtain the eigenvalue matrix:

[0038]

[0039] Among them, X N×J Let represent the eigenvalue matrix, N represent the number of targets, and J represent the number of amplitude-frequency features, specifically including the mean, variance, median, range, and coefficient of variation features. i,j This represents the j-th amplitude-frequency feature of the i-th target. This represents the feature vector of the i-th analysis object.

[0040] Based on the above technical solutions,

[0041] The mean is calculated as follows:

[0042]

[0043] Where, x i,1 t represents the mean. ik This represents the number of times the i-th target is detected within the k-th time interval;

[0044] The variance is calculated as follows:

[0045]

[0046] Where, x i,2 Indicates variance;

[0047] The median is calculated as follows:

[0048]

[0049] Where, x i,3 This represents the median;

[0050] The range is calculated as follows:

[0051] x i,4 =max{m i,t}-min{m i,t}

[0052] Where, x i,4 The range is represented by 'max', which indicates the calculation of the maximum value, and the minimum value is represented by 'min'.

[0053] The coefficient of variation is calculated as follows:

[0054]

[0055] Where, x i,5 This represents the coefficient of variation.

[0056] Based on the above technical solution, the specific steps of classifying features and obtaining integrated evaluation results based on the integrated evaluation model include:

[0057] Based on the quality model, features are classified, specifically:

[0058]

[0059] Among them, s i For the quality model f i The output score of (x);

[0060] A decision tree is used to group each feature into a single group to obtain the classification result. The normalization formula is as follows:

[0061]

[0062] Based on the integrated evaluation model and the output scores, the evaluation results are obtained, specifically:

[0063]

[0064] Where y represents the overall score. This indicates an integrated evaluation model.

[0065] This invention provides a data analysis system for the quality of transportation engineering materials with complex data sources, comprising:

[0066] The extraction module is used to extract data items related to material quality analysis based on standardized testing reports, and to define the attributes of the extracted data items;

[0067] The processing module is used to clean up null, duplicate, synonym, and invalid values ​​in the defined data items, as well as to merge data types and perform tagging, transforming the original data into a computable dataset.

[0068] The calculation module is used to determine the evaluation scope based on the quality data analysis objectives, determine the time interval according to the time range of the objectives, perform feature calculations on the data within the time interval, and form a feature set.

[0069] The evaluation module is used to classify features and obtain integrated evaluation results based on the integrated evaluation model, and outputs the dataset, feature set and integrated evaluation results.

[0070] Compared with the prior art, the advantages of the present invention are as follows:

[0071] 1. Material quality analysis is performed using test reports as the original data source, covering the entire process from extracting original data items from test reports to the quality analysis results, realizing dynamic quality analysis of multi-dimensional material testing parameters; for the purpose of test quality analysis, simplified data item attributes are defined, along with data cleaning and preprocessing methods, to form a computable and reliable material test dataset, simplifying the complex data relationships and data types of complex data sources;

[0072] 2. Dynamic time window segmentation is adopted to solve the problem of complex time spans in the original data, while refining the time intervals to constrain the influence range of individual random values ​​under complex conditions; the coefficient of variation feature of material test values ​​is defined and integrated with conventional features to make the material quality feature set more in line with the needs of engineering quality analysis; a JSON format data interface is adopted, and its dataset, feature value set and evaluation results are output in JSON format, which can be well connected with the test and inspection business system, making it easy to achieve data sharing and access control, and expanding the application of data analysis. Attached Figure Description

[0073] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0074] Figure 1 This is a flowchart illustrating a method for analyzing the quality data of transportation engineering materials with a complex data source, as described in an embodiment of the present invention. Detailed Implementation

[0075] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of this application, but not all embodiments.

[0076] This invention, based on feature engineering methods in data science, proposes a method for analyzing the quality data of transportation engineering materials with complex data sources to solve the problem of quality analysis of complex test and inspection data. This method enables a complete quality analysis process and achieves accurate, complete, and highly interpretable multi-dimensional quality evaluation results.

[0077] See Figure 1 As shown in the figure, the present invention provides a method for analyzing the quality data of transportation engineering materials with complex data sources, which specifically includes the following steps:

[0078] S1: Based on standardized testing reports, extract data items related to material quality analysis and define the attributes of the extracted data items;

[0079] In this invention, based on standardized testing reports, data items related to material quality analysis are extracted, and the attributes of the extracted data items are defined. Specific steps include:

[0080] S101: Extract the relevant data items for material quality analysis sequentially from the original data document of the standardized testing report; the data items include the name of the testing unit, report number, project name, project location / purpose, testing basis, judgment basis, supplier, test date, testing parameters, test value, test result and report date.

[0081] S102: Define the data type and data length for each data item;

[0082] This involves extracting the following data from the six main parts and over thirty raw data items in the testing report: testing unit name, report number, project name, project location / purpose, testing basis, judgment basis, supplier, test date, testing parameters, test value, test result, and report date. These streamlined data items are then used as attributes for the testing data analysis, covering the essential raw data items required for quality analysis objectives such as the testing source, project, location / purpose, and manufacturer. The data type and length of each data item are defined. Nested data items in the original report are then broken down into multiple testing data records.

[0083] S2: Clean up null, duplicate, synonym, and invalid values ​​in the defined data items, and perform data type merging and tagging to transform the original data into a computable dataset;

[0084] In this invention, null values ​​are values ​​that are ignored and not filled in during the detection process, such as null detection criteria. In this case, they can be filled in based on the detection value, detection result, and the context of the report. However, when the detection value is null, it is treated as invalid data.

[0085] When determining duplicate values ​​based on the detection date and time and the report number of the detection report, detections with the same date and report number are considered to be from the same detection.

[0086] Invalid values ​​are those where text descriptions replace numerical records (e.g., "good" or "medium" instead of numerical values) or those where there are input errors (e.g., incorrect decimal point placement or significant changes in the order of magnitude of integer values). Invalid values ​​are considered invalid data and are cleared.

[0087] Synonyms appear in project name, project part purpose, testing unit name, and manufacturer name. For project name, they are grouped according to road section number and construction or operation stage. For project part purpose, they are grouped according to three types: roadbed, pavement, and others. For testing unit name and manufacturer name, abbreviations or simplified names are converted to full names.

[0088] In this invention, the labeling process for data items is specifically as follows: the name of the testing unit, the name of the manufacturer, and the name of the project are identified and distinguished by using code labels; the label of the testing unit name is used for data analysis of the testing source; the label of the manufacturer name is used for quality data analysis of the manufacturer; and the label of the project name is used for material quality data analysis of the project.

[0089] Data type conversion also involves detection result items. Detection results are usually brief textual descriptions. Based on their actual meaning of "meets design requirements" or "does not meet design requirements", their data types are converted to logical values ​​of 1 or 0, respectively.

[0090] S3: Determine the evaluation scope based on the quality data analysis objectives, and determine the time interval according to the time range of the objectives. Perform feature calculations on the data within the time interval to form a feature set.

[0091] In this invention, the evaluation scope is determined based on the quality data analysis objectives, and the time interval is determined according to the time range of the objectives. Specific steps include:

[0092] S301: Based on the test result data table and the evaluation range determined according to the quality data analysis objectives, select the test value of the corresponding objective within the corresponding test time range;

[0093] S302: Based on the time range of the dataset, the time range of the project duration, and the time range of the material usage period, determine the number of intervals to be divided, and divide the detection time range to obtain the time intervals.

[0094] The test results data table is specifically presented as follows:

[0095]

[0096] Among them, M N×T This represents the test result data table, m i,t Let N represent the t-th measurement value of the i-th target, N represent the number of targets, and T represent the number of days in the date and time range of the target detection.

[0097] The detection time range is divided into time intervals, and the corresponding formula is:

[0098] T0 = ​​TK

[0099] Where T0 represents the number of days in the time interval, and K represents the number of intervals to be divided.

[0100] This involves inputting the test result data table, selecting the test values ​​of a certain test parameter for N analysis targets within a time range of T. For example, selecting N testing units, N manufacturers, N projects, or N types of engineering parts / uses to be analyzed, determining the evaluation range based on the quality data analysis target, and taking the test values ​​within the time range T for each target.

[0101] In the above test result data table, 1≤i≤N, 1≤t≤T. Selecting a time window allows for the analysis of material quality at any time interval within 1≤t≤T, resulting in a dynamic division of the test sequence.

[0102] In this invention, feature calculations are performed on data within a time interval to form a feature set. Specific steps include:

[0103] For the data within each time interval, amplitude-frequency characteristic analysis was performed to obtain the eigenvalue matrix:

[0104]

[0105] Among them, X N×J Let represent the eigenvalue matrix, N represent the number of targets, and J represent the number of amplitude-frequency features, specifically including the mean, variance, median, range, and coefficient of variation features. i,j This represents the j-th amplitude-frequency feature of the i-th target. This represents the feature vector of the i-th analysis object, and correspondingly, This represents the feature vector of the first analysis object. This represents the feature vector of the Nth analysis object.

[0106] In this invention, the mean is calculated as follows:

[0107]

[0108] Where, x i,1 t represents the mean. ik This represents the number of times the i-th target is detected within the k-th time interval;

[0109] The variance is calculated as follows:

[0110]

[0111] Where, x i,2 Indicates variance;

[0112] The median is calculated as follows:

[0113]

[0114] Where, x i,3 This represents the median;

[0115] The range is calculated as follows:

[0116] x i,4 =max{m i,t}-min{m i,t}

[0117] Where, x i,4 The range is represented by 'max', which indicates the calculation of the maximum value, and the minimum value is represented by 'min'.

[0118] The coefficient of variation is calculated as follows:

[0119]

[0120] Where, x i,5 This represents the coefficient of variation.

[0121] S4: Classify the features and evaluate them based on the ensemble evaluation model to obtain the ensemble evaluation results, and output the dataset, feature set and ensemble evaluation results.

[0122] In this invention, features are classified, and an integrated evaluation result is obtained based on an integrated evaluation model. Specific steps include:

[0123] S401: Based on the quality model, features are classified, specifically:

[0124]

[0125] Among them, s i For the quality model f i The output score of (x); from a modeling perspective, clustering, decision trees, random forests, Bayesian methods, support vector machines (SVMs), and neural networks can all be used to model f. i (x).

[0126] S402: Using a decision tree, each feature is grouped into a single group to obtain the classification result. The normalization formula is as follows:

[0127]

[0128] Based on the current raw and feature data of bulk materials, using a decision tree to group each feature can quickly and effectively obtain classification results.

[0129] S403: Based on the integrated evaluation model and the output scores, the evaluation results are obtained, specifically:

[0130]

[0131] Where y represents the overall score. This indicates an integrated evaluation model.

[0132] From a modeling perspective, complex cases can be handled using ensemble learning models, while simpler ones include weighted averages. Currently, based on the raw and feature data of bulk materials, multi-feature weighted averages can quickly and effectively yield ensemble evaluation results with good interpretability.

[0133] The output dataset, feature set, and ensemble evaluation results include labeled and unlabeled datasets, and feature sets including datasets with original feature values ​​and normalized values. The ensemble evaluation results are output in JSON data format.

[0134] The following specific examples illustrate the method for analyzing the quality data of transportation engineering materials according to the present invention.

[0135] Example 1

[0136] Strength and quality analysis of steel mechanical joints from the manufacturer's perspective, explaining the process of manufacturer-level analysis and the role of the coefficient of variation.

[0137] The goal of manufacturer analysis is to assist in material selection for engineering projects by analyzing the classification patterns of material quality from various manufacturers based on the historical material quality data provided by the manufacturers.

[0138] First, the original report data was extracted and data attributes were defined. Material test reports on tensile strength and residual deformation of individual tensile forces from 12 suppliers with comparable test volumes were selected from the period of May 2021 to May 2022, covering the strength and quality testing of steel mechanical joints. Manufacturer names were standardized to full names, and abbreviations and synonyms were merged, with manufacturers labeled S001 to S012. Data preprocessing included handling null values, duplicate values, invalid values, synonyms, and data types. The amount of valid data and the data items extracted from the dataset are shown in Tables 1 and 2 below.

[0139] Table 1. Data Analysis of Grade I Strength Quality of Steel Mechanical Joints from Manufacturer Perspective: Original Report Data Volume and Inspection Volume

[0140]

[0141]

[0142] Table 2. Strength Test Data Attributes of Steel Mechanical Joints

[0143] Attribute Name Data types Length (bytes) Name of testing unit Text 128 Report No. String 16 Testing basis String 32 Judgment basis String 8 Manufacturer Name Text 128 model String 16 grade String 8 Detection parameters String 8 Detection value Float 8 Test date and time Data 12 Result determination Logic 1 Report Date Data 8

[0144] The mean, variance, median, range, and coefficient of variation of the test data from each manufacturer were calculated. The characteristic values ​​are shown in Table 3.

[0145] Table 3. Characteristic Values ​​of Grade I Tensile Strength of Steel Mechanical Joints (Manufacturer Quality Analysis)

[0146]

[0147]

[0148] Feature classification and ensemble evaluation are shown in Table 4. Each feature value is linearly divided into 4 categories according to the decision tree threshold, and the ensemble value is allocated with an average weight of 20% for each of the 5 features.

[0149] Table 4. Analysis, Characteristic Classification, and Integrated Evaluation Results of Steel Mechanical Joint Manufacturers

[0150]

[0151] Compared to the general case, the evaluation using mean and variance features, as well as the feature quality analysis results reflecting the median and range of material quality, shows that the number of A and D categories is relatively large, at 40% or more. Adding the coefficient of variation feature enhances the weight of the material strength and stability quality features, while also increasing the number of manufacturers in the intermediate categories B and C, making the integrated classification result closer to a normal distribution and making the division of A and D categories more accurate.

[0152] Example 2

[0153] Based on the quality analysis of cement concrete strength testing and the dimensions of the testing source, this paper explains the complete process of the method and the detailed role of its dynamic characteristics.

[0154] The goal of the detection source analysis is to analyze whether there are regular differences in detection quality between two types of detection sources: on-site laboratory and non-on-site laboratory.

[0155] First, data extraction and data attribute definition were performed. Strength test reports for C30 and C50 grades, covering a quarterly period from December 21, 2021 to March 20, 2022, were extracted and categorized into two main groups based on the testing unit name: on-site laboratories and non-on-site laboratories. For tests involving both on-site and non-on-site laboratories, the labels were merged. Testing units with names containing "Project Section Laboratory" and "On-site Laboratory" were grouped into "On-site Laboratory." Testing units with names containing parent organization laboratories such as "Supervision Laboratory," "Chief Engineer's Office Laboratory," "Chief Supervisor's Laboratory," "Resident Office Laboratory," or "Central Laboratory" were grouped into "Non-on-site Laboratory." Other data preprocessing included handling null values, duplicate values, invalid values, and data type adjustments. The processed data volume and extracted data attributes are shown in Tables 5 and 6.

[0156] Table 5. Data on the 28-day compressive strength test of C30 and C50 cement concrete.

[0157]

[0158] Table 6. Attribute Table of Cement Concrete Strength Test Data

[0159]

[0160]

[0161] The detection data was divided into 10 time windows. The original data source, a quarter of 90 days, was further divided into 10 time periods, each lasting 9 days. Feature values ​​were calculated for each of the 10 time periods, and the mean, variance, and coefficient of variation were selected for analysis. The results are shown in Table 7.

[0162] Table 7 Characteristic values ​​of C30 strength test results for cement concrete in both on-site and off-site testing laboratories.

[0163]

[0164] Analyzing the above characteristics, a decision tree is used for the single-item evaluation model of dynamic feature values, and a weighted average is used for the ensemble evaluation model, which can achieve rapid evaluation results. Among them, the single indicator s... i Evaluation function f i (x) represents the frequency of occurrence of the better quality. X_i1, X_i2, and X_i5 are the mean, variance, and coefficient of variation evaluation values, respectively, and Y is the integrated evaluation value of the three. The feature evaluation and integrated evaluation results are shown in Table 8.

[0165] Table 8 Evaluation of Detection Sources at Construction Site Laboratories and Non-Construction Site Laboratories

[0166] f_i() X_i1 X_i2 X_i5 Y_i S_1 0.9 0.1 0.1 0.37 S_2 0.1 0.9 0.9 0.63

[0167] From the feature analysis and integrated evaluation of the examples, due to the addition of dynamic segmentation, the differences in mean and stability between the two were refined. Across 10 time slices, the non-site laboratory's test results showed a frequency of 0.9 compared to the site laboratory, while the stability quality was the opposite; the site laboratory outperformed the non-site laboratory by 0.9 frequencies. Specific evaluation scores are shown in Table x. Considering the actual situation, the non-site laboratory, which includes testing units from multiple parties involved in the project, is more susceptible to changes in personnel, machinery, materials, methods, and environment. Therefore, its test values ​​exhibit more significant fluctuations and greater instability than those from the site laboratory.

[0168] Example 3

[0169] Based on the quality analysis and engineering dimensions of cement concrete strength testing, this paper explains the complete process of the method and its detailed role in dynamic characteristics.

[0170] The goal of engineering analysis is to analyze whether there are regular differences in the quality of materials used in different projects.

[0171] First, data extraction and data attribute definition. Intensity testing reports for one quarter with C30 and C50 grades, aged 28 days, were extracted and labeled as Project A and Project B based on the project name. The project name was a merged label of the original project name, and the road section number was standardized after data cleaning, resulting in two codes: Project "A" and Project "B". Other data preprocessing included handling null values, duplicate values, invalid values, and data type adjustments. The processed data volume is shown in Table 9, and the extracted data attributes are the same as in Example 2, as shown in Table 6.

[0172] Table 9. Test data for compressive strength of C30 and C50 cement concrete.

[0173]

[0174] The detection data was divided into 10 time windows. The original data source, a quarter of 90 days, was further divided into 10 time periods, each lasting 9 days. Feature values ​​were calculated for each of the 10 time periods, and the mean, variance, and coefficient of variation were selected for analysis. The results are shown in Table 10 below.

[0175] Table 10 Characteristic values ​​of C30 strength test results for cement concrete in Road Engineering A and Road Engineering B

[0176]

[0177] Analyzing the above characteristics, a decision tree is used for the single-item evaluation model of dynamic feature values, and a weighted average is used for the ensemble evaluation model, which can achieve rapid evaluation results. Among them, the single indicator s... i Evaluation function f i (x) represents the frequency of occurrence of the better quality. X_i1, X_i2, and X_i5 are the mean evaluation value, variance evaluation value, and coefficient of variation evaluation value, respectively, and Y is the integrated evaluation value of the three. The feature evaluation and integrated evaluation results are shown in Table 11.

[0178] Table 11 Strength and Quality Evaluation of C30 Cement Concrete in Road Engineering A and Road Engineering B

[0179] f_i() X_i1 X_i2 X_i3 Y S_1 1 0.4 0.7 0.7 S_2 0 0.6 0.3 0.3

[0180] From the feature matrix of the example, due to the addition of dynamic time interval division, the differences between the two are refined. In the 10 time slices, the mean feature of project A is better than that of project B in all cases, the variance feature is better than that of B in a frequency of 0.4, and the coefficient of variation is better than that of B in a frequency of 0.7. Therefore, in terms of the stability of material quality, A is still better than B, which is consistent with the mean quality evaluation results of the project.

[0181] This invention provides a data analysis system for the quality of transportation engineering materials with complex data sources, including an extraction module, a processing module, a calculation module, and an evaluation module.

[0182] The extraction module extracts data items related to material quality analysis based on standardized testing reports and defines the attributes of these data items. The processing module cleans up null, duplicate, synonym, and invalid values ​​from the defined data items, merges data types, and performs labeling to transform the raw data into a computable dataset. The calculation module determines the evaluation scope based on the quality data analysis objectives and defines the time interval according to the objective's time range. It then performs feature calculations on the data within the time interval to form a feature set. The evaluation module classifies the features and evaluates them based on an integrated evaluation model to obtain integrated evaluation results. Finally, it outputs the dataset, feature set, and integrated evaluation results.

[0183] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

[0184] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

Claims

1. A traffic engineering material quality data analysis method with complex data sources, characterized in that, Specifically comprising the following steps: Based on the normalized detection report, the data items related to material quality analysis are extracted, and the attributes of the extracted data items are defined; The defined data items are cleaned of null values, duplicate values, synonymous values, and invalid values, and the data types and labels are merged to convert the original data into a computable data set; According to the quality data analysis target, the evaluation range is determined, and according to the time range of the target, the time interval is determined, and the data in the time interval is calculated to form a feature set; Classify the features, and evaluate the integrated evaluation results based on the integrated evaluation model, and output the data set, feature set, and integrated evaluation results; Wherein, the null value is a value ignored during detection and not filled in, when the detection value is null, it is treated as invalid data; The duplicate value is the same detection through the detection date and time and the report number of the detection report, and the same date and report number are consistent; The invalid value is a case where a numerical value is replaced by a text description, and an input error, for invalid values, is considered invalid data and is cleaned up; The synonymous value appears in the engineering name, engineering part purpose, detection unit name and manufacturer name, for the engineering name, according to the road section number and construction or operation stage, for the engineering part purpose, according to the roadbed, pavement and other three kinds of merging, for the detection unit name and manufacturer name, the abbreviation or abbreviation is converted to full name.

2. The traffic engineering material quality data analysis method with complex data sources as claimed in claim 1, wherein, Based on the normalized detection report, the data items related to material quality analysis are extracted, and the attributes of the extracted data items are defined, the specific steps comprising: From the normalized detection report, the data items related to material quality analysis are extracted in sequence; Define the data type and data length of each data item; Wherein, the data items include detection unit name, report number, engineering name, engineering part / purpose, detection basis, judgment basis, supplier, test date, detection parameter, detection value, detection result and report date.

3. The traffic engineering material quality data analysis method with complex data source according to claim 2, characterized in that: For data item label processing, specifically: detection unit name, manufacturer name and engineering name are identified and distinguished by using code labels; The label of the detection unit name is used for data analysis of the detection source; The label of the manufacturer name is used for quality data analysis of the manufacturer; The label of the engineering name is used for material quality data analysis of the engineering.

4. The traffic engineering material quality data analysis method with complex data sources as claimed in claim 1, wherein, The evaluation range is determined according to the quality data analysis target, and the time interval is determined according to the time range of the target, the specific steps comprising: According to the detection result data table and the evaluation range determined according to the quality data analysis target, the detection values corresponding to the target in the corresponding detection time range are selected; Based on the time range of the data set, the time range of the engineering period, and the time range of the material use period, the number of intervals to be divided is determined, the detection time range is divided, and the time interval is obtained.

5. The traffic engineering material quality data analysis method with complex data sources according to claim 4, characterized in that: the detection result data table is specifically represented as: wherein, represents a detection result data table, represents the number of targets, represents the number of measurements of the target, represents the number of targets, represents the number of days of the detection date time range of the target; the detection time range is divided to obtain a time interval, and the corresponding formula is: wherein, denotes the number of days of the time interval, denotes the number of intervals to be divided.

6. The traffic engineering material quality data analysis method with complex data sources as claimed in claim 5, wherein, the data in the time interval is subjected to feature calculation to form a feature set, and the specific steps include: the data in each time interval is subjected to amplitude-frequency feature analysis to obtain a feature value matrix: in, Represents the eigenvalue matrix. Indicates the number of targets. This indicates the number of amplitude-frequency features, specifically including the mean, variance, median, range, and coefficient of variation. Indicates the first The first goal Amplitude-frequency characteristics, Indicates the first The feature vector of the analysis object.

7. The traffic engineering material quality data analysis method with complex data sources according to claim 6, characterized in that: the mean value is calculated in the following manner: wherein, denotes the mean, denotes the number of detections of the target in the time interval. the variance is calculated in the following manner: wherein denotes the variance; the median is calculated in the following manner: wherein indicates the median; the range is calculated in the following manner: wherein, represents a range, represents a maximum value calculation, represents a minimum value calculation; the coefficient of variation is calculated in the following manner: wherein denotes the coefficient of variation.

8. The traffic engineering material quality data analysis method with complex data sources as claimed in claim 7, wherein, the features are classified, and an integrated evaluation result is obtained based on an integrated evaluation model, and the specific steps include: the features are classified based on the quality model, and the specific steps include: wherein, is an output score for the quality model . each feature is classified by using a decision tree to obtain a classification result, wherein the normalization formula is: ; an evaluation result is obtained based on the integrated evaluation model and the output score, and the specific steps include: wherein, represents the integrated score, represents the ensemble evaluation model.

9. A traffic engineering material quality data analysis system with complex data sources, characterized by, including: an extraction module for extracting material quality analysis related data items based on the normalized detection report, and defining the attributes of the extracted data items; a processing module for cleaning null values, duplicate values, synonymous values, invalid values, merging data types and marking processing of the defined data items, and converting the original data into a calculable data set; a calculation module for determining an evaluation range according to the quality data analysis target, determining a time interval according to the time range of the target, and calculating the features of the data in the time interval to form a feature set; an evaluation module for classifying the features and obtaining an integrated evaluation result based on an integrated evaluation model, and outputting the data set, feature set and integrated evaluation result; wherein the null value is a value ignored and not filled in during the detection process, and when the detection value is a null value, it is treated as invalid data; the duplicate value is a detection with the same date and report number belonging to the same detection when the detection date and time and the report number of the detection report are judged and the data are merged; the invalid value is a case where a numerical value is replaced by a text description or a case of input error, and for the invalid value, it is treated as invalid data and cleaned up; the synonymous value appears in the engineering name, engineering site purpose, detection unit name and manufacturer name, and for the engineering name, it is merged according to the road section number and construction or operation stage, for the engineering site purpose, it is merged according to the roadbed, pavement and other three, and for the detection unit name and manufacturer name, abbreviations or abbreviations are converted into full names.

Citation Information

Patent Citations

  • Oil well supply and production matching degree quantitative evaluation method based on multi-source data

    CN113807671A