Road section data anomaly detection processing method and electronic device

By clustering and anomaly detection of road segment data, abnormal road condition index values ​​are identified and corrected, solving the problem of reduced accuracy caused by abnormal data in the pavement performance prediction model, and achieving more efficient pavement performance prediction and scientific maintenance decisions.

CN121564952BActive Publication Date: 2026-07-21ROADMAINT CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ROADMAINT CO LTD
Filing Date
2025-10-14
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing pavement performance prediction models contain abnormal data in the detection data, which reduces the prediction accuracy. Furthermore, traditional methods are difficult to effectively handle various complex anomalies, lack multi-factor consideration and have insufficient intelligence, resulting in inappropriate maintenance plans and wasted funds.

Method used

By acquiring the initial road segment dataset and target road condition indicators, clustering and anomaly detection modules are used to identify and correct abnormal road condition indicator values, forming the target road segment dataset and improving the prediction accuracy of the road performance prediction model.

Benefits of technology

It improves the prediction accuracy of pavement performance prediction models, provides strong technical support for scientific and economical highway maintenance decisions, reduces data processing volume, and improves the efficiency of anomaly detection and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564952B_ABST
    Figure CN121564952B_ABST
Patent Text Reader

Abstract

The present disclosure provides a road section data anomaly detection processing method and an electronic device, obtaining an initial road section data set and a target road condition index; determining a target influencing factor corresponding to the target road condition index, clustering all road section data in the initial road section data set according to the target influencing factor to obtain a plurality of road section groups; for each road section group, performing anomaly detection processing on the initial road condition index values corresponding to all road section data in the road section group to determine abnormal road condition index values; determining a target abnormal type corresponding to the abnormal road condition index values, correcting the abnormal road condition index values according to the target abnormal type to obtain target road condition index values; counting all target road condition index values, and determining a target road section data set according to all target road condition index values and the initial road section data set. The target road section data set is a road section data set after the abnormal road condition index values are corrected, and the target road section data set is used for subsequent training of a road surface performance prediction model, thereby improving the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing, and in particular to a method and electronic device for detecting and processing road segment data anomalies. Background Technology

[0002] Through construction in recent years, my country has built a massive highway network. As a linear structure, highway infrastructure requires substantial investment to operate and maintain such a large scale of highway assets, to preserve and increase their value, and to better serve social and economic development.

[0003] Highway maintenance projects are one of the important means of operating and maintaining the value of highway assets. By carrying out maintenance projects on the most suitable road sections, it is possible to effectively save on capital costs and carbon emission costs, improve road conditions, and maximize the benefits of maintenance funds.

[0004] Currently, road condition data is typically analyzed through regular monitoring to identify road maintenance needs and formulate reasonable maintenance plans. However, in actual monitoring data, various abnormal data often arise due to factors such as measurement equipment errors, human operational mistakes, changes in monitoring trajectories, or special weather events, which leads to a decrease in the model's prediction accuracy. Summary of the Invention

[0005] In view of this, the purpose of this disclosure is to propose a method and electronic device for detecting and processing road segment data anomalies, in order to solve or partially solve the above-mentioned problems.

[0006] To achieve the above objectives, the first aspect of this disclosure provides a method for detecting and processing road segment data anomalies, including: Obtain an initial road segment dataset and target road condition indicators, wherein each road segment in the initial road segment dataset includes multiple initial influencing factors and initial road condition indicator values; Determine the target influencing factors corresponding to the target road condition index, and perform clustering processing on all road segment data in the initial road segment dataset based on the target influencing factors to obtain multiple road segment groups; For each road segment group, anomaly detection processing is performed on the initial road condition index values ​​corresponding to all road segment data within the road segment group to determine the abnormal road condition index values ​​within the road segment group. Determine the target anomaly type corresponding to the abnormal road condition index value, and correct the abnormal road condition index value according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group. Statistically calculate all target road condition index values ​​corresponding to all road segment groups, and determine the target road segment dataset based on the all target road condition index values ​​and the initial road segment dataset.

[0007] Based on the same inventive concept, a second aspect of this disclosure proposes a road segment data anomaly detection and processing device, comprising: The data acquisition module is configured to acquire an initial road segment dataset and target road condition indicators, wherein each road segment data in the initial road segment dataset includes multiple initial influencing factors and initial road condition indicator values; The clustering module is configured to determine the target influencing factors corresponding to the target traffic condition index, and to perform clustering processing on all road segment data in the initial road segment dataset based on the target influencing factors to obtain multiple road segment groups. The anomaly detection module is configured to perform anomaly detection processing on the initial traffic condition index values ​​corresponding to all road segment data within each road segment group, and determine the abnormal traffic condition index values ​​within the road segment group. An anomaly correction module is configured to determine the target anomaly type corresponding to the abnormal road condition index value, and to correct the abnormal road condition index value according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group. The target road segment dataset determination module is configured to statistically analyze all target road condition index values ​​corresponding to all road segment groups, and determine the target road segment dataset based on the all target road condition index values ​​and the initial road segment dataset.

[0008] Based on the same inventive concept, a third aspect of this disclosure proposes an electronic device including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0009] Based on the same inventive concept, a fourth aspect of this disclosure provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the methods described above.

[0010] As can be seen from the above, this disclosure proposes a method and electronic device for detecting and processing road segment data anomalies. It acquires an initial road segment dataset and a target road condition index. Each road segment in the initial road segment dataset includes multiple initial influencing factors and initial road condition index values. The initial influencing factors are factors considered by the pavement performance prediction model during prediction, and the initial road condition index values ​​are specific values ​​corresponding to the target road condition index, i.e., the initial road condition index values ​​are the prediction results of the pavement performance prediction model. The initial influencing factors affect the initial road condition index values. The method determines the target influencing factors corresponding to the target road condition index and clusters all road segment data in the initial road segment dataset based on these target influencing factors to obtain multiple road segment groups. By clustering the road segment data in the initial road condition dataset, all road segment data is divided into multiple data groups, allowing for separate processing of each data group, reducing data processing volume, and improving the efficiency of subsequent identification and processing of abnormal road condition index values. For each road segment group, anomaly detection processing is performed on the initial road condition index values ​​corresponding to all road segment data within the group to determine the abnormal road condition index values ​​within the group. The target anomaly type corresponding to the abnormal road condition index value is determined, and the abnormal road condition index value is corrected according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group. By identifying the abnormal road condition index value and correcting it according to its corresponding target anomaly type, the correction of the abnormal road condition index value is made more accurate. All target road condition index values ​​corresponding to all road segment groups are counted, and the target road segment dataset is determined based on the all target road condition index values ​​and the initial road segment dataset. At this time, the target road segment dataset is the road segment dataset after correcting the abnormal road condition index values. Subsequently, the target road segment dataset is used to train the pavement performance prediction model, which improves the prediction accuracy of the pavement performance prediction model and provides strong technical support for realizing scientific and economical highway maintenance decisions. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of the road segment data anomaly detection and processing method according to an embodiment of this disclosure; Figure 2 This is a schematic diagram of the structure of the road segment data anomaly detection and processing device according to an embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0014] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0015] Through construction in recent years, my country has built a massive highway network. As a linear structure, highway infrastructure requires substantial investment to operate and maintain such a large scale of highway assets, to preserve and increase their value, and to better serve social and economic development.

[0016] Highway maintenance projects are one of the important means of operating and maintaining the value of highway assets. By carrying out maintenance projects on the most suitable road sections, it is possible to effectively save on capital costs and carbon emission costs, improve road conditions, and maximize the benefits of maintenance funds.

[0017] In the field of road engineering and maintenance, the use of machine learning technology to predict road surface conditions has become a trend. In recent years, a large amount of historical data has been accumulated from numerous road network technical condition inspection and evaluation works.

[0018] Currently, road condition data is typically collected and analyzed regularly to determine road maintenance needs and formulate reasonable maintenance plans. During the maintenance decision-making process, historical data on various road performance indicators can be used to accurately predict the road's development trend over the next 3 to 5 years, significantly improving the scientific rigor of maintenance analysis and decision-making compared to existing models.

[0019] Currently, commonly used pavement performance prediction models include multiple regression, support vector machines, and neural networks, which have been widely applied to pavement performance prediction modeling and have achieved high accuracy. However, these models are highly dependent on data quality. In actual detection data, various anomalous data often arise due to factors such as measurement equipment errors, human operational mistakes, changes in detection trajectories, or special weather events. These anomalous data violate the changing patterns of domain knowledge, and if they are not accurately identified and properly handled, they will seriously affect the model training effect and reduce the model's prediction accuracy.

[0020] Regarding the quality control of pavement inspection data, some existing solutions are beginning to focus on the screening and cleaning of outlier data. For example, one method utilizes the historical development patterns of pavement performance over many years to identify anomalies: by traversing the inspection sequences of a road segment over many years, the longest subsequence that maintains a monotonically increasing or decreasing trend is found, and data outside this subsequence are marked as outliers. Based on this, the detected outlier data is cleaned; for example, outliers are directly deleted for certain comprehensive performance indices, while for continuous indicators such as rut ​​depth, smoothness, and skid resistance coefficient, linear interpolation is used to replace them with adjacent normal data points. This method considers historical trends while preserving the original data as much as possible, and can perform quality assessment and review on both new and old data, thus possessing certain practical value.

[0021] Other studies have employed machine learning-based residual prediction to locate outliers. Because numerous factors influence road surface performance, it's difficult to determine which outliers need correction or removal through simple statistical analysis. Therefore, the prediction model is first trained using initial data, and outliers are identified by the deviation between predicted and actual values. The model is then repeatedly trained with corrected data to iteratively improve prediction accuracy. These methods demonstrate the crucial role of more intelligent anomaly detection in enhancing model accuracy.

[0022] Existing methods often employ fuzzy processing, simple elimination, or ordinary interpolation techniques, which cannot effectively handle various complex anomalies and fail to systematically consider the combination of key factors affecting pavement performance prediction. This leads to biases in the assessment of current road conditions and the prediction of future performance, making it difficult to guarantee the reliability of the processed data. Consequently, inappropriate maintenance plans may be formulated, failing to effectively address pavement defects and resulting in wasted maintenance funds.

[0023] In summary, the current data on detected road surface performance has the following shortcomings: Firstly, anomaly detection is incomplete. Current methods for identifying outliers in pavement performance data rely on a single criterion, such as statistical thresholds or monotonic trends, which cannot effectively handle anomalies from diverse sources and in various forms. For instance, relying solely on the 3σ principle or simple trend analysis makes it difficult to discover anomalies hidden in complex contexts.

[0024] Secondly, the lack of multi-factor consideration means that road performance is affected by multiple factors such as structure, traffic, climate, and maintenance. Traditional methods often fail to classify and process data according to the influencing factors, and cannot formulate appropriate anomaly identification standards for different situations, resulting in some anomalies not being identified or normal data being misjudged.

[0025] Furthermore, the processing methods are simplistic. Currently, most outliers are handled by either direct deletion or linear interpolation, without selecting the optimal completion strategy based on different indicators and anomaly scenarios. This may make it difficult to strike a balance between preserving valid data information and eliminating anomalies.

[0026] Finally, the level of intelligence is not high. Currently, the detection of outliers in road performance data usually requires manual setting of rules or thresholds, such as manually judging data trends and subjectively deciding whether to delete data. It lacks intelligent learning capabilities and cannot adapt to large-scale data and changing road conditions.

[0027] To address the aforementioned problems, this embodiment proposes a method for detecting and processing road segment data anomalies, such as... Figure 1 As shown, the method includes: Step 101: Obtain the initial road segment dataset and target road condition indicators, wherein each road segment data in the initial road segment dataset includes multiple initial influencing factors and initial road condition indicator values.

[0028] In specific implementation, an initial road segment dataset and target road condition indicators are obtained. The target road condition indicators are the prediction results output by the road performance prediction model after the initial road performance prediction model is trained based on the corrected target road segment dataset. That is, the road performance indicators predicted by the road performance prediction model.

[0029] In this embodiment, the target road condition indicators include at least one of the following: Pavement Damage Index (PCI), Road Quality Index (RQI), Road Rutting Index (RDI), Road Skid Resistance Index (SRI), Road Wear Index (PWI), Road Bump Index (PBI), Road Deflection Index (PSSI), and distress value indicators, etc.

[0030] In this embodiment, the initial road segment dataset is the collected existing road segment data. Each road segment data in the initial road segment dataset includes multiple initial influencing factors and initial road segment index values. The initial road segment index values ​​are the specific values ​​corresponding to the target road condition index of the road segment data determined based on the multiple initial influencing factors.

[0031] In this embodiment, the initial influencing factors are all the factors of the multi-dimensional data required for the pavement performance prediction model to make predictions. The multi-dimensional data includes pavement structure parameters, traffic load, climate environment, maintenance history, etc. The initial influencing factors specifically include at least one of the following: pavement structure type, pavement structure gradation type, pavement structure material type, annual average daily traffic volume, number of passenger cars, number of trucks, average temperature, minimum temperature, precipitation, maintenance time of maintenance history, material type of maintenance history, gradation type of maintenance history, etc.

[0032] In this embodiment, the gradation type of the road structure refers to the combination and arrangement of particles of different sizes in the road material. It is mainly classified according to the particle composition, mineral composition and porosity, nominal maximum particle size of aggregate and structural type. These classification methods jointly determine the density, strength, stability, drainage performance and other characteristics of the road material, thereby affecting the performance and service life of the road.

[0033] In this embodiment, the gradation type of the maintenance history refers to the gradation type of the pavement structure after maintenance when the pavement was maintained at a historical time.

[0034] Step 102: Determine the target influencing factors corresponding to the target road condition index, and perform clustering processing on all road segment data in the initial road segment dataset based on the target influencing factors to obtain multiple road segment groups.

[0035] In practice, the target influencing factors corresponding to the target road condition index are determined. These target influencing factors are those selected from the initial influencing factors that have a significant impact on the target road segment index. Based on the target influencing factors, all road segment data in the initial road segment dataset are clustered, i.e., all road segment data are classified according to the target influencing factors to obtain multiple road segment groups.

[0036] Step 103: For each road segment group, perform anomaly detection processing on the initial road condition index values ​​corresponding to all road segment data within the road segment group to determine the abnormal road condition index values ​​within the road segment group.

[0037] In practice, for each road segment group, anomaly detection processing is performed on the initial road condition index values ​​corresponding to all road segment data within the group. That is, within each road segment group, anomaly detection is performed on the initial road condition index values ​​corresponding to all road segment data within the group to obtain the abnormal road condition index values ​​for that road segment group.

[0038] Step 104: Determine the target anomaly type corresponding to the abnormal road condition index value, and correct the abnormal road condition index value according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group.

[0039] In practice, for each road segment group, after determining the abnormal road condition index value corresponding to the road segment group, the target abnormality type corresponding to the abnormal road condition index value is determined, wherein the target abnormality type is a local abnormality or a range abnormality.

[0040] In this embodiment, since each road segment group includes at least one road segment data, when there are multiple road segment data in a road segment group, the number of abnormal road condition index values ​​determined within the road segment group may also be multiple. When the number of abnormal road condition index values ​​is multiple, the target anomaly type corresponding to each abnormal road condition index value should be determined separately. That is, the target anomaly types corresponding to different abnormal road condition index values ​​within the same road segment group may be different.

[0041] The abnormal road condition index value is corrected according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group.

[0042] Step 105: Calculate all target road condition index values ​​corresponding to all road segment groups, and determine the target road segment dataset based on the all target road condition index values ​​and the initial road segment dataset.

[0043] In practice, for each road segment group, the road segment data within the road segment group before correction is divided into normal road segment data and abnormal road segment data. Each normal road segment data includes multiple initial influencing factors and initial road condition index values, and each abnormal road segment data includes multiple initial influencing factors and abnormal road condition index values.

[0044] After correcting the abnormal road condition index values ​​within the road segment group to obtain the target road condition index values, there are no abnormal road condition index values ​​within the road segment group at this point. That is, the road segment data within the road segment group is divided into normal road segment data and corrected abnormal road segment data. Each normal road segment data includes multiple initial influencing factors and initial road condition index values, while each corrected abnormal road segment data includes multiple initial influencing factors and target road condition index values.

[0045] All target road condition index values ​​corresponding to all road segment groups are statistically analyzed. Based on these all target road condition index values ​​and the initial road segment dataset, a target road segment dataset is determined. That is, the target road segment dataset includes multiple road segment groups, each including normal road segment data and corrected abnormal road segment data. Each normal road segment data includes multiple initial influencing factors and initial road condition index values, and each corrected abnormal road segment data includes multiple initial influencing factors and target road condition index values.

[0046] After determining the target road segment dataset, the pavement performance prediction model can be trained based on the target road segment dataset. The trained pavement performance prediction model can then be used to predict the performance of the target pavement, and the maintenance strategy corresponding to the target pavement can be determined based on the prediction results.

[0047] The above scheme obtains an initial road segment dataset and target road condition indicators. Each road segment in the initial road segment dataset includes multiple initial influencing factors and initial road condition indicator values. The initial influencing factors are those considered by the pavement performance prediction model during prediction, and the initial road condition indicator values ​​are the specific values ​​corresponding to the target road condition indicators, i.e., the prediction results of the pavement performance prediction model. The initial influencing factors affect the initial road condition indicator values. The target influencing factors corresponding to the target road condition indicators are determined. Based on the target influencing factors, all road segment data in the initial road segment dataset are clustered to obtain multiple road segment groups. By clustering the road segment data in the initial road condition dataset, all road segment data are divided into multiple data groups, allowing for separate processing of each data group, reducing data processing volume, and improving the efficiency of subsequent identification and processing of abnormal road condition indicator values. For each road segment group, anomaly detection processing is performed on the initial road condition indicator values ​​corresponding to all road segment data within the group to determine the abnormal road condition indicator values ​​within the group. The target anomaly type corresponding to the abnormal road condition index value is determined, and the abnormal road condition index value is corrected according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group. By identifying the abnormal road condition index value and correcting it according to its corresponding target anomaly type, the correction of the abnormal road condition index value is made more accurate. All target road condition index values ​​corresponding to all road segment groups are counted, and the target road segment dataset is determined based on the all target road condition index values ​​and the initial road segment dataset. At this time, the target road segment dataset is the road segment dataset after correcting the abnormal road condition index values. Subsequently, the target road segment dataset is used to train the pavement performance prediction model, which improves the prediction accuracy of the pavement performance prediction model and provides strong technical support for realizing scientific and economical highway maintenance decisions.

[0048] In some embodiments, step 102 involves determining the target influencing factors corresponding to the target traffic condition index, and clustering all road segment data in the initial road segment dataset based on the target influencing factors to obtain multiple road segment groups. Specifically, this includes: Step 1021: Obtain the pavement base structure type and traffic flow level corresponding to each road segment data in the initial road segment dataset; Step 1022: Divide all road segment data according to the road base structure type and the traffic flow level to obtain multiple road segment data subsets; Step 1023: For each road segment data subset, determine the target influencing factors corresponding to the target road condition index in the road segment data subset, and perform clustering processing on all road segment data in the road segment data subset according to the target influencing factors to obtain multiple road segment groups.

[0049] In specific implementation, the pavement base structure type and traffic flow level corresponding to each road segment in the initial road segment dataset are obtained. The pavement base structure type indicates the type of base material used to compose the base structure of that road segment, and includes semi-rigid base, composite base, and other types. The traffic flow level indicates the traffic flow intensity of that road segment, and can be divided into light traffic, moderate traffic, heavy traffic, and extremely heavy traffic.

[0050] Based on the road surface base structure type and the traffic flow level, all road segment data are divided into multiple road segment data subsets. Specifically, the specific road surface base structure type and traffic flow level of each road segment data can be determined, and road segment data with the same road surface base structure type and the same traffic flow level are grouped together, that is, divided into a road segment data subset.

[0051] For example, if the pavement base structure type of road segment data A is semi-rigid base and the traffic flow intensity is medium traffic, the pavement base structure type of road segment data B is semi-rigid base and the traffic flow intensity is light traffic, and the pavement base structure type of road segment data C is semi-rigid base and the traffic flow intensity is medium traffic, then road segment data A and road segment data C are classified into the same road segment data subset, and road segment data B is classified into a separate road segment data subset.

[0052] For each road segment data subset, the target influencing factors corresponding to the target road condition indicators in the road segment data subset are determined. Then, based on the target influencing factors, all road segment data in the road segment data subset are clustered to obtain multiple road segment groups.

[0053] In this embodiment, each road segment group contains at least one road segment data. When clustering all road segment data in the subset of road segment data according to the target influencing factors, a feature vector is constructed for each road segment data based on the target influencing factors. K-means clustering is used for analysis, the elbow method is selected to determine the optimal number of clusters, and the within-group variance and mean distribution of the clustering results are checked. If necessary, the number of clusters is adjusted and re-clustering is performed. Then, the cluster labels are mapped to specific road segment data to form a dataset with similar conditions for each label, that is, each label corresponds to a road segment group, ultimately resulting in multiple road segment groups.

[0054] In some embodiments, determining the target influencing factors corresponding to the target traffic condition index in the road segment data subset in step 1023 specifically includes: Step 10231: Obtain all road segment data included in the road segment data subset, and determine multiple target feature labels based on the all road segment data, wherein each target feature label contains an initial road condition index value and an initial influencing factor; Step 10232: Input the multiple target feature labels into the random forest model, process them through the random forest model, and output the importance weight corresponding to each initial influencing factor; Step 10233: Obtain the initial influencing factors whose importance weight is greater than the preset weight threshold, and obtain the candidate factor sequence matrix; Step 10234: Input the candidate factor sequence matrix into the principal component analysis model, process it through the principal component analysis model, and output the target influencing factors.

[0055] In practice, all road segment data included in the road segment data subset are obtained, and multiple target feature labels are determined based on all road segment data. Each target feature label contains an initial road condition index value and an initial influencing factor.

[0056] Taking one road segment from the road segment data subset as an example, this road segment data includes initial influencing factors and initial road condition index values. The initial influencing factors are pavement structure type, average temperature of the climate environment, and number of passenger cars. The initial road condition index value is PCI of 80. Therefore, the target feature labels corresponding to this road segment data are determined to be: pavement structure type is flexible pavement with PCI value of 80, average temperature of the climate environment is 20 degrees Celsius with PCI value of 80, and number of passenger cars is 1000 with PCI value of 80.

[0057] The multiple target feature labels are input into a random forest model. The random forest model processes the data and outputs the importance weights for each initial influencing factor corresponding to each road segment. Specifically, the importance weights for each initial influencing factor are calculated using the `feature_importances_` attribute during the random forest model processing.

[0058] For each road segment data, the importance weights of all initial influencing factors are traversed, and initial influencing factors with importance weights greater than a preset weight threshold are selected to form a candidate factor sequence matrix for that road segment data. For example, based on the importance weights sorted from high to low, initial influencing factors with importance weights greater than 85% of the preset weight threshold are initially selected as initial influencing factors with a high contribution to the target road condition index.

[0059] The entire candidate factor sequence matrix is ​​input into the principal component analysis model, and after processing by the principal component analysis model, the target influencing factors are output.

[0060] In this embodiment, the candidate factor sequence matrix corresponding to each road segment data can be normalized by mean and variance, and then connected to the principal component analysis model to calculate each principal component and its variance contribution rate. The principal component with a variance contribution rate of 90% is selected as the target influencing factor.

[0061] For example, the initial road segment dataset is divided into three subsets: A, B, and C, based on the road surface base structure type and traffic flow level. For each subset, such as A, the target influencing factors are determined based on all road segment data within A. Then, the data from all road segments in subset A is clustered according to these target influencing factors, resulting in road segment groups A1, A2, and A3. Subsequently, anomaly detection is performed on the initial traffic condition index values ​​corresponding to all road segment data within each group. Taking road segment group A1 as an example, abnormal traffic condition index values ​​within A1 are identified.

[0062] The above scheme, when determining the target influencing factors corresponding to the target road condition indicators, firstly divides all road segment data according to the road surface base structure type and traffic flow level, obtaining multiple road segment data subsets. Then, the road segment data in each subset is clustered. Finally, the initial influencing factors are screened based on the random forest model and principal component analysis model to obtain the target influencing factors. This process of classification, clustering, and screening ensures that subsequent detection of abnormal road condition indicator values ​​is carried out on homogeneous road segment data with similar decay backgrounds, thereby greatly improving the sensitivity and accuracy of detection.

[0063] In some embodiments, step 103 involves performing anomaly detection processing on the initial traffic condition index values ​​corresponding to all road segment data within each road segment group to determine abnormal traffic condition index values ​​within the road segment group. This specifically includes: Step 1031: For each road segment group, perform statistical analysis on the initial road condition index values ​​corresponding to all road segment data within the road segment group to obtain the first abnormal road condition index value. Step 1032: Take all initial road condition index values ​​except the first abnormal road condition index value as the first initial road condition index value, obtain the target abnormal rule corresponding to the target road condition index, and take the first initial road condition index value that satisfies the target abnormal rule as the second abnormal road condition index value. Step 1033: Take the other initial road condition index values ​​in the first initial road condition index values, except for the second abnormal road condition index value, as the second initial road condition index value. Input the second initial road condition index value into the anomaly identification model. After processing by the anomaly identification model, output the third abnormal road condition index value. Step 1034: Summarize the first abnormal road condition index value, the second abnormal road condition index value, and the third abnormal road condition index value to obtain the abnormal road condition index value within the road segment group.

[0064] In practice, for each road segment group, the initial road condition index values ​​corresponding to all road segment data within the road segment group are statistically analyzed to obtain a first abnormal road condition index value, wherein the first abnormal road condition index value is at least one of the initial road condition index values.

[0065] Take all initial road condition index values ​​except the first abnormal road condition index value as the first initial road condition index value, obtain the target abnormal rule corresponding to the target road condition index, and take the first initial road condition index value that satisfies the target abnormal rule as the second abnormal road condition index value.

[0066] The target anomaly rules are pre-defined rules developed using domain knowledge of pavement performance evolution. Typically, the performance indicators of asphalt pavements tend to decline or remain stable over time, and are unlikely to increase spontaneously unless maintenance is implemented. However, the area of ​​pavement distress (cracks and strip repairs are collectively defined as one type of distress) generally accumulates year by year and does not decrease naturally. Different target pavement condition indicators correspond to different target anomaly rules; the specific target anomaly rule can be found by searching for the target pavement condition indicator.

[0067] For example, if the target road condition index is PCI, the corresponding target anomaly rule is an increase of 2 points or a decrease of 5 points in any two-year period. Simultaneously, when the target road condition index is PCI, a change of no more than 10% in the total area of ​​cracks and repair defects over two years can be considered. If the target road condition index is RQI, the corresponding target anomaly rule is a fluctuation of 2 points within any two-year period. If the target road condition index is RDI, the corresponding target anomaly rule is a fluctuation of 2 points within any two-year period. If the target road condition index is a defect value, the corresponding target anomaly rule is a decrease in the defect area within any two-year period. Furthermore, when the target road condition index is a defect value, transverse cracks, longitudinal cracks, and strip repairs are all calculated as crack-type defects, while other defect types are calculated separately.

[0068] It is understandable that since the second abnormal road condition index value is the abnormal road condition index value determined when the target abnormal rule corresponding to the target road condition index is met, there is no second abnormal road condition index value when there is no initial road condition index value that meets the target abnormal rule.

[0069] The other initial road condition index values ​​in the first initial road condition index value, excluding the second abnormal road condition index value, are used as the second initial road condition index value. The second initial road condition index value is input into the anomaly identification model, and after processing by the anomaly identification model, the third abnormal road condition index value is output.

[0070] The abnormal road condition index values ​​of the first abnormal road condition index, the second abnormal road condition index, and the third abnormal road condition index are summarized to obtain the abnormal road condition index values ​​within the road segment group.

[0071] The above scheme uses statistical analysis, target anomaly rules corresponding to target road condition indicators, and anomaly identification models to identify abnormal road condition indicator values. It can not only use statistics and machine learning to discover anomalies in data distribution, but also use domain knowledge rules to filter out pseudo-variables that do not conform to physical laws. This enables the accurate identification of various complex anomalies, including global statistical anomalies, time series logical anomalies, multi-dimensional feature combination anomalies, and local density anomalies, with higher accuracy and comprehensiveness.

[0072] In some embodiments, step 1031 involves statistically analyzing the initial road condition index values ​​corresponding to all road segment data within the road segment group to obtain a first abnormal road condition index value, specifically including: Step 10311: Obtain the initial road condition index values ​​corresponding to all road segment data within the road segment group, and determine the target mean and target standard deviation corresponding to all initial road condition index values; Step 10312: Sum the target mean with the target standard deviation of a preset multiple to obtain the maximum road condition index value; Step 10313: Subtract the target mean from the target standard deviation by a preset multiple to obtain the minimum road condition index value; Step 10314: In response to the existence of an initial road condition index value that is greater than the maximum road condition index value or less than the minimum road condition index value, the initial road condition index value is determined to be the first abnormal road condition index value.

[0073] In specific implementation, the initial road condition index values ​​corresponding to all road segments within the road segment group are obtained, and the target mean and target standard deviation corresponding to all initial road condition index values ​​are determined. A preset multiple is obtained, and the target mean is summed with the target standard deviation of the preset multiple to obtain the maximum road condition index value. The difference between the target mean and the target standard deviation of the preset multiple is calculated to obtain the minimum road condition index value.

[0074] The normal range is determined based on the minimum and maximum road condition index values. If the initial road condition index value is less than or equal to the maximum road condition index value and greater than or equal to the minimum road condition index value, then the initial road condition index value meets the normal range. In this case, the initial road condition index value is the normal road condition index value, and the corresponding road segment data is the normal road segment data.

[0075] If the initial road condition index value is greater than the maximum road condition index value or less than the minimum road condition index value, it is determined that the initial road condition index value is abnormal, and the initial road condition index value is taken as the first abnormal road condition index value.

[0076] For example, the preset multiplier is 3, and the initial road condition index value corresponding to the road segment data is the annual value of the road condition index for that road segment. A frequency distribution histogram is plotted or statistical characteristics are calculated to determine the target mean μ and target standard deviation σ corresponding to all initial road condition index values. The normal fluctuation range is set to μ ± 3σ, meaning 99.7% of the data should fall within this range. If the initial road condition index value is not within the normal fluctuation range, then the initial road condition index value is determined to be the first abnormal road condition index value.

[0077] In some embodiments, the anomaly detection model includes an isolated forest model and an outlier factor model. In step 1033, the second initial road condition index value is input into the anomaly detection model, processed by the anomaly detection model, and a third abnormal road condition index value is output, specifically including: For each second initial road condition index value: Step 10331: Input the second initial road condition index value into the isolated forest model, process it through the isolated forest model, and output the first anomaly score corresponding to the second initial road condition index value. Step 10332: Input the second initial road condition index value into the outlier factor model, process it through the outlier factor model, and output the second anomaly score corresponding to the second initial road condition index value. Step 10333: Perform weighted processing on the first abnormal score and the second abnormal score to obtain the target abnormal score corresponding to the second initial road condition index value; Step 10334: In response to the target abnormal score being greater than a preset score threshold, the second initial road condition index value is determined to be the third abnormal road condition index value.

[0078] In practice, all the second initial road condition index values ​​are input into the isolated forest model. The isolated forest model processes these values ​​and outputs a first anomaly score corresponding to each second initial road condition index value. Specifically, the isolated forest algorithm in the isolated forest model constructs multiple binary trees using random subsampling and random partitioning to determine the first anomaly score corresponding to each second initial road condition index value. The first anomaly score represents the probability that each second initial road condition index value is an anomalous road condition index value; the higher the first anomaly score, the more likely the corresponding second initial road condition index value is an anomalous road condition index value.

[0079] All initial second road condition index values ​​are input into the outlier model. The outlier model processes these values ​​and outputs a second anomaly score for each initial second road condition index value. Specifically, the outlier model uses a local outlier algorithm to compare the local reachability density of each initial second road condition index value with its neighborhood. Outliers with densities significantly lower than the neighborhood mean are identified, thus better identifying isolated points of anomalousness within the same category. The second anomaly score represents the probability that each initial second road condition index value is an anomalous value; the higher the second anomaly score, the more likely the corresponding initial second road condition index value is to be an anomalous value.

[0080] A target fusion weight is determined, and the first and second abnormal scores are weighted to obtain the target abnormal score corresponding to the second initial road condition index value. If the target abnormal score is greater than a preset score threshold, the second initial road condition index value is determined to be the third abnormal road condition index value.

[0081] In this embodiment, the target fusion weights are weight values ​​determined by an adaptive model. The Isolation Forest model can be considered as a global sparsity strategy, and the Outlier Factor model as a local sparsity strategy. If the global variance is large, the weights are biased towards global sparsity, meaning the weight value corresponding to the first outlier score is higher. If the intra-group density gradient is significant, the weights are biased towards local sparsity, meaning the weight value corresponding to the second outlier score is higher.

[0082] In some embodiments, determining the target anomaly type corresponding to the abnormal road condition index value in step 104 specifically includes: Step 1041: Obtain the number of index values ​​for the abnormal road condition index; Step 1042: In response to the fact that the number of indicator values ​​is one, determine that the target anomaly type corresponding to the abnormal road condition indicator value is a local anomaly; or, Step 1043: In response to the fact that there are multiple index values, obtain the road segment identifier of the road segment data corresponding to each abnormal road condition index value, determine that the target abnormality type corresponding to the abnormal road condition index value corresponding to the consecutive road segment identifiers that are greater than a preset number threshold is a range abnormality, and determine that the target abnormality type corresponding to other abnormal road condition index values ​​other than the abnormal road condition index values ​​included in the range abnormality is a local abnormality.

[0083] In specific implementation, the number of abnormal road condition index values ​​is obtained, that is, the number of abnormal road condition index values ​​within the road segment group is obtained.

[0084] If the number of indicator values ​​is only one, meaning there is only one abnormal road condition indicator value within the road segment group, and only one road segment has abnormal data, then the target anomaly type corresponding to the abnormal road condition indicator value is determined to be a local anomaly.

[0085] If there are multiple index values, that is, there are multiple abnormal road condition index values ​​in the road segment group, obtain the road segment identifier of the road segment data corresponding to each abnormal road condition index value. The road segment identifier is used to distinguish different road segments, that is, each road segment has a unique road segment identifier.

[0086] After obtaining the segment identifier of the segment data corresponding to each abnormal traffic condition index value, it is determined whether multiple segment identifiers are consecutive segment identifiers, that is, whether the segment identifiers correspond to adjacent segment segments.

[0087] If multiple road segment identifiers are consecutive, and the number of consecutive road segment identifiers exceeds a preset threshold, then the target anomaly type for the abnormal road condition indicator values ​​corresponding to the consecutive road segment identifiers is determined to be a local anomaly. Other abnormal road condition indicator values, excluding those included in the range anomaly category, are also classified as local anomalies.

[0088] For example, the preset quantity threshold is 5, and the length of each road segment data is 1km. If the number of abnormal road condition index values ​​is determined to be 9, and the road segment identifiers corresponding to these abnormal road condition index values ​​are 101, 102, 103, 104, 105, 106, 123, 135, and 136, then road segment identifiers 101, 102, 103, 104, 105, and 106 are determined to be consecutive road segment identifiers, corresponding to consecutive road segments, and the number is greater than 5. That is, the length of the road segment with abnormal road condition index values ​​is 6km, which is greater than the road segment length of 5km corresponding to the preset quantity threshold. Therefore, the target anomaly type for the abnormal road condition index values ​​corresponding to road segment identifiers 101, 102, 103, 104, 105, and 106 is determined to be a range anomaly. The target anomaly type for the abnormal road condition index values ​​corresponding to road segment identifiers 123, 135, and 136 is determined to be a local anomaly.

[0089] The above approach first determines the anomaly type for each abnormal road condition index value, and then determines the corresponding correction method based on the anomaly type. This avoids the permanent loss of valuable information caused by simple rejection, and also avoids the introduction of systematic biases by simple interpolation, which would pollute the original dataset.

[0090] In some embodiments, step 104, which corrects the abnormal road condition index value according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group, specifically includes: Step 104A: In response to the target anomaly type being a local anomaly, determine the target completion index value according to a preset data completion algorithm, delete the abnormal road condition index value, and use the target completion index value as the target road condition index value corresponding to the road segment group; or, Step 104B: In response to the target anomaly type being a range anomaly, obtain the target year corresponding to the abnormal road condition index values ​​included in the range anomaly, obtain the actual road condition index values ​​of the road segment identifiers corresponding to the range anomaly in the adjacent years of the target year, and perform correction processing on the abnormal road condition index values ​​based on the actual road condition index values ​​to obtain the target road condition index values ​​corresponding to the road segment group.

[0091] In practice, if the target anomaly type is determined to be a local anomaly, the target completion index value can be determined according to the preset data completion algorithm, the abnormal road condition index value can be deleted, and the target completion index value can be used as the target road condition index value corresponding to the road segment group, that is, the abnormal road condition index value can be replaced by the target completion index value.

[0092] In this embodiment, a corresponding preset data completion algorithm can be determined based on the target road condition index. Specifically, for the road driving quality index (RQI), road rutting index (RDI), road skid resistance index (SRI), road wear index (PWI), road bounce index (PBI), and road deflection index (PSSI), since their interannual fluctuations are small, linear interpolation can be directly used to fill out outliers, which is simple and reliable.

[0093] For the pavement damage index (PCI) and the distress value index, since they fluctuate greatly over time and are affected by many factors, a single method is difficult to apply to all situations. Therefore, a three-stage broken line method or curve fitting method can be used to determine the target supplementary index value.

[0094] Specifically, the three-stage broken-line method divides the pavement performance degradation process into three stages based on expert experience or large amounts of data. Using a broken-line approximation, it defines reference values ​​for target pavement condition indicators for typical road sections at three key service life nodes, such as 99 points for year 0, 92 points for year 5, and 82 points for year 10. The method then determines the target year for the road section data corresponding to abnormal pavement condition indicator values, and uses the broken-line to find the target complete indicator value for that year.

[0095] The complete lifecycle degradation of road surface performance often exhibits an S-shaped curve characteristic. Based on spatiotemporal big data, road segments with high data quality in the road segment dataset are selected as references, and the parameters of the S-shaped curve are fitted, including the upper limit (initial value), lower limit (extreme value), and curve shape parameters a and b. The target year of the road segment data corresponding to the abnormal road condition index value is determined, and the curve is found according to the target year to obtain the target completion index value corresponding to the target year.

[0096] In this embodiment, for local anomalies, methods such as multinomial fitting, local weighted regression, and Kriging interpolation can be combined to smooth and complete complex fluctuating data. The specific method can be flexibly selected according to the data characteristics of the abnormal road condition index values ​​to achieve the best repair effect.

[0097] If the target anomaly type is determined to be a range anomaly, the target year corresponding to the abnormal road condition index value included in the range anomaly is obtained, the actual road condition index value of the road segment identifier corresponding to the range anomaly in the adjacent years of the target year is obtained, and the abnormal road condition index value is corrected according to the actual road condition index value to obtain the target road condition index value corresponding to the road segment group.

[0098] For example, if the target anomaly type is a range anomaly, the target road condition index is PCI, and the road condition identifiers corresponding to the range anomaly are 101, 102, 103, 104, 105, and 106, the target year corresponding to the abnormal road condition index values ​​included in the range anomaly is determined to be 2015. The actual road condition index values ​​of the road segments corresponding to the range anomaly in the adjacent years to the target year are obtained, that is, the actual road condition index value PCI2014 corresponding to road condition identifiers 101, 102, 103, 104, 105, and 106 in 2014, and the actual road condition index value PCI2016 corresponding to road condition identifiers 101, 102, 103, 104, 105, and 106 in 2016 are obtained. The target road condition index value is determined by interpolating the actual road condition index values ​​of 2014 and 2016, which involves averaging the actual road condition index values ​​of PCI2014 and PCI2016.

[0099] It is understandable that, in the example above, when the target anomaly type is range anomaly, the actual road condition index values ​​of adjacent years for each road condition identifier can be obtained, and then the difference can be processed for each road condition identifier based on the actual road condition index value to obtain the target road condition index value corresponding to that road condition identifier.

[0100] The above scheme corrects the abnormal road condition index values ​​according to the target anomaly type by applying different correction methods to obtain the target road condition index values. It adaptively selects the optimal scheme from a variety of advanced interpolation methods to fill the data, thereby preserving the true decay trend and effective information contained in the data to the greatest extent and ensuring the high fidelity of the processed dataset.

[0101] In some embodiments, after determining the target road segment dataset based on all target traffic condition index values ​​and the initial road segment dataset in step 105, the method further includes: Step 10A: Input the target road segment dataset into the initial road surface performance prediction model; Step 10B: Train the initial pavement performance prediction model using the target road segment dataset until the preset training termination condition is met, and obtain the pavement performance prediction model.

[0102] In practice, the target road segment dataset is input into the initial road surface performance prediction model, and the initial road surface performance prediction model is trained using the target road segment dataset until the preset training termination condition is met, thus obtaining the road surface performance prediction model.

[0103] In this embodiment, the preset training termination condition includes at least one of the following: determining that all data in the training dataset has been input into the initial pavement performance prediction model for training, determining that the loss function of the initial pavement performance prediction model has converged to a preset convergence threshold, or determining that the initial pavement performance prediction model has been iterated for training to a preset number of iterations.

[0104] For example, the preset training termination condition is that all data in the training dataset has been input into the initial road performance prediction model for training: The training dataset contains fifty sets of data, each set including a target road segment dataset and target road condition index values. The preset training termination condition is that all data in the training dataset has been input into the initial pavement performance prediction model for training. That is, when all fifty sets of data have been input into the initial pavement performance prediction model, there is no training data in the training dataset that has not yet been input into the initial pavement performance prediction model. At this point, the initial pavement performance prediction model training is considered complete, and the pavement performance prediction model is obtained.

[0105] Another example is that the preset training termination condition is to determine that the loss function of the initial road performance prediction model converges to a preset convergence threshold: Training data from the training dataset is input into the initial pavement performance prediction model for training, and the training results are output. A loss function is determined based on the training results and the target pavement condition index value. The loss function may include at least one of the following: mean squared error loss function, cross-entropy loss function, logarithmic loss function, exponential loss function, squared loss function, or absolute value loss function, etc. When the loss function converges to a preset convergence threshold, the preset training termination condition is satisfied, and the pavement performance prediction model is obtained.

[0106] Another example is that the preset training termination condition is to determine the initial road performance prediction model to be iterated and trained to a preset number of iterations.

[0107] The training data in the training dataset is input into the initial road performance prediction model for iterative training. The number of iterations is recorded. When the number of iterations equals the preset number of iterations, the preset training termination condition is met, and the road performance prediction model is obtained.

[0108] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0109] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0110] Based on the same inventive concept, another embodiment of this disclosure proposes a road segment data anomaly detection and processing device, such as... Figure 2 As shown, it specifically includes: The data acquisition module 201 is configured to acquire an initial road segment dataset and target road condition indicators, wherein each road segment data in the initial road segment dataset includes multiple initial influencing factors and initial road condition indicator values. Clustering processing module 202 is configured to determine the target influencing factors corresponding to the target traffic condition index, and to perform clustering processing on all road segment data in the initial road segment dataset according to the target influencing factors to obtain multiple road segment groups; The anomaly detection module 203 is configured to perform anomaly detection processing on the initial traffic condition index values ​​corresponding to all road segment data within each road segment group, and determine the abnormal traffic condition index values ​​within the road segment group. Anomaly correction module 204 is configured to determine the target anomaly type corresponding to the abnormal road condition index value, and correct the abnormal road condition index value according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group. The target road segment dataset determination module 205 is configured to statistically analyze all target road condition index values ​​corresponding to all road segment groups, and determine the target road segment dataset based on the all target road condition index values ​​and the initial road segment dataset.

[0111] In some embodiments, the clustering processing module 202 is specifically configured as follows: Obtain the pavement base structure type and traffic flow level corresponding to each road segment in the initial road segment dataset; The data of all road segments are divided according to the road base structure type and the traffic flow level to obtain multiple road segment data subsets; For each road segment data subset, the target influencing factors corresponding to the target road condition index in the road segment data subset are determined. Based on the target influencing factors, all road segment data in the road segment data subset are clustered to obtain multiple road segment groups.

[0112] In some embodiments, the clustering processing module 202 is specifically configured as follows: Obtain all road segment data included in the road segment data subset, and determine multiple target feature labels based on the all road segment data, wherein each target feature label contains an initial road condition index value and an initial influencing factor; The multiple target feature labels are input into a random forest model, and after processing by the random forest model, the importance weight corresponding to each initial influencing factor is output. Obtain initial influencing factors whose importance weight is greater than a preset weight threshold, and obtain a candidate factor sequence matrix; The candidate factor sequence matrix is ​​input into the principal component analysis model, and after processing by the principal component analysis model, the target influencing factors are output.

[0113] In some embodiments, the anomaly detection module 203 is specifically configured as follows: For each road segment group, statistical analysis and processing are performed on the initial road condition index values ​​corresponding to all road segment data within the road segment group to obtain the first abnormal road condition index value. Take all initial road condition index values ​​except the first abnormal road condition index value as the first initial road condition index value, obtain the target abnormal rule corresponding to the target road condition index, and take the first initial road condition index value that satisfies the target abnormal rule as the second abnormal road condition index value. The other initial road condition index values ​​in the first initial road condition index value, except for the second abnormal road condition index value, are used as the second initial road condition index value. The second initial road condition index value is input into the anomaly identification model, processed by the anomaly identification model, and the third abnormal road condition index value is output. The abnormal road condition index values ​​of the first abnormal road condition index, the second abnormal road condition index, and the third abnormal road condition index are summarized to obtain the abnormal road condition index values ​​within the road segment group.

[0114] In some embodiments, the anomaly detection module 203 is specifically configured as follows: Obtain the initial road condition index values ​​corresponding to all road segment data within the road segment group, and determine the target mean and target standard deviation corresponding to all initial road condition index values; The target mean is summed with the target standard deviation by a preset multiple to obtain the maximum road condition index value; The minimum road condition index value is obtained by subtracting the target mean from the target standard deviation by a preset multiple. In response to the existence of an initial road condition index value that is greater than the maximum road condition index value or less than the minimum road condition index value, the initial road condition index value is determined to be the first abnormal road condition index value.

[0115] In some embodiments, the anomaly detection module 203 is specifically configured as follows: The second initial road condition index value is input into the isolated forest model, and after processing by the isolated forest model, the first anomaly score corresponding to the second initial road condition index value is output. The second initial road condition index value is input into the outlier factor model, and after processing by the outlier factor model, the second anomaly score corresponding to the second initial road condition index value is output. The first abnormal score and the second abnormal score are weighted to obtain the target abnormal score corresponding to the second initial road condition index value; In response to the target abnormal score being greater than a preset score threshold, the second initial road condition index value is determined to be the third abnormal road condition index value.

[0116] In some embodiments, the anomaly correction module 204 is specifically configured as follows: Obtain the number of index values ​​for the abnormal road condition index; In response to the fact that the number of the indicator values ​​is one, the target anomaly type corresponding to the abnormal road condition indicator value is determined to be a local anomaly; or, In response to the fact that there are multiple index values, the road segment identifier of the road segment data corresponding to each abnormal road condition index value is obtained, and the target abnormality type corresponding to the abnormal road condition index value corresponding to the consecutive road segment identifiers that are greater than a preset number threshold is determined to be a range abnormality, and the target abnormality type corresponding to other abnormal road condition index values ​​other than those included in the range abnormality is determined to be a local abnormality.

[0117] In some embodiments, the anomaly correction module 204 is specifically configured as follows: In response to the target anomaly type being a local anomaly, a target completion index value is determined according to a preset data completion algorithm, the abnormal road condition index value is deleted, and the target completion index value is used as the target road condition index value corresponding to the road segment group; or... In response to the target anomaly type being a range anomaly, the target year corresponding to the abnormal road condition index values ​​included in the range anomaly is obtained, the actual road condition index values ​​of the road segment identifiers corresponding to the range anomaly in the adjacent years of the target year are obtained, and the abnormal road condition index values ​​are corrected based on the actual road condition index values ​​to obtain the target road condition index values ​​corresponding to the road segment group.

[0118] In some embodiments, the apparatus further includes a model training module, which is specifically configured to: The target road segment dataset is input into the initial road surface performance prediction model; The initial pavement performance prediction model is trained using the target road segment dataset until the preset training termination condition is met, thus obtaining the pavement performance prediction model.

[0119] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0120] The apparatus described above is used to implement the corresponding road segment data anomaly detection and processing method in any of the following embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0121] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the road segment data anomaly detection and processing method described in any of the above embodiments.

[0122] Figure 3 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0123] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0124] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0125] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0126] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0127] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0128] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0129] The electronic devices described above are used to implement the corresponding road segment data anomaly detection and processing methods in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0130] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the road segment data anomaly detection and processing method as described in any of the above embodiments.

[0131] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0132] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the road segment data anomaly detection and processing method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0133] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0134] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0135] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0136] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0137] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0138] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0139] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0140] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for detecting and processing road segment data anomalies, characterized in that, include: Obtain an initial road segment dataset and target road condition indicators, wherein each road segment in the initial road segment dataset includes multiple initial influencing factors and initial road condition indicator values; Determine the target influencing factors corresponding to the target road condition index, and perform clustering processing on all road segment data in the initial road segment dataset based on the target influencing factors to obtain multiple road segment groups; For each road segment group, anomaly detection processing is performed on the initial road condition index values ​​corresponding to all road segment data within the road segment group to determine the abnormal road condition index values ​​within the road segment group. Determine the target anomaly type corresponding to the abnormal road condition index value, and correct the abnormal road condition index value according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group. Statistically calculate all target road condition index values ​​corresponding to all road segment groups, and determine the target road segment dataset based on the all target road condition index values ​​and the initial road segment dataset; For each road segment group, anomaly detection processing is performed on the initial traffic condition index values ​​corresponding to all road segment data within the road segment group to determine abnormal traffic condition index values ​​within the road segment group, including: For each road segment group, statistical analysis and processing are performed on the initial road condition index values ​​corresponding to all road segment data within the road segment group to obtain the first abnormal road condition index value. Take all initial road condition index values ​​except the first abnormal road condition index value as the first initial road condition index value, obtain the target abnormal rule corresponding to the target road condition index, and take the first initial road condition index value that satisfies the target abnormal rule as the second abnormal road condition index value. The other initial road condition index values ​​in the first initial road condition index value, except for the second abnormal road condition index value, are used as the second initial road condition index value. The second initial road condition index value is input into the anomaly identification model, processed by the anomaly identification model, and the third abnormal road condition index value is output. By summing the first abnormal road condition index value, the second abnormal road condition index value, and the third abnormal road condition index value, the abnormal road condition index value within the road segment group is obtained. The step of statistically analyzing and processing the initial road condition index values ​​corresponding to all road segment data within the road segment group to obtain the first abnormal road condition index value includes: Obtain the initial road condition index values ​​corresponding to all road segment data within the road segment group, and determine the target mean and target standard deviation corresponding to all initial road condition index values; The target mean is summed with the target standard deviation by a preset multiple to obtain the maximum road condition index value; The minimum road condition index value is obtained by subtracting the target mean from the target standard deviation by a preset multiple. In response to the existence of an initial road condition index value that is greater than the maximum road condition index value or less than the minimum road condition index value, the initial road condition index value is determined to be the first abnormal road condition index value.

2. The method according to claim 1, characterized in that, The step involves determining the target influencing factors corresponding to the target road condition index, and then clustering all road segment data in the initial road segment dataset based on the target influencing factors to obtain multiple road segment groups, including: Obtain the pavement base structure type and traffic flow level corresponding to each road segment in the initial road segment dataset; The data of all road segments are divided according to the road base structure type and the traffic flow level to obtain multiple road segment data subsets; For each road segment data subset, the target influencing factors corresponding to the target road condition index in the road segment data subset are determined. Based on the target influencing factors, all road segment data in the road segment data subset are clustered to obtain multiple road segment groups.

3. The method according to claim 2, characterized in that, The determination of the target influencing factors corresponding to the target road condition index in the road segment data subset includes: Obtain all road segment data included in the road segment data subset, and determine multiple target feature labels based on the all road segment data, wherein each target feature label contains an initial road condition index value and an initial influencing factor; The multiple target feature labels are input into a random forest model, and after processing by the random forest model, the importance weight corresponding to each initial influencing factor is output. Obtain initial influencing factors whose importance weight is greater than a preset weight threshold, and obtain a candidate factor sequence matrix; The candidate factor sequence matrix is ​​input into the principal component analysis model, and after processing by the principal component analysis model, the target influencing factors are output.

4. The method according to claim 1, characterized in that, The anomaly detection models include the isolated forest model and the outlier factor model. The step of inputting the second initial road condition index value into the anomaly identification model, processing it through the anomaly identification model, and outputting the third abnormal road condition index value includes: The second initial road condition index value is input into the isolated forest model, and after processing by the isolated forest model, the first anomaly score corresponding to the second initial road condition index value is output. The second initial road condition index value is input into the outlier factor model, and after processing by the outlier factor model, the second anomaly score corresponding to the second initial road condition index value is output. The first abnormal score and the second abnormal score are weighted to obtain the target abnormal score corresponding to the second initial road condition index value; In response to the target abnormal score being greater than a preset score threshold, the second initial road condition index value is determined to be the third abnormal road condition index value.

5. The method according to claim 1, characterized in that, Determining the target anomaly type corresponding to the abnormal road condition index value includes: Obtain the number of index values ​​for the abnormal road condition index; In response to the fact that the number of the indicator values ​​is one, the target anomaly type corresponding to the abnormal road condition indicator value is determined to be a local anomaly; or, In response to the fact that there are multiple index values, the road segment identifier of the road segment data corresponding to each abnormal road condition index value is obtained, and the target abnormality type corresponding to the abnormal road condition index value corresponding to the consecutive road segment identifiers that are greater than a preset number threshold is determined to be a range abnormality, and the target abnormality type corresponding to other abnormal road condition index values ​​other than those included in the range abnormality is determined to be a local abnormality.

6. The method according to claim 5, characterized in that, The step of correcting the abnormal road condition index value according to the target anomaly type to obtain the target road condition index value corresponding to the road segment group includes: In response to the target anomaly type being a local anomaly, a target completion index value is determined according to a preset data completion algorithm, the abnormal road condition index value is deleted, and the target completion index value is used as the target road condition index value corresponding to the road segment group; or... In response to the target anomaly type being a range anomaly, the target year corresponding to the abnormal road condition index values ​​included in the range anomaly is obtained, the actual road condition index values ​​of the road segment identifiers corresponding to the range anomaly in the adjacent years of the target year are obtained, and the abnormal road condition index values ​​are corrected based on the actual road condition index values ​​to obtain the target road condition index values ​​corresponding to the road segment group.

7. The method according to claim 1, characterized in that, After determining the target road segment dataset based on all the target traffic condition index values ​​and the initial road segment dataset, the process also includes: The target road segment dataset is input into the initial road surface performance prediction model; The initial pavement performance prediction model is trained using the target road segment dataset until the preset training termination condition is met, thus obtaining the pavement performance prediction model.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 7.