Cost risk grading early warning and dynamic adjustment method based on deviation degree clustering analysis
Patent Information
- Application Number
- CN202611255925.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-18
AI Technical Summary
[0002]环保EPC工程项目普遍具有非标设备占比高、工艺系统复杂、设计变更频繁、耗材价格波动大的特点,造价管控难度远高于常规土建工程
[0044](1) This application adopts a global clustering analysis architecture to automatically classify all cost risks of the project according to the deviation pattern. It can not only identify individual isolated risk points, but also discover batch and systemic cost risks caused by the same type of reasons. For group risks such as the general increase in the price of consumables for the whole project and the superposition of multiple sub-item design changes, they can be captured in the early stage through cluster density characteristics, which allows sufficient time for batch control and early response. It fundamentally makes up for the shortcomings of traditional single-point assessment that cannot identify systemic risks, and greatly enhances the foresight and comprehensiveness of risk control.
Smart Images

Figure CN122779901A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of engineering cost risk technology, and in particular relates to a method for cost risk classification, early warning and dynamic adjustment based on deviation clustering analysis. Background Technology
[0002] Environmental protection EPC projects are generally characterized by a high proportion of non-standard equipment, complex process systems, frequent design changes, and large fluctuations in consumable prices, making cost control far more difficult than in conventional civil engineering projects. Existing engineering cost risk management technologies are mainly divided into three categories: The first category relies on 3D models and dynamic simulations to independently calculate the probability and impact of individual risks and then classify them into levels. This is highly dependent on model accuracy and cannot identify group-based systemic risks. The second category constructs multiple indicators for specific industries, calculates single-point composite deviations through deep learning, and updates the benchmark globally. However, this approach has poor adaptability to different risk types. The third category uses fixed indicator weighted scoring, combined with statistical methods to screen for single-point anomalies. However, the indicator dimensions are fixed and there is no dynamic optimization mechanism for classification.
[0003] Overall, existing technologies all employ an underlying architecture of independent quantification of single risk points, hierarchical control of single points, and unified global parameter tuning. This architecture suffers from three major bottlenecks: First, it focuses solely on single-point numerical deviations, failing to identify group-wide and systemic cost control risks. Second, it directly classifies risks based on single-point values, neglecting the overall deviation characteristics of risk clusters and the group transmission effect. Third, the globally unified correction of model parameters cannot accurately adapt to the evolution patterns of different types of risks, leading to a continuous decline in model accuracy over long-term operation. These shortcomings are particularly pronounced in environmental EPC projects with strong non-standard characteristics and high change frequency, necessitating a breakthrough at the underlying architecture level. Summary of the Invention
[0004] This application addresses the problems existing in the prior art by proposing a cost risk classification, early warning, and dynamic adjustment method based on deviation clustering analysis. It constructs a complete technical system from five levels: risk feature extraction, global clustering and classification, two-layer precise classification, differentiated early warning push, and cluster adaptive iteration. This completely abandons the rigid model of traditional single-point independent assessment and globally unified parameter adjustment, achieving group identification, precise classification, differentiated control, and adaptive optimization of cost risks. It comprehensively improves the comprehensiveness, accuracy, and long-term operational stability of cost risk management in complex engineering projects, and is particularly suitable for environmental engineering EPC projects with numerous non-standard equipment, frequent changes, and complex cost fluctuation patterns.
[0005] To achieve the above objectives, this application provides the following technical solution:
[0006] A method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis includes the following steps: S1. Collect original multi-dimensional cost indicators throughout the entire project lifecycle, and calculate the relative entropy deviation of each cost risk object's corresponding dimension using the industry benchmark probability distribution as a reference, obtaining the multi-dimensional relative entropy deviation vector corresponding to each risk object; S2. Summarize the multi-dimensional relative entropy deviation vectors of all cost risk objects, and perform global clustering using an adaptive truncation distance density peak clustering algorithm to obtain multiple risk clusters with similar deviation characteristics, and identify outlier extreme risk points; S3. Calculate the comprehensive risk entropy value of each risk cluster, and based on the comprehensive risk... The entropy value determines the basic early warning level of the corresponding risk cluster, and the individual early warning level is fine-tuned by combining the vector differences of single risk objects within the cluster to obtain the final graded early warning result for each cost risk object; S4, according to the final graded early warning result of each risk cluster and outlier extreme risk point, the differentiated early warning push path and risk handling strategy are matched, graded early warning instructions and corresponding handling plans are generated and issued for execution; S5, the execution feedback and handling effect data of the corresponding handling plan for each risk cluster are collected, an independent state transition model is constructed for each risk cluster, and cluster iterative optimization is initiated with preset trigger conditions to correct the clustering model parameters and realize the dynamic adjustment of cost risk control.
[0007] Optionally, step S1 includes:
[0008] S11. Retrieve the industry benchmark probability distribution corresponding to each dimension of cost indicators, and map the single-dimensional actual indicators of each risk object to the actual probability distribution value.
[0009] S12. Calculate the distribution difference contribution of positive deviation and negative deviation respectively, and combine them to obtain the relative entropy deviation of this dimension;
[0010] S13. Concatenate the relative entropy deviations of all dimensions of a single risk object in a preset order to form a multidimensional relative entropy deviation vector of the risk object.
[0011] Optionally, step S2 includes:
[0012] S21. Calculate the Euclidean distance between each pair of the multidimensional relative entropy deviation vectors of all risk objects, and construct a distance set;
[0013] S22. Sort the distance set in ascending order of numerical values, and take the distance value corresponding to the preset quantile after sorting as the cutoff distance.
[0014] Optionally, step S2 further includes:
[0015] S23. Calculate the local density of each risk object, count the number of adjacent risk objects whose distance is less than the cutoff distance, and use this as the local density of the risk object.
[0016] S24. Calculate the relative distance of each risk object. For the risk object with the highest local density, take the maximum distance between all objects as its relative distance; for the other risk objects, take the minimum distance between them and all objects with higher local density as their relative distance.
[0017] S25. Determine the cluster center and complete the classification. Based on the product of local density and relative distance, select risk objects whose product is greater than a preset threshold as cluster centers. The remaining risk objects are assigned to the risk cluster corresponding to the nearest cluster center. Risk objects whose local density is less than a preset density threshold and whose relative distance is greater than a preset distance threshold are identified as outlier extreme risk points.
[0018] Optionally, step S3 includes:
[0019] S31. Calculate the mean deviation, intra-cluster dispersion, and inter-cluster transmission coefficient for each individual risk cluster. The mean deviation is the average value of the vector magnitudes of all risk objects within the cluster, the intra-cluster dispersion is the standard deviation of the vector magnitudes within the cluster, and the inter-cluster transmission coefficient is the average vector similarity between the risk cluster and all other risk clusters.
[0020] S32. The objective weights corresponding to the mean deviation, intra-cluster dispersion, and inter-cluster transmission coefficient are calculated using the entropy weight method.
[0021] S33. Sum the three parameters with their corresponding weights to obtain the comprehensive risk entropy value of the risk cluster.
[0022] Optionally, in step S3, determining the cluster-based early warning level based on the comprehensive risk entropy value includes:
[0023] Two-level entropy thresholds are preset, with the lower threshold being less than the higher threshold;
[0024] If the comprehensive risk entropy value is greater than or equal to the high threshold, the cluster-based early warning level is high risk;
[0025] If the comprehensive risk entropy value is greater than or equal to the low threshold and less than the high threshold, the cluster basic warning level is medium risk.
[0026] If the comprehensive risk entropy value is less than the low threshold, the cluster basic early warning level is low risk.
[0027] Optionally, in step S3, fine-tuning the individual early warning level by combining the vector differences of single-risk objects within the cluster includes:
[0028] Calculate the ratio of the vector magnitude of a single risk object to the mean of the average deviation of its risk cluster;
[0029] If the ratio is greater than or equal to 1.2, the individual early warning level of the risk object will be increased by one level above the cluster base level;
[0030] If the ratio is less than or equal to 0.8, the individual early warning level of the risk object will be adjusted down one level from the cluster base level.
[0031] After the adjustment, the highest level is high risk, and after the adjustment, the lowest level is low risk.
[0032] Optionally, step S4 includes:
[0033] High-risk clusters and outlier extreme risk points are pushed to the project decision-making level terminal and trigger the emergency response process, while simultaneously generating batch disposal plans.
[0034] For medium-risk clusters, push them to the cost management department's terminal and match them with the corresponding optimized handling plan for the risk type;
[0035] For low-risk clusters, only periodic summary reports are generated, and routine monitoring is performed.
[0036] Optionally, step S5 includes:
[0037] S51. Define a set of risk states for a single risk cluster, including three types of states: low risk, medium risk, and high risk.
[0038] S52. Based on the historical execution data of the corresponding disposal plan for the risk cluster, the frequency of transitions between different states is counted, and an initial state transition probability matrix is constructed, with the matrix elements corresponding to the probability of transitioning from one state to another.
[0039] S53. Introduce the treatment feedback gain coefficient and modify the state transition probability matrix in combination with the actual treatment effectiveness of the treatment plan to obtain the state transition probability matrix after intervention.
[0040] Optionally, in step S5, the parameters for clustering iterative optimization include:
[0041] The weights of each dimension of the multidimensional relative entropy deviation, the quantile of the cluster cutoff distance, and the entropy weight coefficient of the comprehensive risk entropy value;
[0042] The iteration rule is as follows: if the risk level prediction deviation rate of a single risk cluster exceeds the preset threshold for three consecutive periods, only the model parameters corresponding to that risk cluster will be corrected, while the parameters of the other risk clusters will remain unchanged.
[0043] The beneficial effects of this application are as follows:
[0044] (1) This application adopts a global clustering analysis architecture to automatically classify all cost risks of the project according to the deviation pattern. It can not only identify individual isolated risk points, but also discover batch and systemic cost risks caused by the same type of reasons. For group risks such as the general increase in the price of consumables for the whole project and the superposition of multiple sub-item design changes, they can be captured in the early stage through cluster density characteristics, which allows sufficient time for batch control and early response. It fundamentally makes up for the shortcomings of traditional single-point assessment that cannot identify systemic risks, and greatly enhances the foresight and comprehensiveness of risk control.
[0045] (2) The two-tiered classification model constructed in this application, which combines cluster-level basic classification with individual fine-tuning, first assesses the average deviation level, dispersion, and transmission effect from the overall dimension of the risk cluster to determine the basic risk level, fully considering the group transmission and chain reaction effects of risks; then, it performs individual fine-tuning based on the actual deviation of individual risks to ensure that the classification results not only conform to the overall attributes of the risk type but also match the specific situation of a single point. Compared with the traditional method of directly classifying based on the numerical value of a single point, the matching degree between the classification results and the actual severity of risks is significantly improved, providing a more reliable basis for subsequent resource allocation.
[0046] (3) This application matches differentiated push paths and handling resources for different levels of risk. High-risk risks are directly pushed to the decision-making level and trigger an emergency response, medium-risk risks are pushed to professional management departments, and low-risk risks are only summarized periodically. This mechanism effectively avoids the problem of redundant and excessive early warning information, allowing managers to focus their energy on handling core risks, and control resources are precisely tilted towards high-risk links. The overall risk handling response efficiency and resource utilization efficiency are significantly improved, avoiding the waste of management resources caused by traditional indiscriminate early warning.
[0047] (4) This application uses relative entropy to measure the degree of deviation of cost indicators. Compared with the traditional linear deviation rate, it can more profoundly reflect the differences at the probability distribution level. It is especially suitable for indicators such as cost risk that are non-linear fluctuations, and avoids the structural bias caused by the insufficient sensitivity of traditional deviation calculation to early distribution anomalies. At the same time, the time-series deviation increment dimension is introduced to realize the fusion representation of static state and dynamic trend, which further improves the early risk identification capability and provides a more reliable feature basis for subsequent clustering and classification, thereby improving the accuracy of the entire control system from the source. Attached Figure Description
[0048] Figure 1 This is a schematic diagram illustrating the application of traditional engineering cost risk management processes.
[0049] Figure 2 This is a schematic diagram of the cost risk classification, early warning, and dynamic adjustment method based on deviation clustering analysis in Embodiment 1 of this application; Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely one preferred embodiment of this application and are only used to explain this application. They do not limit the scope of protection of this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0051] like Figure 1 The diagram illustrates a scenario application of traditional engineering cost risk management processes. Existing technologies generally employ a single-point identification, independent assessment, unified early warning, and global parameter adjustment management architecture. Its execution process is as follows: first, individual cost risks are identified one by one; then, each risk point is independently quantitatively assessed and its level determined; and dynamic adjustments to cost risk management are made based on execution feedback and handling effects. The model optimization stage uses a globally unified parameter adjustment feedback mechanism, with all risks sharing a single set of parameters for synchronous correction. This architecture can only handle isolated single-point risks, cannot identify group-based systemic risks, and suffers from insufficient accuracy in classification and model adaptability, making it difficult to meet the cost management needs of complex environmental EPC projects.
[0052] Example 1:
[0053] like Figure 2 As shown, a method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis includes the following steps:
[0054] S1. Collect original cost indicators of multiple dimensions throughout the entire life cycle of the project, and calculate the relative entropy deviation of each cost risk object in the corresponding dimension with reference to the industry benchmark probability distribution, so as to obtain the multi-dimensional relative entropy deviation vector corresponding to each risk object.
[0055] S2. Summarize the multidimensional relative entropy deviation vectors of all cost risk objects, and use the density peak clustering algorithm with adaptive truncation distance to perform global clustering to obtain multiple risk clusters with similar deviation characteristics, and identify and obtain outlier extreme risk points.
[0056] S3. Calculate the comprehensive risk entropy value of each risk cluster, determine the basic early warning level of the corresponding risk cluster based on the comprehensive risk entropy value, and fine-tune the individual early warning level by combining the vector difference of single risk objects within the cluster to obtain the final graded early warning result of each cost risk object.
[0057] S4. Based on the final graded early warning results of each risk cluster and outlier extreme risk point, match differentiated early warning push paths and risk handling strategies, generate graded early warning instructions and corresponding handling plans, and issue them for execution.
[0058] S5. Collect the execution feedback and handling effect data of the corresponding handling plan for each risk cluster, build an independent state transition model for each risk cluster, start cluster iterative optimization with preset trigger conditions, correct the clustering model parameters, and realize dynamic adjustment of cost risk control.
[0059] In step S1, the multi-dimensional cost indicators include six core dimensions: equipment, civil engineering, installation, consumables, changes, and schedule. Each major category can be further refined into specific sub-indicators based on the project type. For environmental protection EPC projects, equipment indicators include deviations in the procurement cost of core non-standard equipment, deviations in the procurement cost of auxiliary equipment, and deviations in tariffs and logistics for imported components; civil engineering indicators include the overrun of the main civil engineering work, additional costs for foundation treatment, and the proportion of on-site visa costs; installation indicators include deviations in installation labor costs, deviations in commissioning costs, and deviations in special operation costs; consumables indicators include the price fluctuation range of process consumables and deviations in consumable usage; changes indicators include the proportion of design change amounts, change frequency, and cost losses caused by delays in change approval; and schedule indicators include the cascading cost losses caused by delays in key milestones and the risk amount of schedule penalty.
[0060] All raw indicator data is automatically collected through the project cost management system, procurement management system, and engineering management system, while also supporting manual data entry. The collected raw data undergoes a cleaning process to remove outliers, missing values, and duplicate data. Indicators with different dimensions are preliminarily normalized to ensure all indicators are on the same numerical scale, preventing dimensional imbalances in subsequent calculations. The preprocessed data is then grouped according to risk objects, with each risk object corresponding to a set of multi-dimensional raw indicator data, forming a structured risk indicator dataset.
[0061] The data collection cycle is dynamically adjusted according to project phases: data is collected every 15 days during the design and procurement phase, every 7 days during peak construction periods, and every 30 days during the final settlement phase. Each batch of raw data undergoes outlier removal (using the 3σ principle), missing value imputation (using linear interpolation), and dimensional normalization (using Min-Max standardization) to ensure all indicators are on the same numerical order of magnitude. The preprocessed data is then grouped by risk object (at the granularity of individual projects or unit projects), with each risk object corresponding to a set of six-dimensional raw indicator data.
[0062] Specifically, step S1 includes:
[0063] S11. Retrieve the industry benchmark probability distributions corresponding to each dimension of cost indicators, and map the single-dimensional actual indicators of each risk object to actual probability distribution values. The industry benchmark probability distribution is fitted using the kernel density estimation method to obtain a continuous probability distribution function, which can fully reflect the central tendency, dispersion, and tail characteristics of the indicators under normal industry levels. Substitute the single-dimensional actual indicator values of a single risk object into the benchmark distribution to obtain the corresponding actual probability distribution values, providing input for subsequent relative entropy calculations.
[0064] S12. Calculate the distributional differences contributing to positive and negative deviations separately, and combine them to obtain the relative entropy deviation for this dimension. Relative entropy (KL divergence) is used to measure the degree of difference between the actual distribution and the benchmark distribution, replacing the traditional linear deviation rate calculation method. Relative entropy can measure the overall difference between two probability distributions, reflecting not only the deviation from the numerical mean but also capturing anomalies in distribution patterns, and its sensitivity to early risk signals is far higher than that of linear deviation. Specifically, the formula for calculating the single-dimensional relative entropy deviation is:
[0065]
[0066] in, Let be the relative entropy deviation of the j-th dimension indicator for the i-th risk object; The actual probability distribution value of the j-th dimension indicator of the i-th risk object is obtained by mapping the actual value of the indicator to the baseline distribution. Let be the industry benchmark probability distribution value corresponding to the j-th dimension indicator; i be the index of the risk object; and j be the index of the cost indicator dimension. This formula contains two terms: the first corresponds to the contribution of positive deviation, and the second corresponds to the contribution of negative deviation. Together, they constitute the complete relative entropy deviation. Relative entropy is non-negative. When the actual distribution is completely consistent with the benchmark distribution, the relative entropy is 0, representing no deviation; the greater the degree of deviation, the higher the relative entropy value.
[0067] S13. Concatenate the relative entropy deviations of all dimensions of a single risk object in a preset order to form a multidimensional relative entropy deviation vector for that risk object. After calculating the relative entropy of all dimensions of a single risk object, concatenate the deviation values of each dimension in a preset fixed dimensional order to form an n-dimensional vector:
[0068]
[0069] in, Let be the multidimensional relative entropy deviation vector of the i-th risk object, corresponding to a data point in the high-dimensional feature space; , where represents the relative entropy deviation value corresponding to the 1st to nth dimensions of the cost indicators for the i-th risk object; n is the total number of dimensions of the cost indicators. Each vector corresponds to a data point in the high-dimensional feature space, the magnitude of the vector reflects the overall degree of risk deviation, and the direction of the vector reflects the deviation pattern characteristics of the risk.
[0070] Furthermore, step S1 also includes updating the multi-dimensional relative entropy deviation vector of each risk object according to a preset period. Each update introduces a temporal deviation increment, incorporating the difference in relative entropy deviation between the current period and the previous period as a new dimension into the vector. This upgrades the vector from a static snapshot to a dynamic trajectory, containing both the current deviation status and evolutionary trend information. For example, in this embodiment, the initial vector is a 6-dimensional basic deviation vector, which is expanded to a 12-dimensional vector after introducing the temporal increment. The first 6 dimensions represent the relative entropy deviation of the current period, and the last 6 dimensions represent the difference in deviation between the current period and the previous period. The update cycle can be flexibly adjusted according to the project stage, with increased update frequency during peak construction periods and appropriately reduced frequency during the design and procurement stages to ensure that the data always closely reflects the real-time status of the project.
[0071] In this embodiment, step S2 includes:
[0072] S21. Calculate the Euclidean distance between each pair of the multidimensional relative entropy deviation vectors of all risk objects, and construct a distance set;
[0073] S22. Sort the distance set in ascending order of numerical values, and take the distance value corresponding to the preset percentile after sorting as the cutoff distance. The value range of the percentile is 1% to 2%.
[0074] Among them, cutoff distance Density peak clustering is a core parameter that defines the neighborhood range of a sample point and directly affects the local density calculation results and clustering effect. This application uses an adaptive quantile method to automatically determine the cutoff distance: First, the Euclidean distance between all risk objects is calculated. The Euclidean distance is the spatial distance between two multidimensional vectors, reflecting the similarity between two risk deviation patterns. The closer the distance, the more similar the patterns. After sorting all distances in ascending order, the value corresponding to the percentile is taken as the cutoff distance. The percentile value is 1% to 2%. This range has been verified by a large amount of project data and can ensure that each sample has an appropriate number of neighbors, making the density calculation stable and reliable.
[0075] By adopting an adaptive truncation distance mechanism, the clustering algorithm can be deployed with zero parameter tuning, which greatly reduces the application threshold. The truncation distance determined by the data's own distribution is more in line with the actual characteristics of the data than manual setting, and the clustering results are more accurate and stable. It can automatically adapt to projects of different sizes and types.
[0076] Specifically, step S2 also includes:
[0077] S23. Calculate the local density of each risk object, count the number of adjacent risk objects whose distance is less than the cutoff distance, and use this as the local density of the risk object.
[0078] S24. Calculate the relative distance of each risk object. For the risk object with the highest local density, take the maximum distance between all objects as its relative distance; for the other risk objects, take the minimum distance between them and all objects with higher local density as their relative distance.
[0079] Local density This reflects the density of similar risks surrounding the i-th risk object. A higher value indicates a more prevalent deviation pattern and a greater likelihood of group-wide risk. Local density is calculated using a truncated kernel function, with the following formula:
[0080]
[0081] in, Let be the local density of the i-th risk object; Let Euclidean distance be the multidimensional vector between the i-th risk object and the j-th risk object; This is the cutoff distance for the clustering algorithm, used to define the range threshold of the sample's neighborhood; To truncate the kernel function, when x < 0, the function... =1, when x≥0, the function =0. Local density is essentially the number of sample points within the cutoff distance neighborhood; the more points, the higher the density.
[0082] relative distance This reflects the distance between a high-risk object and higher-density regions. For the object with the highest local density in the entire dataset, its relative distance is the maximum distance among all samples to ensure sufficient discriminative power. For other objects, the relative distance is the minimum distance between that object and all objects with even higher local density. The larger the relative distance, the farther away the object is from other high-density regions, and the more likely it is to become the center of an independent cluster or an isolated outlier.
[0083] Therefore, by using the two dimensions of local density and relative distance, the degree of clustering and isolation of risks can be characterized simultaneously, providing a two-dimensional basis for subsequent cluster center identification and outlier detection. Compared with single-dimensional clustering features, it can more accurately divide the risk cluster structure and adapt to the irregular distribution of cost risks.
[0084] S25. Determine the cluster center and complete the classification. Based on the product of local density and relative distance, select the risk object whose product is greater than the preset threshold as the cluster center, and classify the remaining risk objects into the risk cluster corresponding to the nearest cluster center.
[0085] With local density and relative distance product This is a cluster center determination metric that considers both density and distance dimensions; samples with significantly high product values are considered cluster centers. Selection... Risk objects with values greater than a preset threshold are designated as cluster centers, and the remaining risk objects are grouped according to their nearest cluster center, thus completing the risk clustering. All samples... The values are sorted in descending order, exhibiting a clear cliff-like distribution. A threshold is automatically determined using the inflection point method; samples exceeding this threshold are identified as cluster centers, eliminating the need for manual pre-setting of cluster numbers. Local density calculation employs a truncated kernel function, where the number of neighboring sample points represents the density value. In relative distance calculation, the relative distance of the sample with the highest global density is taken as the maximum value of the Euclidean distances among all samples. Outlier extreme risk points are identified using a dual-condition joint determination: a local density below 1 / 3 of the average density of all samples, and a relative distance exceeding twice the average relative distance of all samples. Meeting both conditions simultaneously qualifies as an outlier. After determining all cluster centers, each remaining risky object is assigned to the nearest cluster center, thus belonging to its corresponding risk cluster, ultimately resulting in multiple risk clusters with similar deviation characteristics.
[0086] The clustering process does not require pre-setting the number of clusters, can automatically identify the number of risk types, is completely data-driven, is not affected by subjective assumptions, and can truly reflect the type structure of project cost risks; it has a good fit for risk clusters of any shape, breaking through the limitation of traditional K-means clustering that only fits spherical clusters, and can accurately identify risk categories with complex shapes.
[0087] Specifically, in step S2, the rule for identifying outlier extreme risk points is: for a single risk object, if its local density... Less than the preset density threshold, and the relative distance If the distance exceeds a preset threshold, the risk object is determined to be an outlier extreme risk point, classified separately, and directly triggered with the highest level of warning. During clustering, samples with extremely low local density and extremely large relative distances exhibit deviation patterns that differ from all other risks. These are low-probability, high-impact extreme risks, few in number but extremely harmful, and easily overwhelmed by the clustering process. This application uses a dual threshold rule of density and distance to identify such outliers. After determination, they are not included in any risk cluster, but are managed separately and directly marked with the highest priority.
[0088] In this embodiment, step S3 includes:
[0089] S31. Calculate the mean deviation, intra-cluster dispersion, and inter-cluster transmission coefficient for each individual risk cluster. The mean deviation is the average magnitude of the vectors of all risk objects within the cluster; the intra-cluster dispersion is the standard deviation of the magnitudes of all vectors within the cluster; and the inter-cluster transmission coefficient is the average vector similarity between this risk cluster and all other risk clusters. The comprehensive risk entropy value integrates these three dimensions of overall cluster characteristics to fully reflect the severity of the risk cluster.
[0090] First, the mean deviation , that is, the average value of all vector magnitudes in the k-th risk cluster, which reflects the overall deviation level of this type of risk. The higher the average value, the more serious the overall deviation from the benchmark.
[0091] Second, intra-cluster dispersion The standard deviation of the magnitudes of all vectors within a cluster reflects the volatility and uncertainty of this type of risk. The higher the dispersion, the greater the difference in risk performance and the stronger the evolutionary uncertainty.
[0092] Thirdly, the inter-cluster transmission coefficient The coefficient is the mean cosine similarity between the risk cluster and all other risk clusters. It reflects the degree of correlation between this type of risk and other categories. The higher the coefficient, the easier it is to trigger cross-category chain reactions and the stronger the transmission effect.
[0093] S32. The objective weights corresponding to the mean deviation, intra-cluster dispersion, and inter-cluster transmission coefficient are calculated using the entropy weight method. To avoid subjective bias in weighting, the entropy weight method is used to objectively calculate the weights of the three parameters. The core principle of the entropy weight method is that the higher the degree of data dispersion, the greater the amount of effective information contained in the indicator, and the higher the corresponding weight. During calculation, the three parameters are first standardized, then the information entropy of each parameter is calculated separately, and finally the weight coefficients are derived from the information entropy. The specific calculation process is as follows:
[0094] Construct the original evaluation matrix, assuming there are m risk clusters, each cluster corresponding to three parameter values (mean deviation, intra-cluster dispersion, and inter-cluster transmission coefficient), forming an m×3 matrix. , where i=1,…,m; j=1,2,3 correspond to the three parameters mentioned above.
[0095] Normalization is performed by using the range method to normalize each parameter j, resulting in standardized values: ,in, , These are the maximum and minimum values of the j-th parameter across all risk clusters. Since all parameters are positive indicators (higher values indicate higher risk), a uniform positive normalization formula is used.
[0096] Calculate the proportion matrix, and for each parameter j, calculate the proportion of the i-th cluster with respect to that parameter. .
[0097] The information entropy of the j-th parameter is calculated using the following formula:
[0098]
[0099] in, Let m be the information entropy of the j-th evaluation parameter; m be the total number of risk clusters. The value of the i-th risk cluster on the j-th parameter is the proportion of the sum of that parameter for all clusters. This is the normalization coefficient. When hour, The value is 0.
[0100] The weights are calculated using the following formula:
[0101]
[0102] in, Let be the objective weight corresponding to the j-th evaluation parameter. The sum of the three weights is 1. The lower the information entropy, the higher the dispersion and the stronger the discriminative ability of the parameter among the clusters, and the higher the corresponding weight, thus avoiding the human bias introduced by subjective weighting.
[0103] S33. The three parameters are summed with their corresponding weights to obtain the comprehensive risk entropy value of the risk cluster. The formula for calculating the comprehensive risk entropy value is:
[0104]
[0105] in, The mean deviation of the k-th risk cluster is the average value of the vector magnitude of all risk objects within the cluster. Let be the intra-cluster dispersion of the k-th risk cluster, which is the standard deviation of the magnitudes of all risk object vectors within the cluster; Let be the inter-cluster transmission coefficient of the k-th risk cluster, which is the average vector similarity between this cluster and all other risk clusters; This represents the overall risk entropy value of the k-th risk cluster; a larger value indicates a higher overall risk level for the cluster. The weighting coefficient corresponding to the mean of the average deviation; These are the weighting coefficients corresponding to the intra-cluster dispersion. is the weighting coefficient corresponding to the inter-cluster transmission coefficient; k is the index of the risk cluster.
[0106] Specifically, in step S3, determining the cluster-based early warning level based on the comprehensive risk entropy value includes:
[0107] Two-level entropy thresholds are preset, with the lower threshold being less than the higher threshold;
[0108] If the comprehensive risk entropy value is greater than or equal to the high threshold, the cluster-based early warning level is high risk;
[0109] If the comprehensive risk entropy value is greater than or equal to the low threshold and less than the high threshold, the cluster basic warning level is medium risk.
[0110] If the comprehensive risk entropy value is less than the low threshold, the cluster basic early warning level is low risk.
[0111] Preset two-level entropy threshold (Low threshold) and (High Threshold) The threshold is determined based on historical project data in the industry and corresponds to the dividing point of cost impact for different risk levels. Through threshold mapping, continuous comprehensive risk entropy values are divided into three basic levels: high-risk, medium-risk, and low-risk. The high-risk cluster represents severe overall deviation, large fluctuations, and strong transmission, which is highly likely to cause significant cost overruns; the medium-risk cluster represents some deviation with the potential for further deterioration; and the low-risk cluster represents the overall situation within the normal fluctuation range, with controllable risk. The threshold is determined based on industry data, ensuring the scientific and universal nature of the level classification. The level determination standards are unified across different projects, providing horizontal comparability. High-deviation, high-transmission risk clusters have a comprehensive risk entropy value greater than the high threshold, and are classified as high-risk; deteriorating trend risk clusters fall between the two threshold levels, and are classified as medium-risk; the other two risk clusters are below the low threshold, and are classified as low-risk.
[0112] Specifically, in step S3, fine-tuning the individual warning level includes:
[0113] Calculate the ratio of the vector magnitude of a single risk object to the mean of the average deviation of its risk cluster;
[0114] If the ratio is greater than or equal to 1.2, the individual early warning level of the risk object will be increased by one level above the cluster base level;
[0115] If the ratio is less than or equal to 0.8, the individual early warning level of the risk object will be adjusted down one level from the cluster base level.
[0116] After the adjustment, the highest level is high risk, and after the adjustment, the lowest level is low risk.
[0117] The cluster baseline level represents the overall risk profile for that type of risk. However, different risks within the same cluster exhibit varying degrees of deviation. Therefore, individual fine-tuning is employed to account for these individual differences. The ratio of the vector magnitude to the cluster mean is used. As a basis for fine-tuning, a ratio greater than 1.2 indicates a significantly higher deviation than the cluster average, resulting in an upward adjustment by one level; a ratio less than 0.8 indicates a significantly lower deviation than the cluster average, resulting in a downward adjustment by one level. Simultaneously, level boundary constraints are set, with the highest level not exceeding high risk and the lowest level not falling below low risk, maintaining the simplicity of the three-level architecture.
[0118] Through the individual fine-tuning mechanism, a classification effect of unified categories and individual differences is achieved. This ensures that the classification of the same type of risk is consistent with the overall attributes, while avoiding individual misjudgments caused by a one-size-fits-all approach. It avoids wasting resources on minor risks within high-risk clusters and also avoids missing serious risks within low-risk clusters, thus greatly improving the accuracy of classification.
[0119] In this embodiment, step S4 involves accurately pushing out tiered early warnings and matching them with corresponding response strategies, transforming the tiered results into actionable management actions. Setting differentiated push paths for different risk levels specifically includes:
[0120] High-risk clusters and outlier extreme risk points are pushed to the project decision-making level terminal and trigger the emergency response process, and batch disposal plans are generated simultaneously. The emergency response process includes the time limit of telephone notification within 15 minutes and online emergency meeting within 2 hours to ensure that major risks are brought into the decision-making process as soon as possible.
[0121] For medium-risk clusters, push them to the cost management department's terminal and match them with the corresponding optimized handling plan for the risk type;
[0122] For low-risk clusters, only periodic summary reports are generated, and routine monitoring is performed.
[0123] The early warning instructions adopt a standardized format, including elements such as risk number, level, description, handling suggestions, responsible party, and feedback deadline. They are issued through the project management system's workflow engine, automatically linking to responsible parties to form pending tasks. Simultaneously, a risk handling strategy library is pre-built, storing mature handling plans categorized by risk type. During matching, the business semantic tags of the risk cluster are first read, and a set of handling plans for the same type of risk is retrieved from the strategy library. Then, the corresponding intensity of handling measures is selected based on the risk level, ultimately generating standardized handling plan text. High-risk risks are matched with strong handling measures, medium-risk risks with conventional optimized plans, and similar risks generate batch handling plans, significantly improving plan generation efficiency. The entire handling execution process is traceable, automatically collecting handling effect data to form a complete closed loop of early warning, handling, and feedback. Through differentiated and tiered push notifications, management resources are precisely tilted towards high-risk areas, avoiding information redundancy and resource waste caused by indiscriminate early warnings. High-risk risks directly reach the decision-making level to ensure rapid response, while low-risk risks simplify management and free up management energy, significantly improving overall management efficiency and resource utilization efficiency. Batch handling plans are adapted to clustering results, making management efficiency far higher than the traditional model of handling each risk individually.
[0124] The early warning instructions are issued in a JSON structured format, including fields such as risk number, early warning level, risk description, handling strategy ID, responsible entity, and feedback time limit. The risk handling strategy library is pre-built according to a two-dimensional index of risk type tags and levels. High-risk levels are matched with three levels of mandatory measures: freezing related payments, initiating a special audit, and re-inquiring about prices. Medium-risk levels are matched with routine measures such as optimizing contract terms and adjusting procurement plans. Low-risk levels are matched with only routine measures such as strengthening monitoring and inclusion in the monthly tracking ledger.
[0125] In this embodiment, step S5 includes:
[0126] S51. Define a risk state set for a single risk cluster, including three states: low risk (L), medium risk (M), and high risk (H); define an independent state set for each risk cluster, with each of the three states corresponding one-to-one with the warning level, and define the state division boundary and the threshold of the comprehensive risk entropy value. Maintain consistency to ensure that the state definition is fully matched with the hierarchical system, providing a unified state benchmark for state transition analysis.
[0127] S52. Based on the historical execution data of the corresponding handling plan for this risk cluster, calculate the frequency of transitions between different states, construct an initial state transition probability matrix, where each matrix element corresponds to the probability of transitioning from one state to another; retrieve the warning level records for all risk objects within this risk cluster for the most recent 12 periods, and calculate the number of times each period transitions from state a in period t to state b in period t+1. , recorded as For each initial state 'a', calculate the sum of the frequencies of all possible transition targets, and the initial transition probability is calculated as follows:
[0128]
[0129] in, Let be the initial transition probability of the risk moving from state a to state b; This represents the frequency of risk transitions from state a to state b in historical data; 'a' represents the risk state before the transition, categorized as low, medium, or high risk; 'b' represents the risk state after the transition, also categorized as low, medium, or high risk. (All) An initial state transition probability matrix P is constructed, consisting of 3×3 elements, with the sum of each row being 1. When historical data is insufficient, the initial matrix adopts industry experience values: 0.70 for high risk maintaining high probability, 0.25 for medium risk deterioration probability, and 0.15 for low risk deterioration probability. These values are then adjusted periodically using a Bayesian update method as actual data accumulates.
[0130] For a specific state transition scenario (e.g., a→b), all historically executed handling plans are retrospectively analyzed. If no transition to a higher-level state occurs in the next cycle after the handling, the handling is considered effective. Effectiveness rate = Number of effective cases / Total number of handling cases in this scenario. When sufficient historical data has not yet been accumulated for handling plans, the initial effectiveness rate is taken from industry experience values (0.65 for high-risk→medium-risk scenarios, 0.70 for medium-risk→low-risk scenarios, and 0.50 for other scenarios), and is adjusted periodically as actual execution data accumulates.
[0131] S53. Introducing a treatment feedback gain coefficient, and combining it with the actual treatment effectiveness of the treatment plan, the state transition probability matrix is corrected to obtain the state transition probability matrix after intervention. To quantify the intervention effect of treatment measures on risk evolution, a treatment feedback gain coefficient is introduced, and the initial transition matrix is corrected by combining it with historical statistical treatment effectiveness. The calculation formula is as follows:
[0132]
[0133] in, The modified transition probability of risk shifting from state a to state b after intervention is introduced; Let be the initial transition probability of risk from state a to state b under no-intervention conditions; The feedback gain coefficient is used to adjust the magnitude of the impact of the treatment effect on the transition probability. The effectiveness rate of the response plan in the corresponding state transition scenario is the proportion of times that the response measures can prevent the state transition from occurring. After the correction, the probability of a state transition corresponding to an effective response will decrease accordingly. The model can more accurately reflect the impact of control intervention on risk evolution and is more in line with the actual control scenarios of the project.
[0134] By correcting the transfer matrix through feedback, the actual effectiveness of the response measures can be quantitatively evaluated, providing a quantitative basis for optimizing response plans and iterating parameters, thus enabling the control effect to be measurable and optimizable.
[0135] Specifically, in step S5, the parameters for clustering iterative optimization include:
[0136] The weights of each dimension of the multidimensional relative entropy deviation, the quantile of the cluster cutoff distance, and the entropy weight coefficient of the comprehensive risk entropy value;
[0137] The iteration rule is as follows: if the risk level prediction deviation rate of a single risk cluster exceeds a preset threshold for N consecutive periods, only the model parameters corresponding to that risk cluster will be corrected, while the parameters of the other risk clusters will remain unchanged.
[0138] The complete execution flow of cluster iterative optimization includes:
[0139] Triggering conditions are determined by calculating the prediction deviation rate for each risk cluster independently after each control cycle (30 days). The deviation rate is calculated as follows: (Number of risk objects whose predicted and actual levels are inconsistent) / Total number of risk objects in the cluster. If the deviation rate of a single risk cluster exceeds the 10% threshold for three consecutive cycles, the parameter optimization process for that cluster is triggered.
[0140] The range of optimized parameters includes the weights of each dimension of the multidimensional relative entropy deviation (search range [0.05, 0.40]), the quantile of the cluster cutoff distance (search range [1%, 3%]), and the entropy weight coefficient of the comprehensive risk entropy value.
[0141] The optimization objective function is to minimize the average prediction error rate over the next three periods. A Bayesian optimization algorithm (20 initial sampling points and 50 iterations) is used to search for the optimal value in the parameter space.
[0142] After optimization and rollback, all historical data of the cluster are backtested after the parameters are updated. If the backtesting deviation rate decreases by more than 15% compared with before optimization, the update is confirmed. Otherwise, the original parameters are maintained and abnormal events are recorded for manual review.
[0143] The update rules are isolated, and the optimization process is completed independently within a single cluster without involving the parameters of other risk clusters, ensuring that the stability of the already accurate model is not disturbed.
[0144] This application employs a continuous deviation trigger mechanism to initiate iterative optimization. Parameter optimization for a single risk cluster is only triggered when the prediction deviation exceeds the allowable range for multiple consecutive periods; occasional fluctuations do not trigger optimization, ensuring model stability. Optimizable parameters include deviation dimension weights, cluster quantiles, entropy weight coefficients, etc., with the goal of minimizing prediction deviation, and a gradient descent algorithm is used to iteratively correct the parameters. The entire optimization process is completed independently within a single cluster, without involving the parameters of other risk clusters, and will not affect other already accurate models.
[0145] This embodiment, by applying the method of the present invention, successfully identified the systemic risk of a general increase in the price of main materials in advance, and initiated a centralized procurement and price-locking strategy in advance, effectively avoiding batch cost overruns. Extreme outlier risks were also identified in advance, and alternative solutions were initiated, avoiding significant project delays and cost losses. High-risk risks received priority resource allocation, and the response speed was significantly accelerated; low-risk risks had simplified management processes, saving considerable management effort and greatly improving overall management efficiency. After several cycles of clustered iterative optimization, the assessment accuracy of various risks has improved to varying degrees, and the model's adaptability has been continuously enhanced, providing a reusable and accurate model for cost control in subsequent similar projects.
[0146] The above-described specific embodiments are preferred embodiments of a cost risk classification, early warning, and dynamic adjustment method based on deviation clustering analysis in this application. They are not intended to limit the specific scope of this application. The scope of this application includes, but is not limited to, these specific embodiments. All equivalent changes made in accordance with the shape and structure of this application are within the protection scope of this application.
Claims
1. A method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis, characterized in that, Includes the following steps: S1. Collect original cost indicators of multiple dimensions throughout the entire life cycle of the project, and calculate the relative entropy deviation of each cost risk object in the corresponding dimension with reference to the industry benchmark probability distribution, so as to obtain the multi-dimensional relative entropy deviation vector corresponding to each risk object. S2. Summarize the multidimensional relative entropy deviation vectors of all cost risk objects, and use the density peak clustering algorithm with adaptive truncation distance to perform global clustering to obtain multiple risk clusters with similar deviation characteristics, and identify and obtain outlier extreme risk points. S3. Calculate the comprehensive risk entropy value of each risk cluster, determine the basic early warning level of the corresponding risk cluster based on the comprehensive risk entropy value, and fine-tune the individual early warning level by combining the vector difference of single risk objects within the cluster to obtain the final graded early warning result of each cost risk object. S4. Based on the final graded early warning results of each risk cluster and outlier extreme risk point, match differentiated early warning push paths and risk handling strategies, generate graded early warning instructions and corresponding handling plans, and issue them for execution. S5. Collect the execution feedback and handling effect data of the corresponding handling plan for each risk cluster, build an independent state transition model for each risk cluster, start cluster iterative optimization with preset trigger conditions, correct the clustering model parameters, and realize dynamic adjustment of cost risk control.
2. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 1, characterized in that, Step S1 includes: S11. Retrieve the industry benchmark probability distribution corresponding to each dimension of cost indicators, and map the single-dimensional actual indicators of each risk object to the actual probability distribution value. S12. Calculate the distribution difference contribution of positive deviation and negative deviation respectively, and combine them to obtain the relative entropy deviation of this dimension; S13. Concatenate the relative entropy deviations of all dimensions of a single risk object in a preset order to form a multidimensional relative entropy deviation vector of the risk object.
3. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 1, characterized in that, Step S2 includes: S21. Calculate the Euclidean distance between each pair of the multidimensional relative entropy deviation vectors of all risk objects, and construct a distance set; S22. Sort the distance set in ascending order of numerical values, and take the distance value corresponding to the preset quantile after sorting as the cutoff distance.
4. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 3, characterized in that, Step S2 also includes: S23. Calculate the local density of each risk object, count the number of adjacent risk objects whose distance is less than the cutoff distance, and use this as the local density of the risk object. S24. Calculate the relative distance of each risk object. For the risk object with the highest local density, take the maximum distance between all objects as its relative distance; for the other risk objects, take the minimum distance between them and all objects with higher local density as their relative distance. S25. Determine the cluster center and complete the classification. Based on the product of local density and relative distance, select risk objects whose product is greater than a preset threshold as cluster centers. The remaining risk objects are assigned to the risk cluster corresponding to the nearest cluster center. Risk objects whose local density is less than a preset density threshold and whose relative distance is greater than a preset distance threshold are identified as outlier extreme risk points.
5. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 1, characterized in that, Step S3 includes: S31. Calculate the mean deviation, intra-cluster dispersion, and inter-cluster transmission coefficient for each individual risk cluster. The mean deviation is the average value of the vector magnitudes of all risk objects within the cluster, the intra-cluster dispersion is the standard deviation of the vector magnitudes within the cluster, and the inter-cluster transmission coefficient is the average vector similarity between the risk cluster and all other risk clusters. S32. The objective weights corresponding to the mean deviation, intra-cluster dispersion, and inter-cluster transmission coefficient are calculated using the entropy weight method. S33. Sum the three parameters with their corresponding weights to obtain the comprehensive risk entropy value of the risk cluster.
6. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 5, characterized in that, In step S3, determining the cluster-based early warning level based on the comprehensive risk entropy value includes: Two-level entropy thresholds are preset, with the lower threshold being less than the higher threshold; If the comprehensive risk entropy value is greater than or equal to the high threshold, the cluster-based early warning level is high risk; If the comprehensive risk entropy value is greater than or equal to the low threshold and less than the high threshold, the cluster basic warning level is medium risk. If the comprehensive risk entropy value is less than the low threshold, the cluster basic early warning level is low risk.
7. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 6, characterized in that, In step S3, fine-tuning the individual early warning level by combining the vector differences of single-risk objects within the cluster includes: Calculate the ratio of the vector magnitude of a single risk object to the mean of the average deviation of its risk cluster; If the ratio is greater than or equal to 1.2, the individual early warning level of the risk object will be increased by one level above the cluster base level; If the ratio is less than or equal to 0.8, the individual early warning level of the risk object will be adjusted down one level from the cluster base level. After the adjustment, the highest level is high risk, and after the adjustment, the lowest level is low risk.
8. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 1, characterized in that, Step S4 includes: High-risk clusters and outlier extreme risk points are pushed to the project decision-making level terminal and trigger the emergency response process, while simultaneously generating batch disposal plans. For medium-risk clusters, push them to the cost management department's terminal and match them with the corresponding optimized handling plan for the risk type; For low-risk clusters, only periodic summary reports are generated, and routine monitoring is performed.
9. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 1, characterized in that, Step S5 includes: S51. Define a set of risk states for a single risk cluster, including three types of states: low risk, medium risk, and high risk. S52. Based on the historical execution data of the corresponding disposal plan for the risk cluster, the frequency of transitions between different states is counted, and an initial state transition probability matrix is constructed, with the matrix elements corresponding to the probability of transitioning from one state to another. S53. Introduce the treatment feedback gain coefficient and modify the state transition probability matrix in combination with the actual treatment effectiveness of the treatment plan to obtain the state transition probability matrix after intervention.
10. The method for cost risk classification, early warning, and dynamic adjustment based on deviation clustering analysis according to claim 9, characterized in that, In step S5, the parameters for clustering iterative optimization include: The weights of each dimension of the multidimensional relative entropy deviation, the quantile of the cluster cutoff distance, and the entropy weight coefficient of the comprehensive risk entropy value; The iteration rule is as follows: if the risk level prediction deviation rate of a single risk cluster exceeds the preset threshold for three consecutive periods, only the model parameters corresponding to that risk cluster will be corrected, while the parameters of the other risk clusters will remain unchanged.