Risk prediction method, device, and program product for construction objects
By acquiring multi-dimensional feature data of construction objects and automatically extracting and calculating weight coefficients using a risk early warning model, the problems of lag and inconsistent standards in risk assessment in the construction industry have been solved, achieving efficient and accurate risk prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GLODON CO LTD
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-31
AI Technical Summary
Risk assessment in the construction industry relies on human experience and judgment, which leads to delays in risk identification and inconsistent judgment standards. Existing credit scoring models fail to fully consider the special characteristics of the construction industry, resulting in biased assessment results.
By acquiring multi-dimensional feature data of construction objects, using a risk warning model trained based on a preset classification method, extracting target warning indicators and calculating weight coefficients, determining risk probability values and prediction results, and achieving automated risk assessment.
It improves the accuracy and objectivity of risk prediction, avoids errors from subjective human judgment, ensures the comprehensiveness and comparability of assessment results, and supports automated risk prediction processes.
Smart Images

Figure CN122491929A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to risk prediction methods, devices, and program products for building construction objects. Background Technology
[0002] The construction industry is characterized by long project cycles, large capital requirements, complex industrial chains, and significant impact from policy and market fluctuations, making it a typical high-risk industry. Currently, risk assessment of construction projects largely relies on the human experience and judgment of managers. However, assessments based on human experience are limited by the individual knowledge and data processing capabilities of managers. When faced with multi-dimensional and large-scale industry-specific data, it is difficult to achieve comprehensive, objective analysis and continuous tracking, leading to problems such as delayed risk identification and inconsistent judgment standards. Summary of the Invention
[0003] This application provides a method, apparatus, and program for risk prediction of building construction objects to solve the problem of low accuracy in risk prediction.
[0004] Firstly, this application provides a risk prediction method for a construction object, comprising: acquiring multi-dimensional feature data of the construction object, the multi-dimensional feature data including various attribute data affecting the risk status of the construction object; extracting indicator data corresponding to the target warning indicators from the multi-dimensional feature data based on the target warning indicators in the risk warning model; acquiring the weight coefficients corresponding to each target warning indicator in the risk warning model; determining the risk probability value of the construction object based on each indicator data and its corresponding weight coefficients; and determining the risk prediction result of the construction object based on the target interval in which the risk probability value is located; wherein, the risk warning model is trained based on a preset classification method.
[0005] In some optional implementations, the risk prediction result of the construction object is determined based on the target range in which the risk probability value is located, including: if the risk probability value is less than a preset risk probability threshold, then a first warning level of the construction object is determined; if the risk probability value is not less than the preset risk probability threshold, then a second warning level of the construction object is determined; wherein, the risk prediction result includes the first warning level and the second warning level.
[0006] In some optional implementations, the first significance probability value corresponding to each target early warning indicator in the risk early warning model is obtained, as well as the deviation degree of the construction object under each target early warning indicator; from each target early warning indicator, the first early warning indicator with a weight coefficient greater than a preset weight threshold, a first significance probability value less than a first preset probability threshold, and an indicator deviation degree greater than a preset deviation threshold is selected.
[0007] In some optional implementations, the indicator type to which the first early warning indicator belongs, and the risk handling strategy matching the indicator type are obtained; based on the indicator deviation, the correlation parameters of the risk handling strategy are adjusted to obtain the target handling strategy for the construction object.
[0008] In some optional implementations, the training process of the risk warning model includes: acquiring a feature data sample set, which includes multiple feature data samples, each of which includes initial warning indicator samples of multiple dimensions and risk status labels corresponding to the feature data samples; filtering each initial warning indicator sample to obtain target warning indicator samples; performing stratified random sampling on the feature data sample set according to the risk status labels to obtain the training dataset corresponding to the feature data sample set; constructing a risk prediction model with the target warning indicator samples as independent variables and the risk status labels as dependent variables; and training the risk prediction model based on the training dataset.
[0009] In some optional implementations, the process of screening initial warning indicator samples to obtain target warning indicator samples includes: determining the correlation coefficient between each initial warning indicator sample and the risk status label, retaining first warning indicator samples whose correlation coefficient is not less than a preset correlation threshold; performing regression analysis on each first warning indicator sample to obtain the second significance probability value corresponding to each first warning indicator sample, retaining second warning indicator samples whose second significance probability value is less than a second preset probability threshold; performing collinearity verification on each second warning indicator sample to obtain the collinearity measurement value corresponding to each second warning indicator sample; and removing second warning indicator samples whose collinearity measurement value is not less than a preset measurement threshold to obtain target warning indicator samples.
[0010] In some optional implementations, multiple rounds of cross-validation are performed on the training dataset to obtain the multi-round validation results corresponding to the training dataset; based on the average value of the multi-round validation results, the evaluation index of the risk warning model is obtained; wherein, the evaluation index includes at least one of average accuracy, average recall and average precision.
[0011] In some optional implementations, a validation dataset corresponding to the feature data sample set is obtained. The validation dataset is obtained by stratified random sampling of the feature data sample set based on risk status labels. A target feature curve is plotted based on the validation dataset, and the area value under the target feature curve is determined. The validation dataset is input into the risk warning model to obtain the predicted risk status corresponding to each feature data sample in the validation dataset. The predicted risk status is compared with the risk status labels carried in the validation dataset, and the first recall rate corresponding to the feature data sample with the risk status label is determined based on the comparison result. If the area value is not less than a preset area threshold and the first recall rate is not less than a preset recall rate threshold, then the risk warning model is determined to have passed the validation.
[0012] In some optional implementations, if the performance of the risk warning model fails to pass verification, the indicator screening parameters are adjusted. The indicator screening parameters include at least one of a preset relevance threshold, a second preset probability threshold, and a preset measurement threshold. The adjusted indicator screening parameters are then used to re-screen each initial warning indicator sample.
[0013] Secondly, this application provides a risk prediction device for a construction object, comprising: a first acquisition module for acquiring multi-dimensional feature data of the construction object, the multi-dimensional feature data including various attribute data affecting the risk status of the construction object; an extraction module for extracting indicator data corresponding to the target warning indicator from the multi-dimensional feature data based on the target warning indicator in the risk warning model; a second acquisition module for acquiring the weight coefficients corresponding to each target warning indicator in the risk warning model; a first determination module for determining the risk probability value of the construction object based on each indicator data and its corresponding weight coefficients; and a second determination module for determining the risk prediction result of the construction object based on the target interval in which the risk probability value is located; wherein the risk warning model is trained based on a preset classification method.
[0014] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the risk prediction method for construction objects described in the first aspect or any corresponding embodiment.
[0015] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to execute the risk prediction method for a construction object described in the first aspect or any corresponding embodiment.
[0016] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute the risk prediction method for construction objects described in the first aspect or any corresponding embodiment.
[0017] The risk prediction method for construction objects provided in this application comprehensively integrates diverse information affecting risk status by acquiring multi-dimensional feature data of the construction objects, avoiding the one-sidedness of assessment caused by a single data source. It extracts data using target early warning indicators in the risk early warning model, ensuring the model focuses on key risk factors and improving the targeting and efficiency of risk identification. By introducing weight coefficients corresponding to each indicator, it achieves a quantitative distinction of the degree of influence of different risk factors. Based on indicator data and weight coefficients, it calculates risk probability values, integrating multi-dimensional information into a quantifiable risk metric, enhancing the objectivity and comparability of the assessment results. Furthermore, by combining the target interval to determine the risk prediction results, it achieves automated processing from raw data input to risk conclusion output. The entire process does not rely on subjective human judgment, effectively overcoming misjudgments of risk status caused by subjective human factors and improving the accuracy of risk prediction. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the first process of a risk prediction method for a building construction object according to an embodiment of this application; Figure 2 This is a schematic diagram of a second process for a risk prediction method for a building construction object according to an embodiment of this application; Figure 3 This is a schematic diagram of the receiver operating characteristic curve according to an embodiment of this application; Figure 4 This is a structural block diagram of a risk prediction device for a building construction object according to an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] The construction industry is characterized by long project cycles, large capital requirements, complex industrial chains, and significant impact from policy and market fluctuations, making it a typical high-risk industry. The risk of contract default by construction companies is particularly prominent. Currently, the assessment of default risk by construction companies largely relies on the subjective experience of management personnel or directly adopts existing credit scoring models from the financial lending sector.
[0022] For manual experience-based assessments, the comprehensiveness and timeliness of the results are limited by the individual knowledge and data processing capabilities of managers. When faced with large volumes of industry-specific data across multiple dimensions, such as corporate financial status, project payment collection, contract performance, and legal disputes, it is difficult to achieve comprehensive, objective analysis and continuous tracking. This leads to problems such as delayed identification of corporate default risks, strong subjectivity, and inconsistent judgment standards.
[0023] Directly adopting existing credit scoring models from the financial lending sector is problematic because these models were initially designed for standardized lending scenarios, and their internal parameter configurations and logical structures primarily reflect the risk transmission patterns within the financial lending field. When these models are directly applied to the construction industry, the unique operational characteristics and risk evolution features of the construction industry are not adequately considered. This can easily lead to discrepancies between the output assessment results and the actual risk status of the construction project.
[0024] The risk prediction method for construction projects provided in this application extracts target early warning indicators determined by the model directly from the multi-dimensional characteristic data of the construction project, including its financial situation, project payment collection, and contract performance. These indicators are then calculated according to their corresponding weight coefficients to obtain risk probability values, and the risk prediction results are output based on the target range. This process completely replaces manual judgment by management personnel, thus eliminating the problems of strong subjectivity, inconsistent standards, and delayed assessment caused by individual experience differences. Furthermore, since the early warning model used for calculation is trained based on a preset classification method, its internal indicators and weights are derived from learning from real-world difficulties faced by construction companies. This reflects the industry's own risk evolution patterns, rather than the general logic of the financial and credit field. Therefore, it can accurately capture the potential performance difficulties of construction projects during application, avoiding the assessment bias caused by industry mismatch in general credit scoring models.
[0025] According to an embodiment of this application, a risk prediction method for building construction objects is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0026] This embodiment provides a risk prediction method for building construction objects, which can be used on electronic devices such as desktop computers and laptops. Figure 1 This is a flowchart of a risk prediction method for building construction objects according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain multi-dimensional feature data of the building construction object. The multi-dimensional feature data includes various attribute data that affect the risk status of the building construction object.
[0027] The construction project entity refers to the construction company, which is the specific subject of the risk warning model assessment, i.e., the client or potential credit recipient of financial institutions. Multi-dimensional characteristic data is used to assess the credit risk or operational difficulties of the construction project entity through a multi-source heterogeneous data set. This data can include, for example, financial and operational data, project production progress data, industry reference data, and basic enterprise attribute data. Specifically, various information reflecting the construction project entity's operational status and performance capabilities, distributed across different data sources, can be automatically collected through data interfaces or database access. This can also be accomplished through integration with industry regulatory platforms, enterprise resource planning systems, or third-party credit reporting interfaces, using batch synchronous or real-time streaming methods. The data sources can further encompass the construction project entity's project performance records, administrative penalty information, supply chain transaction flows, etc. The collection cycle is not limited to annual collection; incremental updates can be performed quarterly, semi-annually, or event-triggered to adapt to different risk monitoring frequencies.
[0028] Specifically, based on pre-defined data collection rules, multi-source heterogeneous records related to construction production, financial health, market environment, and basic enterprise attributes are extracted from publicly disclosed information, industry statistical platforms, and business management records. These records include not only general financial and operational data but also specific business data reflecting the unique operating models of the construction industry (such as project cycles, capital recovery, and progress control). During the collection process, electronic devices automatically correlate and align data from different sources and in different formats to form a structured set of raw features. Based on this, the raw feature set undergoes adaptive processing and standardization to eliminate interference that may be introduced due to differences in collection criteria, data quality fluctuations, or inconsistencies in units, thereby transforming the raw features into standardized input information with uniform format, standardized values, and suitable for direct use by subsequent models.
[0029] Step S102: Based on the target warning indicators in the risk warning model, extract the indicator data corresponding to the target warning indicators from the multi-dimensional feature data.
[0030] The risk warning model is trained based on a preset classification method.
[0031] A predefined classification method refers to a training mechanism that organizes and utilizes historical sample data according to predefined category attributes during model training. This mechanism categorizes historical samples into different sample sets based on whether they possess certain specific features or outcome attributes, enabling the model to learn the distribution differences and decision boundaries of data features across different sets during training. Examples of predefined classification methods include Logistic Regression, Support Vector Machine (SVM), Random Forest, and Lightweight Gradient Boosting Machine (LightGBM).
[0032] A risk warning model refers to a data analysis model constructed to output an assessment of the risk status of a construction project within a specific future time period. Target warning indicators are a set of key indicators pre-selected from numerous initial indicators in the training samples during the model training phase. This set of indicators has a significant correlation with the risk status of the construction project and specifies which indicator data needs to be extracted when applying the model for prediction. Specifically, a completed risk warning model includes a pre-defined and fixed list of target warning indicators. These indicators are attributes that have been retained after a systematic screening and verification process and have been proven to have key discriminative capabilities for identifying risk status. In different implementations, the method for determining target warning indicators can be flexibly chosen: for example, business experts can pre-define a pool of candidate indicators based on industry experience and then refine them using data-driven feature selection methods; alternatively, they can be directly generated through an automatic feature selection mechanism during model training (such as screening based on regularization or feature importance). By comparing newly collected multi-dimensional feature data with this indicator list, those attribute values that perfectly match the list are extracted, thus obtaining the indicator data corresponding to each target warning indicator.
[0033] Step S103: Obtain the weight coefficients of each target early warning indicator in the risk early warning model.
[0034] Weight coefficients are contribution parameters automatically determined by the risk warning model during training to characterize the relative influence of each target warning indicator on the final prediction conclusion. This parameter is autonomously generated by the model's learning process based on historical sample data, and its specific value depends on the classification algorithm used and is not reliant on manual assignment. The form of the weight coefficients can vary depending on the algorithm framework used for the preset classification method: for example, in a model based on the linear assumption, it represents the weighted multiplier of each indicator; in a tree-based ensemble model, it represents the cumulative importance of features in the decision path; and in a distance-based model, it represents the direction and magnitude parameters of features in the construction of the classification boundary. Specifically, during the construction and training phase of the risk warning model, the electronic device automatically calculates the relative influence of each target warning indicator on the final prediction conclusion based on historical sample data through a preset data-driven learning mechanism. This solution process does not rely on manual experience-based assignment or fixed scoring rules. Instead, the model determines the contribution parameters of each indicator within the model's internal decision structure through iterative optimization or analytical calculation, based on the selected classification learning framework (such as linear mapping, nonlinear ensemble methods, etc.). Finally, the model establishes and stores the correlation between each target early warning indicator and its corresponding contribution parameters. When predicting new building construction projects, these determined parameters can be directly invoked to participate in the generation of subsequent evaluation information.
[0035] Step S104: Based on the data of each indicator and its corresponding weight coefficient, determine the risk probability value of the building construction object.
[0036] The risk probability value refers to the assessment value output by the risk warning model for the current construction project, used to characterize the likelihood of it encountering a specific operational difficulty in the future. This value is an intermediate or final result obtained by the model after calculating the input features. Its specific range and meaning can vary depending on the model type. For example, logistic regression outputs 0-1 values, and SVM outputs distance values after transformation, etc. In other model types, the risk probability value can also be expressed as a weighted average or voting probability of the prediction results of multiple sub-models, or a normalized score after calibration by equidistant mapping. Their common feature is that they provide a continuous value to reflect the risk tendency. Specifically, after obtaining the specific values of the current construction project on the target warning indicators and their corresponding contribution parameters, the risk warning model will perform weighted integration processing on the values of each indicator and contribution parameters according to its internally preset combination operation rules. This integration process is essentially the fusion and comprehensive evaluation of multi-dimensional feature information. Its operation method is determined by the type of classification algorithm selected by the model (e.g., linear weighted sum followed by nonlinear transformation, or aggregation voting through multi-level decision paths, etc.). After this calculation, the model will output an assessment value that comprehensively reflects the risk tendency of the current object. The range and specific meaning of this value depend on the model type, but its core function is to provide a continuous reference for subsequent category determination.
[0037] Step S105: Based on the target range where the risk probability value is located, determine the risk prediction result of the building construction object.
[0038] The target interval refers to a numerical segment within the possible range of risk probability values, defined by one or more reference judgment benchmarks. Each numerical segment corresponds to a specific warning category. The model determines the final warning category by determining which numerical segment the risk probability value falls into. There are several ways to set reference judgment benchmarks: besides using the statistical characteristics of validation samples to determine the judgment threshold, fixed thresholds or multi-level interval boundaries can be set directly based on business risk preferences, such as dividing 0 to 0.3 into a low-risk interval, 0.3 to 0.6 into a concern interval, and 0.6 to 1.0 into a high-risk interval. Alternatively, the benchmarks can be automatically or manually adjusted in stages according to preset rules based on changes in the overall industry risk level or the institution's own risk tolerance. Specifically, the risk warning model pre-establishes one or more reference judgment benchmarks for dividing the continuous assessment value range. These benchmarks divide the possible value space of the assessment value into several distinct numerical segments. After the model calculates the assessment value of the current construction object, it compares this value with the preset reference benchmarks to determine which specific numerical segment the value falls into. Each numerical segment is mapped to a predefined risk situation category. Once the segment in which the value is located is determined, the model automatically outputs the risk category indication information (such as different levels of warning indicators) corresponding to that segment as the final risk prediction result, thereby completing the transformation from continuous assessment information to discrete decision information.
[0039] The risk prediction method for construction objects provided in this application comprehensively integrates diverse information affecting risk status by acquiring multi-dimensional feature data of the construction objects, avoiding the one-sidedness of assessment caused by a single data source. It extracts data using target early warning indicators in the risk early warning model, ensuring the model focuses on key risk factors and improving the targeting and efficiency of risk identification. By introducing weight coefficients corresponding to each indicator, it achieves a quantitative distinction of the degree of influence of different risk factors. Based on indicator data and weight coefficients, it calculates risk probability values, integrating multi-dimensional information into a quantifiable risk metric, enhancing the objectivity and comparability of the assessment results. Furthermore, by combining the target interval to determine the risk prediction results, it achieves automated processing from raw data input to risk conclusion output. The entire process does not rely on subjective human judgment, effectively overcoming misjudgments of risk status caused by subjective human factors and improving the accuracy of risk prediction.
[0040] This embodiment provides a risk prediction method for building construction objects, which can be used on electronic devices such as desktop computers and laptops. Figure 2 This is a flowchart of a risk prediction method for building construction objects according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain multi-dimensional feature data of the building construction object. This multi-dimensional feature data includes various attribute data that affect the risk status of the building construction object. For details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.
[0041] Step S202: Based on the target early warning indicators in the risk early warning model, extract indicator data corresponding to the target early warning indicators from the multi-dimensional feature data. The risk early warning model is trained based on a preset classification method. For details, please refer to [link to relevant documentation]. Figure 1 Step S102 of the illustrated embodiment will not be described again here.
[0042] Step S203: Obtain the weight coefficients of each target early warning indicator in the risk early warning model. For details, please refer to [link to relevant documentation]. Figure 1 Step S103 of the illustrated embodiment will not be described again here.
[0043] Step S204: Based on the data of each indicator and its corresponding weighting coefficient, determine the risk probability value of the building construction object. For details, please refer to [link to relevant documentation]. Figure 1 Step S104 of the illustrated embodiment will not be described again here.
[0044] Step S205: Based on the target range where the risk probability value is located, determine the risk prediction result of the building construction object.
[0045] Specifically, step S205 includes: Step S2051: If the risk probability value is less than the preset risk probability threshold, then determine the first warning level of the building construction object.
[0046] Step S2052: If the risk probability value is not less than the preset risk probability threshold, then determine the second early warning level of the building construction object.
[0047] The risk prediction results include a first warning level and a second warning level.
[0048] A preset risk probability threshold refers to a reference criterion used to map the continuous assessment results represented by the risk probability value to a discrete early warning category. This threshold can be flexibly adjusted and set by experts in the relevant field, taking into account the actual operating scenarios, risk tolerance, industry policy requirements, and project characteristics of the construction project. For example, the initial early warning threshold set by the model by default can be 0.5, and the optimal threshold can be determined based on the receiver operating characteristic (ROC) curve analysis of historical samples, with the probability point corresponding to the maximum Youden index (e.g., 0.645) as the optimal threshold, to achieve the best balance between sensitivity and specificity.
[0049] The first warning level refers to the category indication information output by the model to identify the relatively low-risk state of the construction object when the risk probability value is less than a preset risk probability threshold. The second warning level refers to the category indication information output by the model to identify the relatively high-risk state of the construction object when the risk probability value is greater than or equal to the preset risk probability threshold. Specifically, the risk warning model internally stores a reference judgment benchmark value to distinguish different warning categories. During prediction, the obtained assessment value is compared with this benchmark value: if the assessment value is lower than the benchmark value, it indicates that the probability of the construction object experiencing a concern event in the future is relatively low, and the model automatically marks the object as the first warning category; if the assessment value reaches or exceeds the benchmark value, it indicates that its risk tendency is relatively high, and the model marks it as the second warning category. This comparison and judgment process is completely executed automatically by the computer without human intervention. The reference benchmark value is not fixed and immutable; its specific value can be flexibly set and adjusted during the model deployment or operation phase according to the different risk tolerance requirements of the actual application scenario to adapt to different warning strategies.
[0050] The risk prediction method for construction objects provided in this application transforms continuous risk probability values into discrete first and second warning levels by setting a risk probability threshold. This presents complex quantitative risk assessment results in an intuitive and clear binary form, greatly reducing the understanding and usage threshold for business personnel in the decision-making process. At the same time, this threshold determination mechanism ensures the consistency and comparability of warning level classification standards among different assessment objects and assessment batches, avoiding confusion in level identification caused by differences in subjective human judgment, and providing a clear and stable basis for the rapid initiation of risk response measures and resource allocation.
[0051] In some optional implementations, the above-mentioned risk prediction method for building construction objects further includes: Step a1: Obtain the first significance probability value of each target early warning indicator in the risk early warning model, and the deviation of the construction object from each target early warning indicator.
[0052] The first significance probability value refers to the reference metric used by the risk warning model to assess the statistical reliability of each target warning indicator to the final prediction conclusion. This metric reflects the likelihood that the observed association between the indicator and the prediction target is caused by random factors; the smaller the value, the more reliable the predictive contribution of the indicator is generally.
[0053] The deviation of an indicator refers to the degree of deviation of the actual value of a construction project on a certain target early warning indicator from the benchmark level (such as the mean, median, or reasonable range boundary) of that indicator in a reference group (such as projects in the same industry, of the same size, or in the same region). This metric is used to determine whether the value of the indicator is within an abnormal range.
[0054] Specifically, the first significance probability value is obtained based on the statistical evaluation records of the predictive reliability of each indicator during the training phase of the risk warning model. During model building, a statistical reliability metric is automatically calculated to correlate each indicator with the prediction target, and these metrics are stored as auxiliary attributes of the indicator. When the trained model needs interpretation or secondary screening, these pre-calculated statistical reliability metrics can be directly retrieved from within the model. The indicator deviation is obtained by comparing the actual values of a new object on each indicator with pre-calculated benchmark levels (such as group mean or reasonable range boundaries) obtained from historical samples or industry data when predicting a new object, thus deriving a metric representing the degree of abnormality in its value.
[0055] Step a2: From the various target warning indicators, select the first warning indicator whose weight coefficient is greater than the preset weight threshold, whose first significance probability value is less than the first preset probability threshold, and whose indicator deviation is greater than the preset deviation threshold.
[0056] The preset weight threshold refers to the reference limit set for the weight coefficients of important indicators. When the weight coefficient of an indicator exceeds this reference limit, it can be identified as a key indicator that has a strong influence on the prediction conclusion.
[0057] The first preset probability threshold refers to the reference limit set for the first significance probability value when screening indicators that meet the statistical reliability standard. When the significance probability value of an indicator is lower than this reference limit, it can be considered that it has a reliable ability to distinguish the prediction target in a statistical sense.
[0058] The preset deviation threshold refers to the reference limit set for the degree of deviation of an indicator, used to identify abnormal deviation indicators. When the deviation of an indicator exceeds the reference limit, it can be determined that the construction object is significantly different from the reference group in that indicator dimension.
[0059] The first early warning indicator refers to the key feature that is further identified from the target early warning indicator set based on preset multi-dimensional screening conditions, and has outstanding explanatory power for the risk causes of the current specific construction object. It can also be called a high-risk indicator.
[0060] Specifically, after obtaining the predictive assessment information for the current construction project, a multi-dimensional indicator filtering operation will be performed to further identify the key influencing factors leading to the assessment conclusion. This operation is based on a joint judgment using three parallel screening conditions: checking whether the contribution parameter of each indicator in the model exceeds a preset influence reference limit; checking whether the statistical reliability measure exhibited by each indicator during the model training phase is lower than a preset reliability reference limit; and checking whether the deviation of the actual value of each indicator on the current object from the reference benchmark exceeds a preset anomaly reference limit. Only indicators that simultaneously meet the above three conditions will be judged as key features with significant explanatory power for the current risk status of the object, and will be extracted and retained as the first early warning indicator.
[0061] In the above implementation, by combining the weighting coefficient of the target early warning indicator, the first significance probability value, and the indicator deviation degree for joint screening, it is possible to accurately identify key risk-driving indicators that have a significant impact in the statistical sense of the model and deviate significantly from the normal range in actual business performance as the first early warning indicators. This effectively eliminates the randomness or noise interference that may be introduced by a single-dimensional judgment, and significantly improves the accuracy and credibility of high-risk factor identification. At the same time, the selected first early warning indicators are directly incorporated into the risk prediction results, so that the early warning conclusions have clear attribution orientation and interpretability, providing clear and quantitative action guidelines for subsequent targeted risk investigation and precise prevention and control.
[0062] In some optional implementations, the above-mentioned risk prediction method for building construction objects further includes: Step b1: Obtain the indicator type to which the first early warning indicator belongs, and the risk handling strategy that matches the indicator type.
[0063] The indicator type refers to the attribute classification of the enterprise's business aspects reflected by the first warning indicator, such as the category reflecting profitability, the category reflecting turnover efficiency, the category reflecting project progress, etc. This classification is used to match the corresponding framework of response measure suggestions. The risk handling strategy refers to the framework of response measures or the suggestion template with directional guiding significance preset for the abnormal performance of a specific indicator type. Specifically, an indicator attribute classification mapping table is pre-maintained inside the electronic device, which records the corresponding relationship between each target warning indicator that may be selected and the enterprise's business aspects it reflects (for example, one indicator is classified into the capital efficiency dimension, and another indicator is classified into the project execution dimension, etc.). After the first warning indicator is identified, the type attribution of each indicator can be determined by querying this mapping table. At the same time, the electronic device also stores a set of basic response measure framework libraries corresponding to each indicator type. In this framework library, disposal suggestion templates with directional guiding significance are preset for the abnormal performance of each business aspect. Through the bridge of the indicator type, the basic risk handling strategy framework that matches it can be automatically retrieved and called.
[0064] Step b2, based on the indicator deviation degree, adjust the associated parameters of the risk handling strategy to obtain the target handling strategy for the construction object.
[0065] The associated parameter refers to the specific control element or the level of suggestion intensity that can be dynamically adjusted according to the severity of the indicator abnormality in the risk handling strategy, such as the suggested amount of capital injection, the number of days for account period adjustment, the monitoring frequency level, etc. The target handling strategy refers to the personalized response measure plan for the current specific construction object finally generated after dynamically adapting and adjusting the risk handling strategy framework and the associated parameters. Specifically, after obtaining the basic response measure framework that matches the indicator type, the adjustable control elements in the framework will be dynamically adapted according to the actual severity of the deviation of this indicator from the reference benchmark. The greater the deviation degree, usually the greater the response intensity required or the higher the disposal priority involved. Through the preset mapping rules, the specific value of the deviation degree is converted into an adjustment instruction for the associated parameters (such as the suggested execution frequency, the suggested attention level, or the suggested resource allocation intensity, etc.), and the corresponding part in the basic framework is modified and refined. After this personalized adjustment based on the deviation degree, the basic framework is transformed into a more operable target handling strategy plan specifically for the specific risk situation of the current construction object for the decision maker to refer to and execute.
[0066] In the above implementation, by obtaining the indicator type of the first early warning indicator and matching the corresponding risk handling strategy, the logical connection and knowledge reuse between risk factors and handling methods are realized. On this basis, the correlation parameters of the strategy are dynamically adjusted according to the indicator deviation to generate a target handling strategy for a specific construction object. This makes the output suggestions not only targeted in terms of type, but also adapted to the actual risk deviation in terms of handling intensity and urgency. This effectively avoids the problem of one-size-fits-all strategies and significantly improves the precision and operational feasibility of risk response.
[0067] In some optional implementations, the training process of the risk warning model includes: Step c1: Obtain the feature data sample set, which includes multiple feature data samples. Each feature data sample includes initial early warning indicator samples in multiple dimensions and risk status labels corresponding to the feature data sample.
[0068] The feature data sample set refers to the historical data set used to build and train the risk warning model. This set consists of multiple feature data samples, each containing a set of multi-dimensional feature descriptions related to the construction object, as well as reference information identifying the risk category attribute to which the sample belongs. To ensure the stability and statistical significance of model training, the total number of samples in the collected feature data sample set should generally be no less than 300, of which at least 30 samples should be labeled with a risk status tag.
[0069] A feature data sample refers to a single record unit in a feature data sample set. Each record unit corresponds to a feature description of a specific building construction object at a certain historical observation point in time, including the specific values of the object on multiple preset feature dimensions, as well as the actual risk performance category of the object as subsequently observed.
[0070] Risk status labels are identifiers attached to each feature data sample to distinguish different risk performance categories, such as 0 = normal enterprise and 1 = distressed enterprise. This identifier reflects whether the construction project corresponding to that sample has experienced a pre-defined event of concern during a specific observation period. Events of concern may include situations that seriously affect the ability to fulfill obligations, such as defaults in the open market, persistent overdue commercial acceptance bills, lawsuits by financial institutions, bankruptcy reorganization, or bankruptcy liquidation.
[0071] Specifically, in the model building preparation phase, historical operating information and subsequent performance records of multiple construction projects within a certain time period are collected by accessing various data storage sources or data interfaces, from publicly disclosed corporate documents, industry statistical databases, and business management records. For each construction project at each observation point, a set of feature descriptions covering multiple aspects such as financial health, production operations, market environment, and basic attributes are extracted to form a structured record unit. Simultaneously, based on pre-defined rules for defining events of interest, it is determined whether the project has experienced a corresponding risk event within a specific period after the observation point, and an identifier is attached to this record unit to distinguish different risk performance categories. The historical data set composed of a large number of such record units constitutes the feature data sample set used for subsequent model learning and validation.
[0072] Step c2: Filter each initial warning indicator sample to obtain the target warning indicator sample.
[0073] The initial early warning indicator sample refers to the set of original candidate feature variables directly derived from feature data samples during the model training phase, without prior screening. For example, it can cover multiple dimensions such as financial, production, industry, and basic attributes, forming the initial input space for feature selection by the model. Specifically, this initial early warning indicator sample may include: financial dimensions such as debt-to-asset ratio, current ratio, cash-to-short-term-borrowing ratio, interest coverage ratio, return on assets, accounts receivable turnover days, cash collection cycle, and project collection rate; production dimensions such as completion value recognition time, recognition value receipt time, project schedule deviation rate, construction delay days, and building material price fluctuation coefficient; industry dimensions such as industry average debt-to-asset ratio, operating profit margin, accounts receivable turnover days, and industry project overdue rate; and basic dimensions such as company size, years of establishment, and qualification level.
[0074] After obtaining the feature data sample set, an automated feature optimization process will be initiated to identify a subset of features from the initial feature items (i.e., the initial early warning indicator samples) that contribute substantially to distinguishing different risk performance categories and have low information redundancy. After a series of screening steps automatically executed by electronic devices without human intervention, the final subset of features retained is the target early warning indicator sample, which constitutes the core variable space for subsequent model construction.
[0075] In some alternative implementations, step c2 above includes: Step c21: Determine the correlation coefficient between each initial warning indicator sample and the risk status label, and retain the first warning indicator sample whose correlation coefficient is not less than the preset correlation threshold.
[0076] The correlation coefficient is a statistical measure used to quantify the strength and direction of the association between a single initial warning indicator and a risk status label. The magnitude of this measure reflects the degree of correlation between changes in the indicator's value and changes in the risk category; a larger absolute value indicates a stronger association.
[0077] The preset correlation threshold refers to the reference limit set for the correlation coefficient value for the preliminary screening of indicators that are related to the risk category. For example, it can be 0.3.
[0078] The first early warning indicator sample refers to a subset of indicators that have a certain degree of correlation with the risk category and are retained after correlation screening.
[0079] Specifically, each initial feature item in the feature data sample set is traversed. For each initial feature item, a pre-defined association strength measurement algorithm is used to calculate the accompanying change relationship between its value sequence and the corresponding risk category identifier sequence, resulting in a quantified association metric. The sign of this metric indicates the direction of positive or negative association, while its absolute value reflects the strength of the association. An internal reference threshold is pre-set to distinguish between strong and weak associations. The absolute value of the association metric for each feature item is compared with this reference threshold: if it reaches or exceeds the threshold, it indicates a certain degree of association between the feature item and the risk category, and the feature item is retained and included in the first early warning indicator sample set; if it does not reach the threshold, the feature item is considered to contribute weakly to risk category differentiation and is removed from subsequent screening.
[0080] Step c22: Perform regression analysis on each first warning indicator sample to obtain the second significance probability value corresponding to each first warning indicator sample, and retain the second warning indicator samples whose second significance probability value is less than the second preset probability threshold.
[0081] The second significance probability value is a reference metric used in the indicator selection phase of a risk warning model to assess the statistical reliability of each candidate indicator in explaining the risk category. Its meaning is similar to the first significance probability value mentioned above, but it is used in the indicator selection phase during training.
[0082] The second preset probability threshold refers to the reference limit set for the second significance probability value to retain statistically reliable indicators during the indicator screening stage. For example, it can be 0.05. When the significance probability value of an indicator is lower than this reference limit, it is considered to have a non-accidental explanatory power for the risk category and is retained.
[0083] The second early warning indicator sample refers to the subset of indicators that have been retained after correlation screening and significance screening, but have not yet undergone feature redundancy testing.
[0084] Specifically, after the initial screening of correlation strength, a further statistical reliability assessment is performed on the retained first-warning indicator samples. This assessment typically relies on a feature evaluation model that can simultaneously consider the combined impact of multiple features. This model performs hypothesis testing on the independent contribution of each feature in explaining differences in risk categories. For each first-warning indicator sample, the model outputs a reliability metric reflecting its statistically significant non-zero contribution. The smaller this value, the less likely the predictive contribution of that feature is due to random data fluctuations. The electronic device has a pre-set reference threshold for determining statistical reliability. The reliability metric of each feature is compared to this reference threshold: if it is below the threshold, the feature is considered to have a non-accidental explanatory power for the risk category, is retained, and included in the second-warning indicator sample set; otherwise, it is removed from the candidate set.
[0085] Step c23: Perform collinearity verification on each second early warning indicator sample to obtain the collinearity measurement value corresponding to each second early warning indicator sample.
[0086] Multicollinearity is a statistical measure used to assess the degree of information overlap or substitution among multiple indicators. When an indicator exhibits high multicollinearity with other indicators, it indicates limited incremental information and may lead to decreased model stability. Specifically, after retaining statistically reliable features, it is necessary to evaluate the degree of correlation between features within the second warning indicator sample set to avoid decreased model stability due to high information overlap between features. Using a pre-defined multicollinearity diagnostic technique, a measure is calculated for each second warning indicator sample to assess its linear representation by other features. The higher the measure, the less independent information the feature contains and the more severe its overlap with other features.
[0087] Step c24: Remove the second early warning indicator samples whose collinearity measurement value is not less than the preset measurement threshold to obtain the target early warning indicator samples.
[0088] The preset measurement threshold refers to a reference limit set for collinearity measurement values to eliminate highly redundant indicators; for example, it can be 10. When the collinearity measurement value of an indicator reaches or exceeds this reference limit, the indicator is considered redundant and removed from the candidate set. Specifically, the electronic device internally sets a reference limit value for determining the degree of redundancy of feature items. After obtaining the collinearity measurement values of each second warning indicator sample, the measurement value of each feature item is compared with the reference limit value in turn: if the measurement value of a feature item reaches or exceeds the limit, it indicates that the feature item has serious information overlap with other selected feature items, and its added predictive value is limited; the electronic device will remove it from the candidate set. If the measurement value is below the limit, it indicates that the feature item has relatively independent predictive information and is retained. After this round of redundancy elimination, the final set of retained feature items is the target warning indicator sample.
[0089] In some optional implementations, the processes described in steps c21, c22, c23 to c24 above essentially constitute a three-step progressive indicator screening process. Specifically, the correlation analysis performed in step c21 preferably uses the Pearson correlation coefficient for calculation; the regression analysis and significance test performed in step c22 can be implemented using stepwise regression to automatically screen for indicator combinations that have significant explanatory power for risk status, with the second significance probability value being the commonly used p-value in statistics; the collinearity check and redundancy removal performed in steps c23 and c24 can be performed using the variance inflation factor (VIF) diagnostic technique. This three-step screening process is performed sequentially, progressively purifying the indicator set to ensure that the final selected target early warning indicator samples possess both predictive relevance and statistical reliability.
[0090] If, after the above screening, the number of target early warning indicator samples remaining is too large (e.g., greater than or equal to 15), a second simplification rule is executed: this includes retaining the number of indicators with the smallest P-value (e.g., 10 to 12); under the same conditions, prioritizing the retention of industry-specific indicators before financial indicators; and performing collinearity diagnosis again to ensure that the variance inflation factor (VIF) is less than a more stringent value (e.g., 5). Conversely, if the number of indicators after screening is too small (e.g., less than or equal to 3), a supplementary relaxation rule is executed: this includes relaxing the Pearson correlation coefficient threshold to 0.2, relaxing the stepwise regression P-value threshold to 0.1, and forcibly including some industry-strongly correlated indicators (e.g., collection rate, schedule deviation, construction delay, etc.) before screening again to ensure that the model has basic discriminative ability.
[0091] In the above implementation, by first screening the first warning indicator samples whose correlation with the risk status label reaches a preset threshold based on the correlation coefficient value, irrelevant features with weak contribution to risk prediction are effectively eliminated. Then, regression analysis is performed on the retained indicators, retaining only the second warning indicator samples whose second significance probability value meets the preset probability threshold requirement, ensuring that the selected indicators have statistically significant explanatory power for the risk results. Finally, indicators with excessively high collinearity values are eliminated through collinearity verification to obtain the target warning indicator samples, avoiding the problems of unstable model coefficient estimation and interpretation distortion caused by high linear correlation between indicators. This series of progressive screening mechanisms ensures that the final set of indicators input to the model possesses correlation, significance, and independence, providing a reliable feature foundation for the high-precision prediction and robust generalization of the risk warning model.
[0092] Step c3: Perform stratified random sampling of the feature data sample set according to the risk status label to obtain the training dataset corresponding to the feature data sample set.
[0093] The training dataset refers to a subset of samples extracted from the feature data sample set according to a preset data partitioning rule, specifically used for model parameter learning and optimization. Specifically, before model training, to ensure that the sample distribution of different risk performance categories in the training data remains consistent with the original population distribution, the electronic device employs a sampling strategy that maintains class proportions. Specifically, the feature data sample set is grouped according to the risk category label carried by each sample. Then, within each group, a corresponding number of samples are independently and randomly selected according to a preset overall sampling ratio (e.g., 7:3, 6:4, 8:2, etc.). Finally, the samples extracted from each group are merged together to form the training dataset for model parameter learning. This sampling method ensures that the training set is highly consistent with the original dataset in terms of class composition, avoiding model learning distortion caused by sampling bias. Preferably, the feature data sample set is stratified randomly sampled at a ratio of 7:3, that is, 70% of the samples are used as the training set and 30% as the validation set, with the stratification based on risk status labels.
[0094] Step c4: Using the target early warning indicator sample as the independent variable and the risk status label as the dependent variable, construct a risk prediction model.
[0095] After determining the final set of feature variables used for modeling (i.e., the target early warning indicator samples), the model building process begins. The core of this process is establishing a mathematical framework that describes the mapping relationship between input features and output categories. The target early warning indicator samples are set as the input variables in this framework, and the risk category identifier is set as the desired output target variable. Based on the selected classification learning mechanism, a decision structure with adjustable parameters is initialized. The form and complexity of this structure depend on the type of algorithm used (e.g., linear combination structure, tree-based decision structure, etc.). At this point, although the model structure is established, its internal parameters are not yet determined, and it is in a state of waiting for training.
[0096] For example, when using the Logistic Regression algorithm, the constructed model equation takes the form of: ,in To predict the probability that an object will encounter difficulties in the next year. For constant terms, to The regression coefficients for each indicator are denoted as . to These are the standardized values for the corresponding indicators.
[0097] Step c5: Train the risk prediction model based on the training dataset.
[0098] After the mathematical framework of the risk prediction model is constructed, the previously prepared training dataset is input into this framework. The parameter optimization algorithm inside the model reads the input feature values and corresponding risk category labels of each training sample one by one. By comparing the predicted tendency calculated by the model based on the current parameters with the actual category label, and aiming to minimize this difference, the model repeatedly and iteratively updates and adjusts each adjustable parameter using a preset numerical optimization method. This process continues until the model's prediction performance on the training dataset reaches the preset convergence condition or completes the specified number of iterations. After training is complete, the parameters inside the model are fixed, and the risk prediction model now has the ability to output corresponding risk prediction assessment information based on new input feature values.
[0099] For example, when using Logistic Regression, the training process involves solving for the constant term in the equation using the maximum likelihood estimation method. and each regression coefficient to .
[0100] In the above implementation, by acquiring a feature data sample set containing multi-dimensional initial early warning indicator samples and corresponding risk status labels, and then filtering it to obtain target early warning indicator samples, redundant or irrelevant features are effectively eliminated, reducing model complexity and the risk of overfitting. Simultaneously, by performing stratified random sampling of the sample set according to the risk status labels to form a training dataset, the distribution of samples from different risk categories in the training data is ensured to remain consistent with the overall distribution, avoiding model training distortion caused by sample distribution bias. This guarantees the model's generalization ability and prediction stability in real-world application scenarios. Furthermore, by constructing and training a model using target early warning indicator samples as independent variables and risk status labels as dependent variables, the model can objectively learn the intrinsic relationship between each indicator and risk outcome, providing a solid technical foundation for subsequent automated and accurate risk prediction.
[0101] In some optional implementations, the training process of the risk warning model further includes: Step d1 involves performing multiple rounds of cross-validation on the training dataset to obtain the validation results for the training dataset.
[0102] Multi-round validation results refer to the multiple sets of performance metric records obtained during model training by repeatedly dividing the model into training and validation subsets to evaluate its stable performance on unseen data. Each set of records reflects the model's predictive performance under that round of partitioning. Specifically, during the model training phase, to comprehensively evaluate the model's performance stability on data not involved in parameter learning, a cyclical partitioning and repeated evaluation operation is performed on the prepared training dataset. Specifically, the training dataset is roughly divided into several subsets, and then multiple rounds of model training and validation are performed: in each round, one subset is selected alternately as temporary validation data, while the remaining subsets are combined as training data for that round. Within each round, the model is trained using the training data of that round, and its performance metrics are immediately tested on the validation data of that round. Because this process is repeated for a number of rounds equal to the number of subsets, each round generates a corresponding set of performance metric records. These records from different rounds together constitute the multi-round validation results of the training dataset, reflecting the performance fluctuations of the model under different data partitioning methods.
[0103] Preferably, the multi-round cross-validation is 5-fold cross-validation, and the stratification basis is consistent with that used when the model was built, namely, the risk status label.
[0104] Step d2: Based on the average value of the results of multiple rounds of validation, the evaluation index of the risk warning model is obtained.
[0105] The evaluation metrics include at least one of average accuracy, average recall, and average precision.
[0106] Evaluation metrics refer to a series of calculated indicators used to quantitatively measure the performance of a risk warning model. These metrics reflect the model's actual utility from different perspectives (such as prediction accuracy, ability to identify categories of interest, and prediction coverage). Specifically, after obtaining multiple sets of performance measurement records generated during the aforementioned multi-round validation process, these records are summarized and processed to obtain an evaluation conclusion that comprehensively reflects the overall performance of the model. For each preset performance evaluation dimension (such as a measure of overall prediction accuracy, a measure of coverage of categories of interest, and a measure of the accuracy of identification results), the corresponding measurement values from each round of validation results are extracted. Then, by calculating the central tendency statistic (such as the arithmetic mean) of these values, the dispersed measurement information from multiple rounds is aggregated into a single, representative comprehensive evaluation value. This comprehensive evaluation value is the evaluation metric used to measure the performance of the risk warning model in that performance dimension. By taking the average value, the evaluation fluctuations caused by the randomness of a single data partitioning can be effectively smoothed, making the evaluation results more robust and valuable.
[0107] In the above implementation, by performing multiple rounds of cross-validation on the training dataset and obtaining the model's evaluation metrics based on the average of the multiple rounds of validation results, the evaluation bias and random fluctuations that may be caused by a single random partition are effectively avoided. This allows the calculated average accuracy, average recall, and average precision to more objectively and robustly reflect the true generalization performance of the risk warning model on different data subsets, thereby providing a reliable quantitative basis for the reliable evaluation of model performance and subsequent optimization iterations.
[0108] In some optional implementations, the training process of the risk warning model further includes: Step e1: Obtain the validation dataset corresponding to the feature data sample set. The validation dataset is obtained by stratified random sampling of the feature data sample set based on the risk status label.
[0109] A validation dataset is a subset of samples drawn from the feature data sample set according to rules consistent with those used in the training dataset, and which does not participate in the model parameter learning process. This subset is specifically used to independently test the performance of the trained model to evaluate its generalization ability. Specifically, before model training, in addition to preparing the training data for parameter learning, a portion of data completely independent of the training process needs to be reserved for final performance testing of the trained model. To obtain this independent data, a sampling method that maintains the proportion of class composition is used from the original feature data sample set. First, the original dataset is divided into different groups based on the risk category label carried by each sample. Then, within each group, a subset of samples is randomly selected according to a pre-set retention ratio consistent with the training data partitioning. The samples selected from all groups are combined to form the validation dataset. Because this dataset ensures that the proportion of samples from each class is consistent with the original population during extraction, it can objectively and unbiasedly test the model's ability to distinguish between different categories of objects.
[0110] Step e2: Plot the target feature curve based on the validation dataset and determine the area value under the target feature curve.
[0111] A target feature curve is a graphical analysis tool used to visually illustrate the trade-off between a risk warning model's ability to identify the category of concern and the risk of misjudgment at different judgment thresholds. The shape of the curve reflects the model's ability to distinguish between the two types of samples. For example, this target feature curve can be a receiver operating characteristic (ROC) curve, such as... Figure 3 As shown.
[0112] The area value refers to the quantified area of the region enclosed below the target feature curve. This value comprehensively reflects the model's overall discriminative ability under different decision thresholds; the closer the value is to the maximum value, the better the model's discriminative performance. For example, when the target feature curve is an ROC curve, this area value is the AUC (Area Under Curve) value.
[0113] Specifically, when conducting in-depth analysis of model performance using validation datasets, a graphical evaluation tool is typically used to visually represent the model's overall performance under various judgment criteria. Taking the true risk category identifiers of each sample in the validation dataset and the propensity score calculated by the model as input, the reference judgment criteria used to distinguish different prediction categories are continuously adjusted. The dynamic relationship between the model's success rate in identifying the category of interest and the misclassification rate for the category of non-interest is recorded under each criterion setting, and these relationships are plotted as a feature curve on a two-dimensional plane. The area under this curve is a quantitative measure that varies within a specific range, and its magnitude comprehensively reflects the model's overall discriminative ability when not constrained by a specific judgment criterion. The closer the area value is to its theoretical upper limit, the stronger the model's ability to distinguish different categories of samples.
[0114] Step e3: Input the validation dataset into the risk warning model to obtain the predicted risk status corresponding to each feature data sample in the validation dataset.
[0115] Predicting risk status refers to the risk category classification determined by a pre-trained risk warning model after samples from the validation dataset are input into the model, based on the model's internal decision-making logic. Specifically, after the model is trained and its internal parameters are fixed, each sample from the validation dataset is sequentially input into the learned risk warning model. For each input sample, the model performs a front-to-back evaluation calculation based on the specific values of its features and its fixed internal decision-making logic and calculation rules. This calculation generates a risk propensity assessment for the sample. The model then converts this assessment into a clear predicted category conclusion similar to the sample's original risk label, based on its built-in classification rules. This predicted category conclusion is the predicted risk status of the sample. By performing this process on all samples in the validation set, a complete sequence of model prediction results can be obtained.
[0116] Step e4: Compare the predicted risk state with the risk state labels carried in the validation dataset, and determine the first recall rate corresponding to the feature data sample whose risk state label is a risk state based on the comparison result.
[0117] The first recall rate refers to the percentage of samples in the validation dataset that the model can correctly identify and output the corresponding risk category for samples that actually belong to the risk category of concern. Specifically, after obtaining the predicted risk status of all samples in the validation dataset, these model-generated predictions are compared item by item with the original true risk category labels carried by each sample. Electronic devices will focus on the subset of samples that are identified as belonging to a specific risk category of concern in the true labels (i.e., risk status labels). Within this subset, the number of samples whose predictions are also correctly identified by the model as belonging to that risk category of concern is counted, and the proportion of this number to the total number of samples in this subset is calculated. This proportion reflects the model's success rate in capturing real-world risk objects of concern and is an important reference metric for measuring the model's warning sensitivity and the degree of risk omission.
[0118] Step e5: If the area value is not less than the preset area threshold and the first recall rate is not less than the preset recall rate threshold, then the risk warning model is determined to have passed the verification.
[0119] The preset area threshold refers to a reference limit set for the area value to determine whether the model's discrimination ability meets the standard; for example, it can be 0.8. When the calculated area value is not lower than this reference limit, the model is considered to meet the predetermined requirements in terms of overall discrimination ability.
[0120] The preset recall threshold refers to a reference limit set for the first recall rate to determine whether the model's ability to identify the category of interest meets the standard. For example, it can be 75%. When the calculated recall rate is not lower than this reference limit, the model is considered to be able to effectively capture the preset risk events.
[0121] For example, as shown in Table 1, Table 1 is a cross-statistic table of risk predictions for the model validation set. The rows (Bankrupt) represent the actual risk status of construction companies, where 0 indicates a normal company and 1 indicates a distressed company. The columns (Predicted_Class) represent the predicted risk categories output by the model, where 0 indicates the model predicts a normal company and 1 indicates the model predicts a distressed company. The values in the table represent the number of company samples in the corresponding category. Specific data statistics are as follows: There were a total of 287 real normal enterprises (Bankrupt=0), of which 263 were correctly predicted as normal enterprises by the model (true negatives) and 24 were incorrectly predicted as distressed enterprises (false positives). There were a total of 26 truly distressed companies (Bankrupt=1), of which 21 were correctly predicted as distressed companies by the model (true positives), and 5 were incorrectly predicted as normal companies (false negatives). The model predicts a total of 313 companies, including 268 companies that are in good standing and 45 companies that are in distress.
[0122] Based on the data in the table, the recall rate of distressed enterprises was calculated as follows: Recall rate = number of true positive samples ÷ total number of real distressed enterprises = 21 ÷ 26 ≈ 80.77%. This result meets the validation requirement of ≥75% recall rate for distressed enterprises in this model.
[0123] Table 1
[0124] Specifically, the electronic device pre-stores two reference thresholds for determining whether the model's performance meets the application requirements: one for the area under the aforementioned feature curve, and the other for the success rate of identifying the aforementioned category of interest. After completing the evaluation calculations for the validation dataset, the actually calculated area value and the first recall rate are compared numerically with their corresponding reference thresholds. Only when the area value reaches or exceeds its threshold, and the first recall rate also reaches or exceeds its threshold, is the overall performance of the risk warning model deemed to meet the preset quality requirements and thus pass validation. This dual-condition setting ensures that the model possesses both good overall discrimination ability and effectively controls the risk of missing key risk events.
[0125] In the above implementation, by obtaining a validation dataset obtained through stratified random sampling from the same source as the training data, the consistency between the distribution of validation samples and the distribution of modeling samples is ensured, thereby enabling objective verification of the model's true generalization performance on data not used in the training. Simultaneously, by plotting the target feature curve and calculating the area under the curve, the overall distinguishing ability of the model for samples in different risk states is quantitatively evaluated. Based on this, the validation dataset is input into the model to obtain predicted risk states, which are then compared with the real labels. The first recall rate for risk state samples is specifically calculated to accurately measure the model's coverage of actual distressed enterprises. Meeting both the area value and the recall rate at preset thresholds is a dual necessary condition for the model to pass validation, balancing the model's overall distinguishing performance with its sensitivity to high-risk samples. This ensures that only models with both good comprehensive distinguishing ability and key risk identification capability can be applied in practice, effectively improving the reliability and business usability of the early warning results.
[0126] In some optional implementations, the training process of the risk warning model further includes: Step f1: If the performance of the risk warning model fails to pass the verification, adjust the indicator screening parameters. The indicator screening parameters include at least one of the preset relevance threshold, the second preset probability threshold, and the preset measurement threshold.
[0127] The indicator screening parameters refer to a set of adjustable control parameters used to control the rigor and direction of the screening process during the feature screening stage of the risk warning model. These parameters include, but are not limited to, various reference thresholds related to relevance, significance, and redundancy determination. Specifically, if either the area value or the first recall rate of the model fails to reach the corresponding reference threshold in the above verification stage, it indicates that the performance of the current model does not meet the predetermined standard. At this time, a model reconstruction process will be automatically initiated. The starting point of this process is to adjust the control parameters used in the early feature selection stage of model training. These control parameters include, but are not limited to: reference thresholds for initial screening of feature association strength, reference thresholds for evaluating the statistical reliability of features, and reference thresholds for eliminating redundant features. According to the preset adjustment direction (for example, if the model performance is insufficient, some thresholds will be appropriately relaxed to include more features; if the model performance is unstable, some thresholds will be tightened to reduce feature redundancy), one or more of these control parameters will be numerically modified.
[0128] For example, if the model underfits, the preset correlation threshold can be appropriately relaxed (e.g., reduced from 0.3 to 0.2) or the second preset probability threshold can be relaxed (e.g., reduced from 0.05 to 0.1); if the model overfits, the preset measurement threshold can be tightened (e.g., the Variance Inflation Factor (VIF) threshold can be reduced from 10 to 5) to eliminate more redundant indicators.
[0129] Step f2: Using the adjusted indicator screening parameters, re-screen each initial warning indicator sample.
[0130] After adjusting the feature selection control parameters, the previously selected feature subset based on the old parameters will be discarded. Using the updated control parameters, the original, unselected initial warning indicator sample set will be re-executed with a complete automated feature optimization process. This process follows the same logic as the initial selection, sequentially performing initial screening based on correlation strength, secondary screening based on statistical reliability, and feature redundancy removal. However, because the control parameter values have changed, the composition of the final retained feature subset (i.e., the new target warning indicator sample) will differ from the previous round. Based on this new feature subset, data will be re-partitioned, the model will be rebuilt and trained, and the model will re-enter the validation phase to assess whether the new model meets performance requirements. This closed-loop optimization process can be iterated continuously until the model performance meets the validation criteria.
[0131] In the above implementation, when the performance of the risk warning model does not meet the verification criteria, the preset relevant threshold, the second preset probability threshold, and the preset measurement threshold are adjusted as the index screening parameters. The initial warning index samples are then re-screened using the adjusted parameters. This achieves closed-loop feedback and automatic optimization in the feature selection process during model construction, avoiding the subjective intervention and trial-and-error costs caused by manually redesigning indicators or blindly adding or deleting features. As a result, the model can adaptively converge to a better combination of indicators and performance in a data-driven manner, ensuring the prediction accuracy and business applicability of the final output model.
[0132] In the following embodiment, the above business processing method will be illustrated by taking into account the specific application scenario of "a bank using a risk warning model to conduct batch risk assessments on existing construction enterprise customers in the bank".
[0133] First, the system automatically collects multi-dimensional characteristic data from various construction companies through data interfaces, covering financial and operational data, project production progress data, industry benchmark data, and basic enterprise attribute data. Among these, industry-specific indicators such as completion value confirmation time, confirmation value receipt time, project progress deviation rate, construction delay days, and building material price fluctuation coefficient are the core features that distinguish this model from general credit scoring models.
[0134] After data collection, a preprocessing process is automatically executed, including missing value imputation, outlier correction, and standardization, to form a standardized dataset that can be directly used for model calculation.
[0135] The risk warning model automatically optimizes initial warning indicators based on a pre-set three-step progressive screening mechanism. First, correlation analysis filters indicators significantly associated with the risk status. Next, stepwise regression filters statistically significant indicators. Finally, collinearity diagnosis eliminates highly redundant indicators, ultimately determining a concise and efficient set of target warning indicators. During this process, industry-specific indicators such as the time to receive revenue and project schedule deviation rates are retained, demonstrating the model's accurate capture of industry-specific risk transmission paths.
[0136] After inputting the target early warning indicator data of the enterprise to be evaluated into the already trained Logistic regression model, the model automatically calculates and outputs the probability value of the enterprise's future predicament based on the objective weight coefficients obtained through learning driven by historical sample data.
[0137] The model further compares the risk probability value with a preset warning threshold. This threshold is not fixed and can be flexibly adjusted by experts based on actual risk tolerance. The model also supports automatically optimizing the threshold based on historical sample data to achieve the best balance between sensitivity and specificity. If the risk probability value reaches or exceeds the threshold, the enterprise is classified as high-risk; otherwise, it is classified as low-risk.
[0138] While outputting the warning level, the model also automatically identifies the key influencing factors that lead to the high-risk judgment. By comprehensively measuring the weight contribution of each indicator in the model, its statistical significance, and the deviation of the company's actual value from the industry benchmark, the model selects high-risk indicators and generates personalized warning reports.
[0139] The early warning report not only clearly defines the risk probability value and warning level, but also provides actionable risk management strategy recommendations for high-risk indicators. Strategy generation combines indicator type matching with tiered adjustments based on the degree of abnormal deviation, making the recommendations more targeted and actionable.
[0140] Through the above process, this application achieves fully automated processing from data collection, industry-appropriate indicator selection, model calculation to interpretable early warning output. The entire process does not rely on manual experience to set weights or select indicators, effectively overcoming the shortcomings of traditional assessment methods such as strong subjectivity, insufficient industry adaptability, delayed early warning, and poor interpretability, and significantly improving the ability to identify risks in advance and the efficiency of decision-making response for construction companies.
[0141] This embodiment also provides a risk prediction device for a building construction object, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0142] This embodiment provides a risk prediction device for building construction objects, such as... Figure 4 As shown, it includes: The first acquisition module 401 is used to acquire multi-dimensional feature data of the building construction object. The multi-dimensional feature data includes various attribute data that affect the risk status of the building construction object. Extraction module 402 is used to extract indicator data corresponding to the target warning indicators from multi-dimensional feature data based on the target warning indicators in the risk warning model. The second acquisition module 403 is used to acquire the weight coefficients of each target early warning indicator in the risk early warning model; The first determining module 404 is used to determine the risk probability value of the building construction object based on the data of each indicator and its corresponding weight coefficient. The second determining module 405 is used to determine the risk prediction result of the building construction object based on the target interval where the risk probability value is located; The risk warning model is trained based on a preset classification method.
[0143] In some alternative implementations, the second determining module 405 includes: The first determination submodule is used to determine the first warning level of the building construction object if the risk probability value is less than the preset risk probability threshold. The second determination submodule is used to determine the second early warning level of the building construction object if the risk probability value is not less than the preset risk probability threshold. The risk prediction results include a first warning level and a second warning level.
[0144] In some alternative implementations, the risk prediction device for the building construction object further includes: The third acquisition module is used to acquire the first significant probability value of each target early warning indicator in the risk early warning model, as well as the deviation of the construction object under each target early warning indicator. The first screening module is used to select first warning indicators from various target warning indicators, namely, indicators whose weight coefficient is greater than a preset weight threshold, whose first significance probability value is less than a first preset probability threshold, and whose indicator deviation is greater than a preset deviation threshold.
[0145] In some alternative implementations, the risk prediction device for the building construction object further includes: The fourth acquisition module is used to acquire the indicator type to which the first early warning indicator belongs, and the risk handling strategy that matches the indicator type; The first adjustment module is used to adjust the associated parameters of the risk treatment strategy based on the deviation of the indicators, so as to obtain the target treatment strategy for the construction object.
[0146] In some alternative implementations, the extraction module 402 includes: The first acquisition submodule is used to acquire a feature data sample set, which includes multiple feature data samples. Each feature data sample includes initial warning indicator samples of multiple dimensions and risk status labels corresponding to the feature data sample. The first filtering submodule is used to filter each initial warning indicator sample to obtain the target warning indicator sample; The sampling module is used to perform stratified random sampling of the feature data sample set according to the risk status label to obtain the training dataset corresponding to the feature data sample set. The submodule is used to construct a risk prediction model using the target early warning indicator sample as the independent variable and the risk status label as the dependent variable. The training submodule is used to train the risk prediction model based on the training dataset.
[0147] In some optional implementations, the first filtering submodule includes: The determination unit is used to determine the correlation coefficient between each initial warning indicator sample and the risk status label, and retain the first warning indicator sample whose correlation coefficient is not less than the preset correlation threshold. The analysis unit is used to perform regression analysis on each first warning indicator sample to obtain the second significance probability value corresponding to each first warning indicator sample, and retain the second warning indicator samples whose second significance probability value is less than the second preset probability threshold. The verification unit is used to perform collinearity verification on each second early warning indicator sample to obtain the collinearity measurement value corresponding to each second early warning indicator sample. The elimination unit is used to eliminate second warning indicator samples whose collinearity measurement value is not less than a preset measurement threshold, so as to obtain target warning indicator samples.
[0148] In some optional implementations, the extraction module 402 further includes: The validation submodule is used to perform multiple rounds of cross-validation on the training dataset to obtain the validation results for the training dataset. The evaluation submodule is used to obtain the evaluation index of the risk warning model based on the average value of the results of multiple rounds of validation. The evaluation metrics include at least one of average accuracy, average recall, and average precision.
[0149] In some optional implementations, the extraction module 402 further includes: The second acquisition submodule is used to acquire the verification dataset corresponding to the feature data sample set. The verification dataset is obtained by stratified random sampling of the feature data sample set based on the risk status label. The third determination submodule is used to draw the target feature curve based on the validation dataset and determine the area value under the target feature curve. The prediction submodule is used to input the validation dataset into the risk warning model to obtain the predicted risk status corresponding to each feature data sample in the validation dataset. The comparison submodule is used to compare the predicted risk status with the risk status labels carried in the validation dataset, and determine the first recall rate corresponding to the feature data sample whose risk status label is a risk status based on the comparison result. The fourth determination submodule is used to determine if the risk warning model passes the verification if the area value is not less than the preset area threshold and the first recall rate is not less than the preset recall rate threshold.
[0150] In some optional implementations, the extraction module 402 further includes: The adjustment submodule is used to adjust the indicator screening parameters if the performance of the risk warning model fails to pass the verification. The indicator screening parameters include at least one of the preset relevance threshold, the second preset probability threshold, and the preset measurement threshold. The second screening submodule is used to re-screen each initial warning indicator sample using the adjusted indicator screening parameters.
[0151] The risk prediction device for construction objects provided in this application can execute the risk prediction method for construction objects provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.
[0152] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0153] The following is a detailed reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0154] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0155] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the risk prediction method for construction objects according to embodiments of this application.
[0156] Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0157] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the risk prediction method for construction objects shown in the above embodiments is implemented.
[0158] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0159] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for risk prediction of a building construction object, characterized in that, The method includes: Obtain multi-dimensional feature data of the building construction object, the multi-dimensional feature data including various attribute data that affect the risk status of the building construction object; Based on the target early warning indicators in the risk early warning model, extract indicator data corresponding to the target early warning indicators from the multi-dimensional feature data; Obtain the weight coefficients of each of the target early warning indicators in the risk early warning model; Based on the data of each indicator and its corresponding weight coefficient, the risk probability value of the building construction object is determined; Based on the target range in which the risk probability value falls, the risk prediction result for the building construction object is determined; The risk warning model is trained based on a preset classification method.
2. The method according to claim 1, characterized in that, The process of determining the risk prediction result of the construction object based on the target interval where the risk probability value falls includes: If the risk probability value is less than the preset risk probability threshold, then the first warning level of the building construction object is determined; If the risk probability value is not less than the preset risk probability threshold, then the second early warning level of the building construction object is determined; The risk prediction results include the first warning level and the second warning level.
3. The method according to claim 1, characterized in that, The method further includes: Obtain the first significant probability value of each of the target early warning indicators in the risk early warning model, and the deviation of the construction object under each of the target early warning indicators; From the various target early warning indicators, select the first early warning indicator whose weight coefficient is greater than a preset weight threshold, whose first significance probability value is less than a first preset probability threshold, and whose indicator deviation is greater than a preset deviation threshold.
4. The method according to claim 3, characterized in that, The method further includes: Obtain the indicator type to which the first early warning indicator belongs, and the risk handling strategy that matches the indicator type; Based on the deviation of the aforementioned indicators, the associated parameters of the risk management strategy are adjusted to obtain the target management strategy for the construction object.
5. The method according to claim 1, characterized in that, The training process of the risk warning model includes: Obtain a feature data sample set, which includes multiple feature data samples. Each feature data sample includes multiple dimensions of initial early warning indicator samples and a risk status label corresponding to the feature data sample. By filtering each of the initial early warning indicator samples, the target early warning indicator samples are obtained; The feature data sample set is stratified and randomly sampled according to the risk status label to obtain the training dataset corresponding to the feature data sample set. The risk prediction model is constructed using the target early warning indicator sample as the independent variable and the risk status label as the dependent variable. The risk prediction model is trained based on the training dataset.
6. The method according to claim 5, characterized in that, The step of filtering each of the initial early warning indicator samples to obtain the target early warning indicator samples includes: Determine the correlation coefficient between each of the initial warning indicator samples and the risk status label, and retain the first warning indicator sample whose correlation coefficient value is not less than a preset correlation threshold among the initial warning indicator samples; Regression analysis is performed on each of the first warning indicator samples to obtain the second significance probability value corresponding to each of the first warning indicator samples, and the second warning indicator samples whose second significance probability value is less than the second preset probability threshold are retained. Perform collinearity verification on each of the second early warning indicator samples to obtain the collinearity measurement value corresponding to each of the second early warning indicator samples; The target early warning indicator sample is obtained by removing the second early warning indicator sample whose collinearity metric value is not less than a preset metric threshold.
7. The method according to claim 5, characterized in that, The method further includes: Perform multiple rounds of cross-validation on the training dataset to obtain the multi-round validation results corresponding to the training dataset; The evaluation index of the risk warning model is obtained based on the average value of the multi-round verification results. The evaluation metrics include at least one of average accuracy, average recall, and average precision.
8. The method according to claim 5, characterized in that, The method further includes: Obtain the verification dataset corresponding to the feature data sample set, wherein the verification dataset is obtained by stratified random sampling of the feature data sample set based on the risk status label; Based on the verification dataset, a target feature curve is plotted, and the area value under the target feature curve is determined. Input the verification dataset into the risk warning model to obtain the predicted risk status corresponding to each feature data sample in the verification dataset; The predicted risk state is compared with the risk state label carried in the validation dataset, and the first recall rate corresponding to the feature data sample whose risk state label is a risk state is determined based on the comparison result. If the area value is not less than a preset area threshold and the first recall rate is not less than a preset recall rate threshold, then the risk warning model is determined to have passed the verification.
9. The method according to claim 8, characterized in that, The method further includes: If the performance of the risk warning model fails to pass the verification, the indicator screening parameters are adjusted. The indicator screening parameters include at least one of a preset relevance threshold, a second preset probability threshold, and a preset measurement threshold. Using the adjusted index screening parameters, the initial early warning index samples were re-screened.
10. A risk prediction device for a building construction object, characterized in that, The device includes: The first acquisition module is used to acquire multi-dimensional feature data of the building construction object, the multi-dimensional feature data including various attribute data that affect the risk status of the building construction object; The extraction module is used to extract indicator data corresponding to the target warning indicator from the multi-dimensional feature data based on the target warning indicator in the risk warning model. The second acquisition module is used to acquire the weight coefficients of each of the target early warning indicators in the risk early warning model; The first determining module is used to determine the risk probability value of the building construction object based on each of the indicator data and its corresponding weight coefficient; The second determining module is used to determine the risk prediction result of the building construction object based on the target interval in which the risk probability value is located; The risk warning model is trained based on a preset classification method.
11. A computer program product, characterized in that, It includes computer instructions for causing a computer to execute the risk prediction method for a building construction object as described in any one of claims 1 to 9.