Engineering bidding risk decision-making method and system based on internal audit data fusion
By integrating internal audit data, constructing a risk-response mapping library, and employing machine learning models, the problem of the disconnect between company-level and project-level risks was solved, enabling intelligent decision support for bidding risks and improving the accuracy and efficiency of decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA RAILWAY 23RD BUREAU GRP NO 1 ENG
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies suffer from a disconnect between company-level and project-level risk assessments and insufficient utilization of internal audit information, making it difficult to provide scientific data-driven decision support during the bidding stage.
By acquiring project-level internal audit reports, bidding project texts, and basic structured features of projects, we extract quantitative indicators and process the text to construct a structured audit indicator and standardized audit text corpus. We then use fuzzy hierarchical analysis to generate a company-level risk index, utilize natural language processing technology to mine risk event features and a risk-response mapping library, and employ machine learning models for risk prediction and decision support.
It has achieved the correlation and integration of multi-source risk data, improved the consistency and scientific nature of bidding risk identification, provided data-driven intelligent decision support, and ensured the accuracy and efficiency of bidding decisions.
Smart Images

Figure CN121961704A_ABST
Abstract
Description
A Method and System for Engineering Bidding Risk Decision-Making Based on Internal Audit Data Fusion Technical Field
[0001] This invention relates to the field of contracting enterprise management and engineering project risk control, and in particular to a method and system for engineering bidding risk decision-making based on internal audit data fusion. Background Technology
[0002] As the construction industry enters a stage of high-quality development, construction contracting companies face a more complex market and project environment, making risk prevention and control a core requirement for their operations and management. Construction projects are inherently characterized by their wide scope, large investment amounts, long construction periods, numerous participating entities, and complex contract chains, leading to the coexistence of multiple risks related to quality and safety, contract management, and financial management. Furthermore, the risk transmission of construction projects is strong; loss of control in a single project can trigger cost overruns, broken cash flow, and other problems, ultimately evolving into systemic risks for the company. In addition, with the continuous expansion of audit and supervision coverage, management deficiencies and compliance issues exposed by construction contracting companies during operations are on the rise, necessitating efficient internal control mechanisms to prevent risks. This places higher demands on companies' internal control and risk management capabilities.
[0003] As the starting point for risk management throughout the entire project lifecycle, the bidding stage presents a crucial challenge for construction companies: accurately identifying potential risks and making informed decisions within a limited timeframe. Currently, construction companies generally establish internal audit systems, accumulating a wealth of internal audit reports containing core information such as risk events, management deficiencies, and compliance issues through systematic review and evaluation of business activities, project management practices, and historical cases. These reports reflect the risk profile of past projects and provide important historical references for identifying risks in new projects. However, current internal audit reports are mostly stored in unstructured text format, lacking a standardized reference system to support bidding decisions. With the rapid development of artificial intelligence and the maturity of technologies such as big data, natural language processing, and machine learning, new technical pathways have been provided for the in-depth utilization of internal audit information. Mining and integrating project-level internal audit report text data to construct intelligent bidding risk identification models has become an important direction for engineering informatization and digitalization.
[0004] The existing project bidding risk assessment methods still have significant shortcomings: (1) Company-level risk assessment and project-level risk assessment are separated. Although internal audit reports are generated for completed projects and projects under construction, the report results are mostly used for post-project supervision or management improvement. A risk assessment linkage mechanism combining the company's historical project performance and the characteristics of new projects has not been established at the bidding stage; (2) The utilization efficiency of the text information in project-level internal audit reports is low. The basic project information (such as project type, size, construction period, etc.) and audit results (such as profit margin, property management, personnel allocation, etc.) contained in the reports have important decision-making value. However, the traditional method relies on manual reading to extract information, which is not only inefficient but also greatly affected by subjective factors, making it difficult to achieve standardized extraction and utilization of information; (3) Historical internal audit data has not been transformed into reusable intelligent decision-making resources. A large number of project-level internal audit reports accumulated by enterprises have not yet been used to build effective risk prediction models, which cannot provide data-driven intelligent decision support for new project bidding. In summary, the existing internal audit risk management methods are difficult to meet the actual needs of construction companies to quickly provide scientific support for bidding decisions in the context of digital and refined management.
[0005] Existing patent CN117787569B discloses an intelligent auxiliary bidding evaluation method and system. The method includes: based on a preprocessed bid dataset, using an autoregressive integral moving average model to perform time series analysis on the data, extracting time trends and seasonal characteristics, identifying key stages of the bidding evaluation process, and generating predictions for key time nodes in the bidding evaluation process. However, this invention does not integrate internal audit data to achieve company-level and project-level risk linkage assessment, does not fully explore the structured and unstructured risk characteristics in audit reports, and lacks a risk-response mapping library. Therefore, it is insufficient in the targeted use of bidding risk data, the linkage of risk assessment, and the accuracy of decision support. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of existing technologies, such as the separation of company-level and project-level risks and insufficient utilization of internal audit information, and to provide a method and system for engineering bidding risk decision-making based on internal audit data fusion.
[0007] In a first aspect, the present invention provides a method for engineering bidding risk decision-making based on internal audit data fusion. The method includes the following steps: S1, obtaining project-level internal audit reports, bidding project text materials, and basic structured features of the project; the project-level internal audit reports are in at least one of PDF, Word, and plain text formats; the bidding text materials include tender documents and contract texts; the basic structured features of the project include project type, construction scale, investment amount, construction period, and geographical attributes; S2, performing quantitative indicator extraction and text processing on the project-level internal audit reports respectively, outputting structured audit indicators and a standardized audit text corpus; based on the enterprise's preset risk assessment indicator system and the above-mentioned structured audit indicators, using fuzzy hierarchical analysis to calculate and generate company-level risk indices for each period; based on the above-mentioned standardized audit text corpus... The database and the aforementioned bidding project texts are processed sequentially using natural language processing technology, including text preprocessing, risk entity identification, risk keyword matching, risk event coding, rectification measure extraction, and risk-measure association mining, outputting risk event features and a risk-measure mapping library. Based on the aforementioned company-level risk indices for each period, the aforementioned basic structured features of the projects, and the aforementioned risk event features, after feature alignment, normalization, and fusion coding, a machine learning model is used for training and inference based on historical project datasets, outputting predicted risk levels, predicted profit margins, and feature contribution. S3: Based on the aforementioned predicted risk levels and predicted profit margins, combined with a preset standard decision-making system, risk warning results are output; based on the aforementioned feature contribution, a risk source list is output; based on the aforementioned risk source list, combined with the aforementioned risk-measure mapping library, targeted measures are output.
[0008] Based on multi-source data integration and acquisition, dual-dimensional processing of audit data, company-level risk index calculation, full-process text mining, multi-feature fusion machine learning prediction, and full-link decision output, this technology effectively solves the problems of the separation between company-level and project-level risks and the insufficient utilization of audit information in existing technologies. It realizes the integration of multi-source risk data and the joint assessment of enterprise-project level risks, improves the consistency, scientificity, and decision-making efficiency of bidding risk identification, and provides data-driven intelligent support for the bidding stage.
[0009] Preferably, the quantitative indicator extraction in S2 includes: using regular expression matching and table recognition technology to extract profit margin, cost deviation, number of personnel management issues, and material management violation records from the project-level internal audit report, converting them into CSV format structured fields, and generating a structured audit indicator matrix; the text processing in S2 includes: performing segmentation, word segmentation, and stop word filtering operations on the text description part of the project-level internal audit report, constructing a dictionary of engineering-specific terms, and forming a standardized audit text corpus.
[0010] By employing regular expression matching combined with table recognition technology to extract structured indicators and convert them into CSV format, performing segmentation, word segmentation, and stop word filtering on the audit report text, and constructing an engineering terminology dictionary, we can accurately and efficiently extract structured indicators from the audit report, avoiding the inefficiency of manual parsing. At the same time, we can build a standardized audit text corpus to provide high-quality data support for subsequent risk feature extraction and rectification measure mining.
[0011] Preferably, the process of generating the company-level risk index for each period using the fuzzy hierarchical analysis method in S2 above includes: normalizing the above-mentioned structured audit indicators for multiple periods to the [0,1] interval; constructing a triangular fuzzy judgment matrix, solving the fuzzy weights of each level of indicators in the above-mentioned enterprise's preset risk assessment indicator system and performing consistency checks; calculating the quarterly or annual fuzzy risk score using the fuzzy comprehensive evaluation method, and generating the above-mentioned company-level risk index after defuzzification processing.
[0012] By normalizing multi-period structured audit indicators to the [0,1] interval, constructing a triangular fuzzy judgment matrix, solving the fuzzy weights of hierarchical indicators and performing consistency checks, and using the fuzzy comprehensive evaluation method to calculate risk scores and defuzzify them, the differences in the dimensions and scoring scales of different indicators are eliminated, the uncertainty in risk assessment is quantified, and a scientific and credible company-level risk index is generated, providing a macro-risk linkage basis for project-level risk assessment.
[0013] Preferably, when performing the consistency check, the consistency ratio is set to be less than 0.1.
[0014] Based on the technical feature of setting a consistency test standard with a consistency ratio of less than 0.1, the logical rationality boundary of the weight allocation of clear hierarchical indicators is obtained, avoiding logical contradictions in weight calculation, ensuring the rigor and reliability of company-level risk index accounting, and providing reliable macroeconomic risk data for subsequent multi-feature fusion prediction.
[0015] Preferably, the risk event coding in S2 above includes: performing binarization or counting coding on the identified risk events according to a preset classification system.
[0016] Based on a pre-defined classification system, the identified risk events are binarized or encoded using counting methods to transform unstructured risk events into standardized, computable risk event feature vectors. This enables the quantitative expression of risk information, facilitating synergistic integration with company-level risk indices and project basic characteristics, thereby improving the training efficiency and prediction accuracy of machine learning models.
[0017] Preferably, the extraction of rectification measures in S2 above includes: locating the "Audit Opinions", "Rectification Requirements" and "Handling Suggestions" paragraphs in the project-level internal audit report above, and using keyword extraction, dependency analysis and conditional action pattern recognition to extract standardized rectification measure phrases.
[0018] By locating the "Audit Opinions," "Rectification Requirements," and "Handling Suggestions" paragraphs in project-level internal audit reports, and employing keyword extraction, dependency analysis, and conditional action pattern recognition to extract standardized rectification measure phrases, this approach can accurately uncover rectification measures within audit reports. This avoids omissions and non-standardization in manual extraction, creating a high-quality rectification measure resource pool and providing precise materials for the construction of a risk-measure mapping library and the recommendation of targeted measures.
[0019] Preferably, the process of constructing the risk-measure mapping library in S2 above includes: based on the co-occurrence relationship between risk events and rectification measures in the historical project-level internal audit reports, establishing the correspondence between risk events and rectification measures through co-occurrence matrix calculation and Apriori association rule mining.
[0020] Based on the co-occurrence relationship between risk events and corrective measures in historical project-level internal audit reports, a corresponding relationship is established through co-occurrence matrix calculation and Apriori association rule mining. This breaks through the limitations of traditional manual matching, constructs a precise mapping relationship between risks and measures, supports flexible one-to-many and many-to-many matching, provides core technical support for the rapid output of targeted corrective measures for decision-making, and improves the accuracy of risk disposal recommendations.
[0021] Preferably, the machine learning model in S2 is an XGBoost model, a random forest model, or a gradient boosting decision tree model. The training process of the machine learning model optimizes the hyperparameters through 5-fold cross-validation. The hyperparameters include the learning rate and the tree depth. The feature contribution in S2 is obtained by analyzing the feature importance ranking or Shapley value decomposition method.
[0022] By selecting XGBoost, Random Forest, or Gradient Boosting Decision Tree models, and optimizing hyperparameters such as learning rate and tree depth through 5-fold cross-validation, and analyzing feature contribution using feature importance ranking or Shapley value decomposition, we can improve the accuracy and generalization ability of risk level and profit margin predictions, achieve risk source tracing, and make the prediction results interpretable. This solves the black box problem of traditional models and provides a clear direction for risk management.
[0023] Preferably, the targeted measures in S3 are at least one corresponding rectification measure obtained by matching at least one risk event in the risk source list; S3 also includes: integrating the risk warning results, the risk source list and the targeted measures to generate a structured bidding risk decision report.
[0024] By adopting the above matching logic and integrating risk warning results, risk source list and targeted measures to generate a structured bidding risk decision report, a multi-dimensional measure matching that supports complex risk scenarios is obtained, ensuring the comprehensiveness of rectification suggestions. At the same time, the core decision information is presented intuitively, reducing the information screening cost of bidding decisions and improving decision efficiency and operability.
[0025] In a second aspect, the present invention provides an engineering bidding risk decision-making system based on internal audit data fusion, which executes the aforementioned engineering bidding risk decision-making method based on internal audit data fusion when the system is running.
[0026] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a method and system for engineering bidding risk decision-making based on internal audit data fusion. By integrating project-level internal audit reports, bidding documents, and basic structured features of the project, it achieves multi-source risk data correlation and integration; it extracts structured indicators and standardizes text processing of project-level internal audit reports, improving the efficiency of audit data utilization; based on the enterprise's pre-set risk assessment indicator system and structured audit indicators, it uses fuzzy hierarchical analysis to calculate the company-level risk index for each period, quantifying the overall enterprise risk and establishing a foundation for enterprise-project level risk linkage; and it employs natural language processing based on a standardized audit text corpus and bidding documents. This technology mines risk event characteristics and establishes a risk-response mapping library, standardizing unstructured risk information and accurately matching risks with corresponding measures. Based on company-level risk indices, project-level structured features, and risk event characteristics, it integrates multiple features and trains machine learning models for inference, outputting predicted risk levels, predicted profit margins, and feature contribution, achieving quantitative prediction of risk and return and risk tracing. Finally, it outputs risk warning results, a list of risk sources, and targeted measures, accurately addressing the core issues of existing methods such as the separation of company-level and project-level risks and insufficient utilization of audit information. This improves the consistency, scientific rigor, and efficiency of bidding risk decisions, providing data-driven intelligent support for the bidding stage. Attached Figure Description
[0027] Figure 1 is a schematic diagram of the engineering bidding risk decision-making system architecture based on internal audit data fusion in Example 1.
[0028] Figure 2 is a schematic diagram of the internal audit data processing module in Example 1.
[0029] Figure 3 is a schematic diagram of the corporate-level risk assessment module structure in Example 1.
[0030] Figure 4 is a schematic diagram of the feature extraction module for project text risks and rectification measures in Example 1.
[0031] Figure 5 is a schematic diagram of the feature fusion and risk prediction module in Example 1.
[0032] Figure 6 is a schematic diagram of the intelligent decision support module structure in Example 1. Detailed Implementation
[0033] The present invention will now be described in further detail with reference to specific embodiments. However, this should not be construed as limiting the scope of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.
[0034] Unless otherwise specified, the terms "upper," "lower," "left," "right," "center," "inner," and "outer," etc., used in the description of specific embodiments of the present invention to indicate orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings, or the orientation or positional relationship in which the product / equipment / device is usually placed during use. These terms are merely for the purpose of facilitating the description of the present invention or simplifying the description in specific embodiments, and for enabling those skilled in the art to quickly understand the solution, and do not indicate or imply that a particular device / component / element must have a specific orientation, or be constructed and operated in a specific positional relationship. Therefore, they should not be construed as limitations on the present invention.
[0035] Furthermore, the use of terms such as "horizontal," "vertical," "suspended," "parallel," and "coaxial" does not imply that the corresponding device / component / element must be absolutely horizontal, vertical, suspended, parallel, or coaxial. Slight tilt or deviation is permissible, as long as it does not affect the normal function of the relevant component. For example, "horizontal" simply means that its direction is more horizontal relative to "vertical," not that the structure must be perfectly horizontal; a slight tilt is acceptable. "Coaxial" means that two components are arranged as coaxially as possible, allowing them to move coaxially or approximately coaxially when their relative positions change. Alternatively, it can be simplified to mean that the corresponding device / component / element, when arranged in "horizontal," "vertical," "suspended," "parallel," or "coaxial" directions, can have an error / deviation of ±10% relative to the corresponding direction, more preferably within ±8%, more preferably within ±6%, more preferably within ±5%, and more preferably within ±4%. For example, the deviation in the "coaxial" direction is controlled within 0.2-1mm, preferably within 0.2-0.5mm. As long as the corresponding device / component / element is within the error / deviation range, it can still achieve its function in the solution of the present invention.
[0036] Furthermore, the use of terms such as "first," "second," and "third" in terminology is merely for distinguishing descriptions of identical or similar components and should not be interpreted as emphasizing or implying the relative importance of a particular component.
[0037] Furthermore, in the description of the embodiments of the present invention, "several", "more than", and "a number of" represent at least two. The number can be any number, such as two, three, four, five, six, seven, eight, or nine, and can even exceed nine.
[0038] Furthermore, in the description of the technical solution of this invention, unless otherwise explicitly specified / limited / restricted, the terms "set up," "install," "connect," "link," "provided with," "laid out," and "arranged" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to connection methods commonly used in the art, such as welding, riveting, bolting, and threaded connections. Such connections can be mechanical, electrical, or communication connections; they can be direct connections or indirect connections through an intermediate medium; and they can refer to the internal communication between two components.
[0039] Example 1 This example further illustrates the present invention using an engineering bidding risk decision-making system based on internal audit data fusion.
[0040] The engineering bidding risk decision-making method based on internal audit data fusion in this invention has three core steps: S1 acquiring multi-source data, S2 multi-dimensional data processing and feature mining, and S3 intelligent decision output. Its structure can be divided into four layers: data input layer, feature processing layer, model calculation layer, and decision output layer.
[0041] The risk decision-making method here can be implemented through five functional modules: internal audit data processing module, company-level risk assessment module, project-level text risk and rectification measure feature extraction module, feature fusion and risk prediction module, and intelligent decision support module. Each module is connected sequentially through data interfaces to ensure data flow and functional synergy.
[0042] The overall architecture of the system is shown in Figure 1. In the figure, "internal audit data" refers to project-level internal audit reports and bidding project documents, and "project structured information" refers to the basic structured features of the project.
[0043] 1. Relationship between modules: "Internal Audit Data Processing Module" and "Company-level Risk Assessment Module": The structured audit indicators output by the Internal Audit Data Processing Module are transmitted to the Company-level Risk Assessment Module through the data interface as input parameters for "generating company-level risk indices for each period using fuzzy hierarchical analysis" as mentioned above, and are used to calculate the company-level risk index for the quarter or year.
[0044] The "Internal Audit Data Processing Module" and the "Project-Level Text Risk and Rectification Measure Feature Extraction Module" are used to extract standardized audit text corpora output by the Internal Audit Data Processing Module. These corpora are then transmitted to the Project-Level Text Risk and Rectification Measure Feature Extraction Module via a text interface to perform natural language processing operations and generate a risk event feature and risk-measure mapping library.
[0045] The “Project-level Text Risk and Rectification Measures Feature Extraction Module” and the “Feature Fusion and Risk Prediction Module”: The risk event features (T) and risk-measure mapping library (M) output by this module are passed to the feature fusion and risk prediction module through the feature interface. Together with the company-level risk index (Rt) and the project basic structured features (P), they constitute the complete feature set required by the machine learning model.
[0046] "Company-level Risk Assessment Module" and "Feature Fusion and Risk Prediction Module": The company-level risk index output by the company-level risk assessment module is directly input into the feature fusion and risk prediction module through the numerical feature interface to participate in the model training and inference process.
[0047] The “Feature Fusion and Risk Prediction Module” and the “Intelligent Decision Support Module”: The predicted risk level, predicted profit margin and feature contribution output by this module are transmitted to the intelligent decision support module through the model interface to provide data support for risk warning and response suggestion generation.
[0048] The "Project-level Text Risk and Rectification Measures Feature Extraction Module" and the "Intelligent Decision Support Module": The risk-measure mapping library is directly transferred to the intelligent decision support module through a knowledge interface, ensuring that the module can match the corresponding rectification measures based on the type of risk event.
[0049] 2. Processing flow of each module: (1) Internal audit data processing module, the structure of which is shown in Figure 2: This module corresponds to the method of "extracting quantitative indicators and processing text for the project-level internal audit reports respectively, and outputting structured audit indicators and standardized audit text corpus". This module specifically includes an audit report acquisition unit, a format conversion unit, a structured indicator parsing unit, an audit text preprocessing unit and an audit data output unit connected in sequence. Each unit is connected in sequence through a data interface, and the process is as follows: ① Audit report acquisition unit: Obtain the original file of the project-level internal audit report from the enterprise internal audit system, document management system or database. The file format includes at least one of PDF format, Word format and plain text format.
[0050] ② Format Conversion Unit: Converts the original audit report file into a parsable, unified text format, such as PDF and Word, and standardizes character encoding to ensure readability and consistency in subsequent parsing.
[0051] ③ Structured Indicator Parsing Unit: Based on regular expression matching and table recognition technology, and on the basis of a unified text format, it identifies and extracts structured audit indicators such as profit margin, cost deviation, number of personnel management issues, and material management violation records from the audit report, converts them into CSV format structured fields, and generates a structured audit indicator matrix.
[0052] ④ Audit text preprocessing unit: Performs segmentation, word segmentation, and stop word filtering operations on the text description part of the audit report, constructs a dictionary of engineering-specific terms, and forms a standardized audit text corpus.
[0053] ⑤ Audit Data Output Unit: The structured audit indicator matrix and standardized audit text corpus are output to the company-level risk assessment module and the project-level text risk and rectification measure feature extraction module, respectively, to achieve data sharing and transmission.
[0054] (2) Company-level risk assessment module, the structure of which is shown in Figure 3: This module corresponds to "based on the enterprise's preset risk assessment indicator system and the structured audit indicators, the fuzzy hierarchical analysis method is used to calculate and generate the company-level risk index for each period" in this method. This module includes an indicator system loading unit, an audit indicator normalization unit, a fuzzy hierarchical analysis weight calculation unit, a fuzzy comprehensive evaluation unit, and a risk index output unit connected in sequence. The implementation process is as follows: ① Indicator system loading unit: Loads the enterprise's preset three-level FAHP (fuzzy hierarchical analysis method) risk assessment indicator system (including first-level indicators, second-level indicators, third-level indicators and the hierarchical relationship between each indicator), and transmits it to the weight calculation unit in the form of a matrix or tree structure.
[0055] ② Audit indicator normalization unit: Normalizes the multi-period structured audit indicators from the internal audit data processing module and maps them uniformly to the [0,1] interval to eliminate the differences in dimensionality and scoring scale between indicators and meet the input requirements of fuzzy hierarchical analysis.
[0056] ③ Fuzzy Hierarchical Analysis Weight Calculation Unit: Construct a triangular fuzzy judgment matrix based on the indicator system, solve the fuzzy weights of each level of indicators, and perform a consistency test (consistency ratio < 0.1). Obtain a weight vector that can be used for risk calculation through defuzzification operation.
[0057] ④ Fuzzy Comprehensive Evaluation Unit: Based on the indicator weights and normalized scores, the fuzzy comprehensive evaluation method is used to calculate the company risk fuzzy score for each time period (such as quarter or year).
[0058] ⑤ Risk Index Output Unit: Defuzzifies the fuzzy comprehensive evaluation results, generates the final company-level risk index, and transmits it to the feature fusion and risk prediction module through the data interface as the input of macro risk features.
[0059] (3) Project-level text risk and rectification measure feature extraction module, the structure of which is shown in Figure 4: The module corresponds to the "based on the standardized audit text corpus and the bidding project text data, natural language processing technology is used to sequentially perform text preprocessing, risk entity identification, risk keyword matching, risk event coding, rectification measure extraction, and risk-measure association mining operations, and output risk event features and risk-measure mapping library" in this method. This module includes a text preprocessing unit, a risk entity identification unit, a risk keyword matching unit, a risk event coding unit, a rectification measure extraction unit, a risk-measure association mining unit, and a text feature output unit. Each unit is sequentially connected or bidirectionally connected through a data interface, and the process is as follows: ① Text preprocessing unit: The input standardized audit text corpus, bidding documents, contract texts and other bidding project text data are processed by segmentation, sentence segmentation and word segmentation, while stop word removal, part-of-speech tagging and terminology normalization operations are performed, and the processed text is transmitted to the risk entity identification unit and the rectification measure extraction unit.
[0060] ② Risk Entity Identification Unit: Identifies engineering risk-related entities in the text, including key area risks such as construction condition risks (complex geology, limited site, etc.), contract risks (unclear payment terms, vague change terms), and management risks (insufficient manpower, material delays), and generates entity tags.
[0061] ③ Risk Keyword Matching Unit: Based on the preset engineering risk dictionary, it matches risk expressions such as "complex", "delayed", "unclear" and "high risk" in the text to improve the accuracy of risk event identification.
[0062] ④ Risk Event Coding Unit: Integrates risk entity identification and keyword matching results, performs binarization or counting coding on various risk events according to a preset classification system, and forms a risk event feature vector, such as T1=payment risk, T2=personnel and organization risk, etc.
[0063] ⑤ Rectification Measures Extraction Unit: Locate the "Audit Opinions," "Rectification Requirements," and "Handling Suggestions" paragraphs in the project-level internal audit report, and use keyword extraction, dependency analysis, and conditional action pattern recognition (such as "should...", "needs to strengthen...") to extract standardized rectification measure phrases such as "supplementary geological exploration," "strengthen contract review," and "increase staffing."
[0064] ⑥ Risk-Measure Correlation Mining Unit: Based on the co-occurrence relationship between risk events and corrective measures in historical project-level internal audit reports, the corresponding relationship between risk events and corrective measures is established through co-occurrence matrix calculation and Apriori association rule mining, generating a risk-measure mapping library, such as "material risk - it is recommended to strengthen supply chain coordination" and "unclear contract terms - it is recommended to strengthen contract review".
[0065] ⑦ Text Feature Output Unit: Outputs risk event feature vectors and risk-measure mapping library, and passes them to the feature fusion and risk prediction module and the intelligent decision support module.
[0066] (4) Feature Fusion and Risk Prediction Module, the structure of which is shown in Figure 5: This module corresponds to the method of “based on the company-level risk index of each period, the basic structured features of the project and the features of the risk event, after feature alignment, normalization and fusion encoding, the machine learning model is used for training and inference based on the historical project dataset, and the predicted risk level, predicted profit margin and feature contribution are output”. This module includes a feature alignment unit, a feature normalization unit, a feature splicing and encoding unit, a model training unit, a risk prediction and classification unit and a model interpretation and feature contribution analysis unit. The units are connected in sequence, and the implementation process is as follows: ① Feature Alignment Unit: Receives the company-level risk index, the basic structured features of the project and the feature vector of the risk event, performs field alignment, missing value completion and time label matching on various features to ensure data consistency.
[0067] ② Feature normalization unit: Normalizes or standardizes numerical features, count features, and risk event features to ensure consistent feature dimensions and improve model training performance.
[0068] ③ Feature splicing and encoding unit: The company-level risk index, project basic structured features and risk event feature vectors are spliced to form a unified feature matrix. At the same time, one-hot, embedded encoding or binary encoding processing is performed on categorical features and text features.
[0069] ④ Model Training Unit: Based on historical project datasets, XGBoost, Random Forest, or Gradient Boosting Decision Tree models are trained using the mapping relationship of "feature matrix - actual profit margin / actual risk level". During the training process, hyperparameters such as learning rate and tree depth are optimized through 5-fold cross-validation.
[0070] ⑤ Risk Prediction and Classification Unit: Input the feature matrix of the new project into the trained model, and output the project risk level (such as high, medium, low) and the predicted profit rate.
[0071] ⑥ Model Interpretation and Feature Contribution Analysis Unit: Analyzes feature contribution by ranking feature importance or Shapley value decomposition, identifies key factors affecting risk prediction, generates interpretable results, and outputs them to the intelligent decision support module.
[0072] (5) Intelligent Decision Support Module, the structure of which is shown in Figure 6: This module corresponds to the method of “outputting risk warning results based on the predicted risk level and the predicted profit rate, combined with the preset standard decision system; outputting a risk source list based on the feature contribution; and outputting targeted measures based on the risk source list and the risk-measure mapping library”. This module includes a risk judgment unit, a risk warning unit, a risk tracing analysis unit, a rectification measure matching unit, a decision suggestion generation unit, and an output interface unit. Each unit is connected in sequence or transmits information through a data interface. The process is as follows: ① Risk Judgment Unit: Receives the project risk level, profit rate prediction value, and prediction confidence level output by the feature fusion and risk prediction module, and determines whether the project is in a high-risk or low-profit rate range by combining the preset standard decision system (including high / medium / low risk range and profit rate threshold).
[0073] ② Risk warning unit: When the risk assessment unit determines that the project is high-risk or the predicted profit rate is lower than the preset threshold, a risk warning is triggered and a risk prompt message is generated. At the same time, the warning status is transmitted to the risk source analysis unit and the rectification measure matching unit.
[0074] ③ Risk Source Analysis Unit: Receives key risk factors output by the Model Explanation and Feature Contribution Analysis Unit, analyzes textual risk events, structured risk features, or company-level risk indices that affect the prediction results, and forms a list of risk sources.
[0075] ④ Rectification Measures Matching Unit: Based on the risk events in the risk source list, the corresponding rectification measures are automatically matched from the risk-measure mapping library. It supports one-to-many and many-to-many mapping relationship processing and realizes multi-dimensional measure recommendations in complex risk scenarios. For example, when it is identified as "unclear contract payment terms", "strengthen contract terms review" is matched.
[0076] ⑤ Decision Recommendation Generation Unit: Integrates project risk level, risk source list and matching targeted measures to generate a structured bidding risk decision report.
[0077] ⑥ Output Interface Unit: Outputs risk warning results, risk source list, targeted measures, and structured bidding risk decision report, providing complete decision support for bidding managers.
[0078] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A risk decision-making method for engineering bidding based on internal audit data fusion, characterized in that, The method includes the following steps: S1, obtaining project-level internal audit reports, bidding project text materials, and basic structured features of the project; the project-level internal audit report is in at least one of PDF, Word, and plain text formats; the bidding project text materials include tender documents and contract texts; the basic structured features of the project include project type, construction scale, investment amount, construction period, and geographical attributes; S2, extracting quantitative indicators and processing the text of the project-level internal audit reports, outputting structured audit indicators and a standardized audit text corpus; based on the enterprise's preset risk assessment indicator system and the structured audit indicators, using fuzzy hierarchical analysis to calculate and generate company-level risk indices for each period; based on the standardized audit text corpus and the bidding project text materials... Natural language processing (NLP) technology is used to sequentially perform text preprocessing, risk entity identification, risk keyword matching, risk event encoding, rectification measure extraction, and risk-measure association mining operations, outputting a risk event feature and risk-measure mapping library. Based on the company-level risk index of each period, the basic structured features of the project, and the risk event features, after feature alignment, normalization, and fusion encoding, a machine learning model is used for training and inference based on historical project datasets to output predicted risk level, predicted profit margin, and feature contribution. S3: Based on the predicted risk level and predicted profit margin, combined with a preset standard decision-making system, a risk warning result is output; based on the feature contribution, a risk source list is output; based on the risk source list, combined with the risk-measure mapping library, targeted measures are output.
2. The engineering bidding risk decision-making method based on internal audit data fusion according to claim 1, characterized in that, The quantitative indicator extraction in S2 includes: using regular expression matching and table recognition technology to extract profit margin, cost deviation, number of personnel management issues, and material management violation records from the project-level internal audit report, converting them into CSV format structured fields, and generating structured audit indicators; the text processing in S2 includes: performing segmentation, word segmentation, and stop word filtering operations on the text description part of the project-level internal audit report, constructing a dictionary of engineering-specific terms, and forming a standardized audit text corpus.
3. The engineering bidding risk decision-making method based on internal audit data fusion according to claim 1, characterized in that, The process of generating the company-level risk index for each period using the fuzzy hierarchical analysis method in S2 includes: normalizing the structured audit indicators for multiple periods to the [0,1] interval; constructing a triangular fuzzy judgment matrix, solving the fuzzy weights of each level of indicators in the enterprise's preset risk assessment indicator system and performing a consistency test; calculating the quarterly or annual fuzzy risk score using the fuzzy comprehensive evaluation method, and generating the company-level risk index after defuzzification.
4. The engineering bidding risk decision-making method based on internal audit data fusion according to claim 3, characterized in that, When performing the consistency check, the consistency ratio is set to be less than 0.
1.
5. The engineering bidding risk decision-making method based on internal audit data fusion according to claim 1, characterized in that, The risk event coding in S2 includes: performing binarization or counting coding on the identified risk events according to a preset classification system.
6. The engineering bidding risk decision-making method based on internal audit data fusion according to claim 1, characterized in that, The extraction of corrective measures in S2 includes: locating the "Audit Opinions," "Corrective Requirements," and "Handling Suggestions" paragraphs in the project-level internal audit report, and using keyword extraction, dependency analysis, and conditional action pattern recognition to extract standardized corrective measure phrases.
7. The engineering bidding risk decision-making method based on internal audit data fusion according to claim 1, characterized in that, The construction process of the risk-measure mapping library in S2 includes: based on the co-occurrence relationship between risk events and corrective measures in the historical project-level internal audit reports, establishing the correspondence between risk events and corrective measures through co-occurrence matrix calculation and Apriori association rule mining.
8. The engineering bidding risk decision-making method based on internal audit data fusion according to claim 1, characterized in that, The machine learning model in S2 is an XGBoost model, a random forest model, or a gradient boosting decision tree model. The training process of the machine learning model optimizes the hyperparameters through 5-fold cross-validation. The hyperparameters include the learning rate and the tree depth. The feature contribution in S2 is obtained by analyzing the feature importance ranking or Shapley value decomposition.
9. The engineering bidding risk decision-making method based on internal audit data fusion according to claim 1, characterized in that, The targeted measures in S3 are at least one corresponding rectification measure obtained by matching at least one risk event in the risk source list; S3 also includes: integrating the risk warning results, the risk source list and the targeted measures to generate a structured bidding risk decision report.
10. A project bidding risk decision-making system based on internal audit data fusion, characterized in that: When the system is running, it executes the engineering bidding risk decision-making method based on internal audit data fusion as described in any one of claims 1 to 9.