Project audit evaluation method and device based on machine learning
By establishing a full-process tracking audit index system and a random forest algorithm model for the project, the problem of the existing audit system evaluation standards not being energetic is solved, and efficient and accurate audit evaluation of the ‘special bond + PPP’ model projects is achieved, ensuring the legality and efficiency of the project.
Patent Information
- Application Number
- CN202510306666.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-22
AI Technical Summary
The existing audit system evaluation standards have not been energetic, resulting in strong subjective arbitraryness and poor intuitiveness of the evaluation personnel, and the inability to effectively guarantee the legitimacy and efficiency of public infrastructure construction projects under the ‘special bond + PPP’ model.
Establish a comprehensive evaluation index system for the whole process of project tracking and auditing, and use machine-learning random forest algorithm to build an audit evaluation model. Through data preprocessing and model training optimization, standardize and objectify project audit evaluation.
It improves the accuracy and efficiency of audit evaluation, has high sensitivity and specificity of the model, and can effectively supervise projects under the ‘special bond + PPP’ model to ensure legitimacy and efficiency.
Smart Images

Figure CN120355350A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of construction project evaluation, and specifically relates to a project audit evaluation method and device based on machine learning. Background Art
[0002] Due to the advantages of local government special bonds, such as short issuance review cycle and low financing cost, the financing model of public infrastructure construction projects of the "special bond + PPP" mode has emerged as the times require and has become the main mode of government project investment and financing. How to ensure the legality of the operation of the "special bond + PPP" and guarantee the project benefits is crucial, and a set of suitable whole-process tracking audit index system and evaluation method are needed. The current audit system has the situation that the evaluation criteria cannot be quantified, resulting in strong subjectivity of evaluators, and the intuitiveness is poor. Summary of the Invention
[0003] In view of this, this application provides a project audit evaluation method and device based on machine learning, which solves the technical problems that the current audit system has unquantifiable evaluation criteria, resulting in strong subjectivity of evaluators, and the intuitiveness is poor.
[0004] To achieve the above object, the present invention provides the following technical solutions: A project audit evaluation method based on machine learning includes the following steps: First, according to the comprehensive evaluation index system of the whole-process tracking audit of the project established from three aspects: preliminary project establishment decision-making, mid-term supervision and management, and late-stage handover evaluation, obtain the audit evaluation index information corresponding to the project, and preprocess the audit evaluation index information to obtain standardized evaluation index data; The standardized evaluation index data includes the evaluation grade data of each index, which is convenient for subsequent processing by the audit evaluation model, and the project is audited and evaluated according to the processing result of the audit evaluation model.
[0005] Then construct an audit evaluation model based on the random forest algorithm, use all the indexes in the comprehensive evaluation index system of the whole-process tracking audit of the project as the characteristic factors of the random forest algorithm, use the audit evaluation result as the final model prediction result, and use the data set to train and optimize the audit evaluation model to obtain the optimized audit evaluation model; Finally, input the obtained standardized evaluation index data into the optimized audit evaluation model for project audit evaluation to obtain the project audit evaluation result.
[0006] Further, obtaining the audit evaluation index information corresponding to the project and preprocessing the audit evaluation index information includes: Obtain the audit evaluation index information corresponding to the project through a questionnaire, and clean, segment, remove stop words, replace synonyms, and perform word stemming on the descriptive text data in the audit evaluation index information to obtain a word segmentation list; Since the original text data of the expert questionnaire collected by the system is the descriptive answers of experts to a certain index, it is necessary to preprocess the data so that the model can further identify and process it. Text cleaning mainly includes removing HTML tags, special symbols, whitespace, and unifying case and punctuation formats. Word segmentation processing mainly includes using Jieba segmentation and LTP tools for Chinese word segmentation, and removing meaningless words such as "of", "in", "but", etc. through stop word filtering. After the above processing, it is also necessary to replace synonyms for common spoken language expressions, such as replacing "shortcoming" with "deficiency". Word stemming is similar to synonym replacement, such as replacing "outstanding performance" with "excellent performance".
[0007] Construct a domain dictionary by defining evaluation dimensions, establishing a sentiment word library, and setting modifier coefficients. Traverse each word in the processed word segmentation list and compare it with the domain dictionary. If the current word exists in the keyword list of any domain category, create or update the record of this dimension in the feature dictionary and initialize the score of this dimension to 0; The evaluation dimensions correspond to different evaluation indicators, and the number of dimensions is the number of indicators. For example, the descriptive keywords corresponding to the project initiation and screening procedure legality indicators include: ["project", "review", "competitiveness", "public"]. The sentiment word library includes positive: ["outstanding", "excellent", "remarkable"]; and negative: ["deficiency", "lack", "risk"], etc., and the corresponding weights are defined. After constructing the domain dictionary, perform subsequent weight value calculations and obtain the scores of the corresponding indicator dimensions.
[0008] After matching the domain category, continue to check whether the current word belongs to the sentiment word library. If it belongs to a positive or negative sentiment word, obtain the basic weight, read the basic score of this word from the predefined dictionary, and automatically convert the weight of the negative word to a negative value. Store the current word and its basic weight in the details list, and then check whether there is a word in the previous position of the current sentiment word. If the previous word belongs to the modifier library, read the modifier coefficient, and according to the formula: adjusted weight = basic weight x modifier coefficient, perform multiplicative correction on the basic weight; Accumulate the calculated weight values into the total score of the corresponding dimension, and perform feature quantization processing on the total scores of each dimension from high to low according to the preset matching range to obtain an effective sample data set Zm, Zm=(z1,…,zm).
[0009] Furthermore, construct an audit evaluation model based on the random forest algorithm. Training and optimizing the audit evaluation model using the data set includes: Taking the collected historical project audit evaluation index data Xn and the overall evaluation Yn of historical projects as the training sample S, S = (Xn, Yn), and taking the collected audit evaluation index data of the current project as the evaluation sample Zm, training the audit evaluation model based on the random forest algorithm through the training sample; Establishing the relationship between features through sample categories and non-linear fitting, and obtaining the evaluation result using the decision tree.
[0010] Furthermore, training the audit evaluation model based on the random forest algorithm through the training sample includes: Randomly selecting data and given the original training set , , , is the number of training set samples, randomly sampling with replacement times, drawing 1 each time, and the probability of a sample being drawn is 1 / , and generating different training subsets using this bootstrap sampling integration method, denoted as . Samples in different training subsets can be repeated, and samples in the same training subset can also be repeated. Training decision tree models one by one with different training subsets, denoted as . The models are independent of each other, and the samples not drawn form the validation set to detect the generalization ability; Randomly selecting features, each sample has F feature variables, specifying f < F, so that at each node, randomly select f feature variables from the F feature variables, find the best splitting node to split the decision tree, and do not prune the whole process. Then, according to the f features, calculate the best splitting method to improve the classification performance; Voting to integrate the models, using the regression analysis method to obtain the respective prediction results, and integrating the optimal RF model through voting; A project audit evaluation device based on machine learning, implemented using the above-mentioned project audit evaluation method based on machine learning, includes: a data preprocessing module, a model construction module, and an audit evaluation module; The data preprocessing module is used to obtain the audit evaluation index information corresponding to the project according to the comprehensive evaluation index system for the whole process tracking audit of the project, and preprocess the audit evaluation index information to obtain the standardized evaluation index data; The model construction module is used to construct an audit evaluation model based on the random forest algorithm. All the indicators in the comprehensive evaluation index system for the whole-process tracking audit of projects are used as the feature factors of the random forest algorithm, and the audit evaluation results are used as the final model prediction results. The audit evaluation model is trained and optimized using a data set to obtain an optimized audit evaluation model. The audit evaluation module is used to input the obtained standardized evaluation index data into the optimized audit evaluation model for project audit evaluation to obtain project audit evaluation results.
[0011] As can be seen from the above technical solutions, the advantages of the present invention are: 1. This application conducts audit prediction on "special bond + PPP" projects, establishes a compatible comprehensive evaluation index system and a suitable audit prediction model, applies the artificial intelligence algorithm of machine learning to audit evaluation prediction, establishes an RF audit evaluation prediction model, and can quickly obtain the optimal parameters to optimize the model and match them with the evaluation criteria of the index system. The model training and prediction operations are simple and have strong versatility. Moreover, the RF model has a good fitting degree during training, with high model sensitivity, specificity, and prediction accuracy, which can greatly improve the work efficiency of "special bond + PPP" audit evaluation prediction and has certain reference significance for promoting the full coverage of audit work.
[0012] 2. The audit evaluation method based on random forest in this application has high prediction accuracy, and the fitting degree of model training is good. The model sensitivity, specificity, and prediction accuracy are higher than those of audit evaluation prediction models such as SVM, Multinom, and BP, realizing effective supervision and audit prediction. And this application uses a text data processing method to quantitatively process descriptive indicators, making the evaluation results more objective and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application.
[0014] Figure 1 It is a schematic diagram of the steps of a project audit evaluation method based on machine learning according to this application.
[0015] Figure 2 It is a schematic diagram of the steps of step S2 of this embodiment.
[0016] Figure 3 It is a schematic diagram of the structure of the random forest model of this embodiment.
[0017] Figure 4 It is a schematic diagram of the composition structure of a project audit evaluation device based on machine learning according to this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the purpose, technical solutions, and advantages of this application clearer and more understandable, the following further elaborates on this application in combination with the implementation modes and accompanying drawings. Herein, the illustrative implementation modes of this application and their descriptions are used to explain this application, but do not limit this application.
[0019] It is crucial to ensure the legality of the operation of public infrastructure construction projects in the "special bond + PPP" model and guarantee project benefits. A set of compatible whole-process tracking audit index systems and audit evaluation models are required to achieve effective supervision and audit evaluation. Referring to Figures 1 to 4 , such as Figure 1 As shown, this embodiment provides a project audit evaluation method and device based on machine learning. This method is a method that applies text recognition methods and random forest machine learning methods to the whole-process tracking audit evaluation of public infrastructure construction projects in the "special bond + PPP" model (hereinafter referred to as "special bond + PPP"). First, identify the key points of project audit, establish a comprehensive evaluation index system for the whole-process tracking audit, making the indicators fit and cover comprehensively; then apply the random forest machine learning theory and methods to the whole-process tracking audit of "special bond + PPP"; finally, obtain the original sample data through expert evaluation questionnaires, and perform quantitative processing on the input and output data. Train and predict the established random forest (RandomForests, RF) model. The method specifically includes: Step S1: According to the comprehensive evaluation index system for the whole-process tracking audit of the project established from three aspects: preliminary project establishment decision-making, mid-term supervision and management, and post-transfer evaluation, obtain the audit evaluation index information corresponding to the project, and preprocess the audit evaluation index information to obtain standardized evaluation index data.
[0020] In this embodiment, combined with the actual application and work process of "special bond + PPP", based on the principle of full-life-cycle tracking audit coverage, extract the original evaluation indicators from three aspects of preliminary project establishment decision-making, mid-term supervision and management, and post-transfer evaluation as the whole-process tracking audit of "special bond + PPP", and construct a whole-process tracking audit evaluation index system including 68 third-level indicators, as shown in Table 1: Table 1
[0021] The specific steps of Step S1 are as follows. Obtaining the audit evaluation index information corresponding to the project and preprocessing the audit evaluation index information include: Obtain the audit evaluation index information corresponding to the project through a questionnaire, and clean, segment, remove stop words, perform synonym replacement, and lemmatization on the descriptive text data in the audit evaluation index information to obtain a word segmentation list; construct a domain dictionary by defining evaluation dimensions, establishing a sentiment word library, and setting modifier coefficients, traverse each word in the processed word segmentation list, and compare it with the domain dictionary. If the current word exists in the keyword list of any domain category, create or update the record of this dimension in the feature dictionary and initialize the score of this dimension to 0; after matching the domain category, continue to check whether the current word belongs to the sentiment word library. If it belongs to a positive or negative sentiment word, obtain the basic weight, read the basic score of this word from the predefined dictionary, and automatically convert the weight of the negative word to a negative value. Store the current word and its basic weight in the details list, and then check whether there is a word in the position before the current sentiment word. If the previous word belongs to the modifier library, read the modifier coefficient, and according to the formula: adjusted weight = basic weight × modifier coefficient, perform multiplicative correction on the basic weight; accumulate the calculated weight value into the total score of the corresponding dimension, and perform feature quantization processing on the total scores of each dimension from high to low according to the preset matching range to obtain an effective sample data set Zm, Zm=(z1,…,zm). The 69 index data in the collected questionnaire are all descriptive data, and they are subjected to feature quantization analysis and processing from high to low according to the evaluation level. For the C1-C68 indicators, 4 represents excellent, 3 represents good, 2 represents medium, and 1 represents poor.
[0022] Specifically, since the original text data of the expert questionnaire collected by the system is the descriptive answer of an expert to a certain index, it is necessary to preprocess the data so that the model can further identify and process it. Text cleaning mainly includes removing punctuation marks, numbers, meaningless symbols, and unifying case and punctuation formats. For example, if the input descriptive text is: "The compliance of the preliminary design standard of this project is relatively high, but the economic benefits of the project are relatively low", the output after cleaning is: "The compliance of the preliminary design standard of this project is relatively high, but the economic benefits of the project are relatively low". Word segmentation processing mainly includes using Jieba to perform Chinese word segmentation: ["this project", "preliminary design standard", "compliance", "relatively high", "but", "project economic benefits", "relatively low"], and remove meaningless words such as "of", "in", "but" through stop word filtering to get ["preliminary design standard", "compliance", "relatively high", "project economic benefits", "relatively low"]. After the above processing, it is also necessary to perform synonym replacement on common spoken language expressions. For example, replace "relatively high" with "excellent" and "relatively low" with "insufficient" to obtain the corresponding word segmentation list.
[0023] The construction of the domain dictionary includes dividing domain categories according to questionnaire information. For example, the demonstration of financial affordability in Table 1 includes ("financial expenditure", "cost of the entire project life cycle"), etc. Building an emotional word library is to define positive or negative emotional words and their weights, and the weight range is from 0 to 1. For example, positive ("excellent": 0.9, "remarkable": 0.8), negative ("insufficient": 0.7, "lagging": 0.75, "risk": 0.6). In addition, it is necessary to set modifiers and modification coefficients to handle the strengthening or weakening effect of degree adverbs on emotional words. For example, "extremely": 1.5, "very": 1.2, "relatively": 0.8. When performing modification calculations, attention needs to be paid to multiple modification and negation processing. For multiple modifications, only the last effective modifier is taken. For example, "very extremely excellent" only takes the modification coefficient corresponding to "extremely". Negation processing is to reverse the result negatively. For example, "not up to standard", "up to standard" (0.6) × (-0.1) = -0.6.
[0024] After constructing the domain dictionary, obtain the scores of the corresponding evaluation dimensions according to the comparison results between the word segmentation list and the domain dictionary. Feature quantization processing of the total scores of each dimension from high to low according to the preset matching range requires dividing grades according to the score ranges of each dimension. For example, if the score of a single dimension ≥ 0.8, it is classified as excellent; if 0.5 ≤ score < 0.8, it is classified as good; if score < 0.5, it is classified as medium; if the score is negative, it is classified as poor.
[0025] Step S2: Construct an audit evaluation model based on the random forest algorithm. Take all the indicators in the comprehensive evaluation index system of the whole-process follow-up audit of the project as the feature factors of the random forest algorithm, and take the audit evaluation result as the final model prediction result. Use the data set to train and optimize the audit evaluation model to obtain the optimized audit evaluation model.
[0026] In this embodiment, a real project of "special bonds + PPP" deposited in the Ministry of Finance, namely the special freight railway line project in Zouping City, Binzhou City, Shandong Province, is selected as a case. The three-level indicators C1 - C68 and the overall project evaluation, a total of 68 evaluation indicators in the "special bonds + PPP" whole-process tracking audit evaluation index system in Table 1, are used as the content of the questionnaire. Four options of excellent, good, medium, and poor are set for each evaluation indicator. Questionnaires are distributed to the registered supervision engineers of relevant engineering project management companies, and 39 valid questionnaires are retrieved. Analyzing the survey results of the overall project evaluation, 92.31% of the supervision engineers think this project is excellent or good, indicating that the project is operating well. The case data is used for the training of the audit prediction model. Since the 68 index data in the collected questionnaires are all descriptive data, they are subjected to feature quantification analysis and processing in descending order of the evaluation level. For the C1 - C68 indicators, 4 represents excellent, 3 represents good, 2 represents medium, and 1 represents poor; for the overall project evaluation OE, it is represented by four levels of Excellent, Good, Medium, and Poor. The sample data is shown in Table 2: Table 2
[0027] The valid sample data numbered 1 - 39 processed by the above preprocessing method is denoted as , , n = 1,..., 39; the C1 - C68 index data in each sample is denoted as Xn, Xn=(x1,..., x68), which is used as the input data; the overall project evaluation data is denoted as Yn, Yn=(y1,..., y39), which is used as the output data; based on the random forest learning theory and method, an RF model is established for training and prediction. The process of using the valid sample data is as follows: randomly classify the data. Randomly select 20 sample data from 39 data samples to construct a model training set (TRS) for training the model, and use the other 19 sample data to construct a model test set (TES). The test set is data independent of the training and does not participate in the training at all, but is only used for the final evaluation of the model effect. Adopt K-Fold cross-validation. In order to better utilize the existing limited sample data to increase the data samples and ensure the stability of the generalization error, and thus obtain an ideal model, K-fold cross-validation needs to be used. The training set is divided into K folds, where K - 1 folds are used as K - 1 training sets , and a classification model is established by determining the fitting curve parameters, obtaining K different models, denoted as M, M= ; the Kth fold is used as the validation set (VAS) to evaluate the effect of the model M, and the model is optimized through hyperparameters; the above process is verified K times without repeated sampling, and the final required model is obtained; according to the sample quantity situation in this application, K = 5 is set for 5-fold cross-validation to train and optimize the model.
[0028] As Figure 2 shown, the specific steps of step S2 are as follows. Construct an audit evaluation model based on the random forest algorithm. The training and optimization of the audit evaluation model using the data set include: Step S21: Use the collected historical project audit evaluation index data Xn and the historical project overall evaluation Yn as the training sample S, S = (Xn, Yn), and use the collected audit evaluation index data of the current project as the evaluation sample Zm. Train the audit evaluation model based on the random forest algorithm with the training sample.
[0029] As Figure 3 shown, training the audit evaluation model based on the random forest algorithm with the training sample includes: Step S211: Randomly select data. Given the original training , , , as the number of training set samples, randomly draw times with replacement, draw 1 each time, and the probability of a sample being drawn is 1 / . Use this bootstrap sampling integration method to generate different training subsets, denoted as . Samples in different training subsets can be repeated, and samples in the same training subset can also be repeated. Use different training subsets to train decision tree models one by one, denoted as . The models are independent of each other. The samples not drawn form the validation set to check the generalization ability; Step S212: Randomly select features. Each sample has F feature variables. Specify f < F such that at each node, randomly select f feature variables from the F feature variables, find the best splitting node to split the decision tree, and do not prune the whole process. Then, according to the f features, calculate the best splitting method to improve the classification performance; Step S213: Vote to integrate the model. Use the regression analysis method to obtain their respective prediction results, and integrate the optimal RF model through voting.
[0030] In this embodiment, the RF model determines to randomly select f = 35 variables from 68 variables and use their best splitting to split the nodes.
[0031] Step S22: Establish the relationship between features through sample category and non - linear fitting, and use the decision tree to obtain the evaluation result.
[0032] Step S3: Input the obtained standardized evaluation index data into the optimized audit evaluation model for project audit evaluation to obtain the project audit evaluation result.
[0033] This application first establishes a comprehensive evaluation index system that conforms to the whole-process tracking audit of "special bonds + PPP"; secondly, applies text data recognition, random forest machine learning theory and methods to the whole-process tracking audit of "special bonds + PPP"; finally, selects real case evaluation sample data, trains and predicts the established RF model, and uses the RF model with the highest accuracy as the audit evaluation prediction model. This greatly improves the work efficiency of the "special bonds + PPP" audit evaluation prediction and has certain reference significance for promoting the full coverage of audit work.
[0034] This embodiment also discloses a project audit evaluation device based on machine learning, which is implemented by using the above-mentioned project audit evaluation method based on machine learning, and includes: a data preprocessing module, a model construction module, and an audit evaluation module; The data preprocessing module is used to obtain the audit evaluation index information corresponding to the project according to the comprehensive evaluation index system for the whole-process tracking audit of the project, and preprocess the audit evaluation index information to obtain standardized evaluation index data; The model construction module is used to construct an audit evaluation model based on the random forest algorithm, use all the indexes in the comprehensive evaluation index system for the whole-process tracking audit of the project as the feature factors of the random forest algorithm, use the audit evaluation result as the final model prediction result, and use the data set to train and optimize the audit evaluation model to obtain an optimized audit evaluation model; The audit evaluation module is used to input the obtained standardized evaluation index data into the optimized audit evaluation model for project audit evaluation to obtain the project audit evaluation result.
[0035] This invention also discloses an electronic device, including: a memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the above-mentioned project evaluation method.
[0036] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and modifications can be made to the embodiments of this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A project audit and evaluation method based on machine learning, characterized in that, It includes the following steps: According to the comprehensive evaluation index system for the whole-process follow-up audit of the project, obtain the audit evaluation index information corresponding to the project, and preprocess the audit evaluation index information to obtain standardized evaluation index data; Construct an audit evaluation model based on the random forest algorithm. Take all the indicators in the comprehensive evaluation index system for the whole-process follow-up audit of the project as the characteristic factors of the random forest algorithm, and take the audit evaluation result as the final model prediction result. Use the data set to train and optimize the audit evaluation model to obtain an optimized audit evaluation model; Input the obtained standardized evaluation index data into the optimized audit evaluation model for project audit evaluation to obtain the project audit evaluation result.
2. The method for project audit and evaluation based on machine learning according to claim 1, wherein The obtaining of the audit evaluation index information corresponding to the project and the preprocessing of the audit evaluation index information include: Obtain the audit evaluation index information corresponding to the project through a questionnaire survey, and clean, segment, remove stop words, perform synonym replacement and lemmatization processing on the descriptive text data in the audit evaluation index information to obtain a word segmentation list; Construct a domain dictionary by defining evaluation dimensions, establishing a sentiment word library and setting modifier coefficients. Traverse each word in the processed word segmentation list and compare it with the domain dictionary. If the current word exists in the keyword list of any domain category, create or update the record of this dimension in the feature dictionary and initialize the score of this dimension to 0; After matching the domain category, continue to check whether the current word belongs to the sentiment word library. If it belongs to a positive or negative sentiment word, obtain the basic weight, read the basic score of this word from the predefined dictionary, and automatically convert the weight of the negative word to a negative value. Store the current word and its basic weight in the details list, and then check whether there is a word at the previous position of the current sentiment word. If the previous word belongs to the modifier library, read the modifier coefficient, and according to the formula: adjusted weight = basic weight × modifier coefficient, perform multiplicative correction on the basic weight; Accumulate the calculated weight values into the total score of the corresponding dimension, and perform feature quantization processing on the total scores of each dimension from high to low according to the preset matching range to obtain an effective sample data set Zm, Zm=(z1,…,zm).
3. The project audit and evaluation method based on machine learning according to claim 1, wherein The construction of the audit evaluation model based on the random forest algorithm and the training and optimization of the audit evaluation model using the data set include: Take the collected historical project audit evaluation index data Xn and the historical project overall evaluation Yn as the training sample S, S=(Xn,Yn), and take the collected audit evaluation index data of the current project as the evaluation sample Zm. Train the audit evaluation model based on the random forest algorithm through the training sample; Establish the relationship between features through sample category and non-linear fitting, and use the decision tree to obtain the evaluation result.
4. The method for project audit and evaluation based on machine learning according to claim 3, wherein The training of the audit evaluation model based on the random forest algorithm through the training sample includes: Randomly select data and given the original training set , , , is the number of samples in the training set. Randomly draw with replacement times, draw 1 each time, and the probability of a sample being drawn is 1 / . Use this bootstrap sampling integration method to generate different training subsets, denoted as . Samples in different training subsets can be repeated, and samples in the same training subset can also be repeated. Use different training subsets to train decision tree models one by one, denoted as . The models are independent of each other. The samples not drawn form the validation set to detect 's generalization ability; Randomly select features. Each sample has F feature variables. Specify f < F such that at each node, randomly select f feature variables from the F feature variables, find the best splitting node to split the decision tree, and do not prune the whole process. Then, calculate the best splitting method based on the f features, thereby improving classification performance; The voting integration model uses the regression analysis method to obtain their respective prediction results, and integrates the optimal RF model through voting.
5. A project audit and evaluation device based on machine learning, which adopts the project audit and evaluation method based on machine learning according to any one of claims 1-4, characterized in that, It includes: A data preprocessing module, a model construction module and an audit evaluation module; The data preprocessing module is used to obtain the audit evaluation index information corresponding to the project according to the comprehensive evaluation index system for the whole process tracking audit of the project, and preprocess the audit evaluation index information to obtain the standardized evaluation index data; The model construction module is used to construct an audit evaluation model based on the random forest algorithm. All the indexes in the comprehensive evaluation index system for the whole process tracking audit of the project are used as the feature factors of the random forest algorithm, and the audit evaluation result is used as the final model prediction result. The dataset is used to train and optimize the audit evaluation model to obtain the optimized audit evaluation model; The audit evaluation module is used to input the obtained standardized evaluation index data into the optimized audit evaluation model for project audit evaluation to obtain the project audit evaluation result.