Substation maintenance efficiency optimization method based on big data analysis

By using big data analysis and the C4.5 algorithm to establish a decision tree, the problem of insufficient comprehensive evaluation of factors in substation maintenance was solved, maintenance efficiency was improved, and the safe and stable operation of the power grid was ensured.

CN114398757BActive Publication Date: 2025-10-21WENZHOU ELECTRIC POWER BUREAU +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111505058.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-10-21
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

Existing technologies are unable to conduct comprehensive evaluation and correlation analysis of various factors affecting substation maintenance, making it difficult to improve maintenance efficiency.

Method used

Through big data analysis, relevant information is collected and a decision tree is established using the C4.5 algorithm to conduct data evaluation and optimization to improve maintenance efficiency.

Benefits of technology

It has achieved targeted improvements in maintenance personnel capabilities, ticketing and warehousing systems, strengthened key maintenance elements, improved equipment reliability, and ensured safe and stable operation of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114398757B_ABST
    Figure CN114398757B_ABST
Patent Text Reader

Abstract

The application discloses a substation maintenance efficiency optimization method based on big data analysis, comprising the following steps: data acquisition and data preprocessing; using an evaluation table to convert and evaluate the processed data; using a C4.5 algorithm to analyze the evaluated multi-source data to obtain a decision tree; and optimizing the maintenance efficiency according to the rules obtained from the decision tree. According to the obtained rules, the maintenance personnel's ability can be improved, the ticket and storage system can be developed, and the key elements of maintenance can be strengthened, so that the quality and process specification of the substation maintenance work are strengthened. The application provides an important reference basis and auxiliary decision support for daily maintenance business, improves the equipment reliability, and guarantees the safe and stable operation of the power grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of substation maintenance, and in particular to a method for optimizing substation maintenance efficiency based on big data analysis. Background Art

[0002] With the continuous expansion of power grids and the increasing number of substations, the number of equipment failures and per capita equipment maintenance efforts have also increased significantly. As asset-intensive enterprises, one of the most important tasks of power grid companies is to ensure the safe and stable operation of their equipment assets. To ensure this crucial task, high-level equipment maintenance and efficient defect handling, especially critical and urgent defects, are essential. Maximizing maintenance efficiency while improving the safety and reliability of maintenance work with limited personnel has become a pressing challenge for power companies. Key safeguards include highly skilled maintenance personnel, an adequate supply of maintenance equipment and spare parts, accurate defect assessment, and pre-assessment of maintenance equipment. To maximize and effectively improve fault elimination efficiency within limited human and material resources, a comprehensive assessment and correlation analysis of the various factors influencing fault elimination is essential. Summary of the Invention

[0003] In view of the problem that the existing technology is unable to conduct comprehensive evaluation and correlation analysis of various factors affecting fault elimination, resulting in difficulty in improving fault elimination efficiency, the present invention provides a substation maintenance efficiency optimization method based on big data analysis. By collecting relevant information and using algorithm models to evaluate the data, and using C4.5 to mine multi-source data and establish a decision tree, relevant rules are obtained and targeted improvement and optimization are carried out to improve fault elimination efficiency.

[0004] The following are the technical solutions of the present invention.

[0005] The substation maintenance efficiency optimization method based on big data analysis includes the following steps:

[0006] Perform data collection and data preprocessing;

[0007] Use the evaluation form to convert and evaluate the processed data;

[0008] The C4.5 algorithm is used to analyze the evaluated multi-source data to obtain a decision tree;

[0009] Maintenance efficiency is optimized based on the rules obtained from the decision tree.

[0010] Based on the patterns obtained, this invention can be used to specifically improve the capabilities of maintenance personnel, develop ticketing and storage systems, and strengthen key maintenance elements, enhancing the quality and process standards of substation maintenance work. This provides important reference and auxiliary decision support for daily maintenance operations, improves equipment reliability, and ensures the safe and stable operation of the power grid.

[0011] Preferably, the data collection includes: collecting data from the defect management platform, PMS system and work area maintenance management system, the collected data including work tickets, defect statistics table and maintenance personnel information, and extracting key data from the collected work tickets, defect statistics table and maintenance personnel information.

[0012] Preferably, the data preprocessing includes: discarding noise data and incomplete data samples appearing in the data, listing the cleaned data, and dividing the cleaned data into a training sample set and a test sample set, wherein the ratio of the number of samples in the training sample set to the number of samples in the test sample set is 2:1.

[0013] Preferably, the processed data is converted and evaluated, including defect status evaluation, main equipment status evaluation, substation status evaluation, personnel capability evaluation, two-ticket execution capability evaluation and maintenance material support status evaluation.

[0014] Preferably, the master device status evaluation includes:

[0015] Through the substation ledger management system and PMS system data, the main equipment status quantities are divided into five categories:

[0016] (1) Intrinsic state quantity K1: derived from the device selection, voltage level, manufacturer type, operating environment and commissioning time;

[0017] (2) Stable state quantity K2: the overall reliability of the same type of equipment and the same model, the family defect rate and impact level;

[0018] (3) Risk-type state quantity K3: obtained based on the equipment's most recent regular inspection cycle test results and historical failure rate;

[0019] (4) Improvement state quantity K4: whether the equipment can be improved or restored to a better performance level after maintenance, transformation, countermeasures, and upgrades;

[0020] The status evaluation of the main equipment is scored according to the above four types of status quantities. The final score PS is obtained by adding different weights using the hierarchical analysis method. The calculation formula is as follows:

[0021] PS=(K1λ1+K2λ2+K3λ3)×K4

[0022] Among them: λ1, λ2, λ3, λ4 are weighting factors.

[0023] Preferably, the substation status evaluation includes:

[0024] The substation inventory information management platform collects data from all substations within the jurisdiction. Core influencing factors, including substation type, commissioning date, regular inspection cycle, voltage level, and integrated system type, are extracted and assigned different weights to obtain the final substation status assessment score TS. The model is shown below:

[0025]

[0026] where a j is the weight factor of the parameter item, ak is the weight factor of the parameter item sub-item, determined by the hierarchical analysis method, n is the number of state parameter evaluation items, l is the number of evaluation sub-items, P k Score the status item.

[0027] Preferably, the C4.5 algorithm is used to analyze the evaluated multi-source data to obtain a decision tree, including: calculating the information entropy of the category attribute, then calculating the expected information entropy of the non-category attribute, and obtaining the information gain rate through information gain and segmentation information, and the attribute with the maximum information gain rate is used as the node of the decision tree to construct the decision tree; for the type attribute D, T is divided into sets T1, T2, ..., Tn according to its value, when all records in each set produce the same result, Info(D, T) is 0, and the gain Gain(D, T) takes the maximum value; so the gain ratio is used instead, that is:

[0028]

[0029] Since T is split based on the value of the type attribute D, SplitInfo(D, T) is the amount of information, and the GainRatio function is used to calculate and compare to construct the corresponding decision tree, and each node is the attribute with the largest gain ratio among the attributes.

[0030] The substantial benefits of this invention include: Based on the patterns obtained, targeted improvements can be made to maintenance personnel capabilities, ticketing and storage systems can be developed, key maintenance elements can be strengthened, and the quality and process standards of substation maintenance can be enhanced. This provides important reference and decision-making support for daily maintenance operations, improves equipment reliability, and ensures the safe and stable operation of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flow chart of an embodiment of the present invention;

[0032] Figure 2 Schematic diagram of a decision tree according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0034] It should be understood that in various embodiments of the present invention, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0035] It should be understood that in the present invention, "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.

[0036] It should be understood that, in the present invention, "B corresponding to A," "B corresponding to A," "A corresponds to B," or "B corresponds to A" means that B is associated with A and B can be determined based on A. Determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information. A and B match when the similarity between A and B is greater than or equal to a preset threshold.

[0037] The technical solution of the present invention is described in detail below with reference to specific embodiments. The embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0038] Example:

[0039] Substation maintenance efficiency optimization methods based on big data analysis include: Figure 1 The following steps are shown:

[0040] Perform data collection and data preprocessing;

[0041] Use the evaluation form to convert and evaluate the processed data;

[0042] The C4.5 algorithm is used to analyze the evaluated multi-source data to obtain a decision tree;

[0043] Maintenance efficiency is optimized based on the rules obtained from the decision tree.

[0044] This embodiment can use the patterns learned to specifically improve maintenance personnel capabilities, develop ticketing and storage systems, and strengthen key maintenance elements, enhancing the quality and process standards of substation maintenance work. This provides important reference and decision-making support for daily maintenance operations, improving equipment reliability and ensuring the safe and stable operation of the power grid.

[0045] Among them, data collection includes: collecting data from the defect management platform, PMS system and work area maintenance management system. The collected data includes work tickets, defect statistics tables and maintenance personnel information, and extracting key data from the collected work tickets, defect statistics tables and maintenance personnel information.

[0046] Data preprocessing includes discarding noise data and incomplete data samples, listing the cleaned data, and dividing the cleaned data into training sample sets and test sample sets, where the ratio of the number of samples in the training sample set to the number of samples in the test sample set is 2:1.

[0047] The processed data is converted and evaluated, including defect status evaluation, main equipment status evaluation, substation status evaluation, personnel capability evaluation, two-ticket execution capability evaluation, and maintenance material support status evaluation.

[0048] Defect status evaluation, including:

[0049] Using data from defect management platforms such as the State Grid PMS system and One-Stop-One-Database, different core status parameters of defects are assigned points to form a defect status evaluation table (see the table below), which is converted into a defect status evaluation score DS using a formula.

[0050]

[0051] Table 1 Defect status evaluation table

[0052] Main equipment status evaluation, including:

[0053] Through the substation ledger management system and PMS system data, the main equipment status quantities are divided into five categories:

[0054] (1) Intrinsic state quantity K1: derived from the device selection, voltage level, manufacturer type, operating environment and commissioning time;

[0055] (2) Stable state quantity K2: the overall reliability of the same type of equipment and the same model, the family defect rate and impact level;

[0056] (3) Risk-type state quantity K3: obtained based on the equipment's most recent regular inspection cycle test results and historical failure rate;

[0057] (4) Improvement state quantity K4: whether the equipment can be improved or restored to a better performance level after maintenance, transformation, countermeasures, and upgrades;

[0058] The status evaluation of the main equipment is scored according to the above four types of status quantities. The final score PS is obtained by adding different weights using the hierarchical analysis method. The calculation formula is as follows:

[0059] PS=(K1λ1+K2λ2+K3λ3)×K4

[0060] Among them: λ1, λ2, λ3, λ4 are weighting factors.

[0061] Substation condition assessment, including:

[0062] The substation inventory information management platform collects data from all substations within the jurisdiction. Core influencing factors, including substation type, commissioning date, regular inspection cycle, voltage level, and integrated system type, are extracted and assigned different weights to obtain the final substation status assessment score TS. The model is shown below:

[0063]

[0064] where a j is the weight factor of the parameter item, ak is the weight factor of the parameter item sub-item, determined by the hierarchical analysis method, n is the number of state parameter evaluation items, l is the number of evaluation sub-items, P k Score the status item.

[0065] Personnel competency assessment, including:

[0066] Based on AHP, a hierarchical function mapping relationship is constructed, and the personnel information archive data and defect elimination efficiency data are correlated. The hierarchical weight is calculated based on the concept of importance inner product, and then a comprehensive evaluation system for personnel maintenance capability and quality is established.

[0067] The comprehensive evaluation system for maintenance capability and quality of personnel is divided into three layers with maintenance and defect elimination as the goal orientation, namely: the top layer is the target layer (A), which is the comprehensive evaluation system for maintenance capability and quality of personnel; the middle layer is the capability layer (B), including personal resume, skill strength, safety capability, innovation capability, etc.; the bottom layer is the indicator layer (C), including length of service, education level, job experience, professional title, star rating, etc.

[0068] The function mapping relationship between the lower layer (X) and the upper layer (Y) is:

[0069] Y=F(X1,X2,......,X n )

[0070] The weight formula is:

[0071] Ω=[ω1ω2......ω n ] T

[0072]

[0073]

[0074] The comprehensive evaluation system model of maintenance capability and quality of personnel is shown in Table 2:

[0075]

[0076]

[0077] Table 2 Comprehensive evaluation system of maintenance capability and quality of personnel

[0078] Based on the comprehensive evaluation system of personnel maintenance ability and quality, a comprehensive evaluation is conducted on the employees who are currently working on the front line of maintenance in the work area for a long time, and the score PA is converted into the personnel ability and quality evaluation score through a formula.

[0079] Two-ticket execution capability evaluation, including:

[0080] The impact of the two-ticket execution on the defect elimination work is mainly reflected in two aspects. On the one hand, as an indispensable part of the maintenance and defect elimination work, the speed of its execution has a direct impact on the efficiency of defect elimination. On the other hand, the correctness of the two-ticket execution and review process will indirectly reflect the control of the safety and reliability of the maintenance and defect elimination work. The final evaluation score PC of the two-ticket execution efficiency is obtained by assigning weights to the two aspects. The evaluation model is as follows:

[0081] PC=ρ i ×(P S ×0.55+P Z ×0.45)

[0082] Among them, P S is the execution speed evaluation score of the two tickets, P Z Score the correctness evaluation for both votes, ρ i For the evaluation coefficient, different star-level managers need to be assigned different values ​​due to the different complexity of the two votes they can execute. Please refer to the following evaluation coefficient division table 3.

[0083] Star rating of person in charge Samsung Two-star One star coefficient 1.25 1.2 1

[0084] Table 3 Evaluation coefficient division table

[0085] In order to obtain two execution speed evaluation scores P SThe relationship between the various links in the two-ticket execution process is described by a bi-directional fuzzy graph, and the time consumption of each link, including the ticket maker's production, the issuer's review, the operation and maintenance review and acceptance, the improvement and modification, and the closing of the loop, is quantitatively described using a generalized fuzzy matrix. The total time consumption of the evaluation process is the cumulative value of each time in the generalized fuzzy matrix, which can be referred to as follows:

[0086]

[0087] T={t ij}

[0088] When i=j, T represents the set of nine time elements of each link in the evaluation process, that is, the time consumed within each link; when i≠j, T represents the time consumed from link i to link j in the evaluation process.

[0089] Two-ticket execution correctness evaluation score P Z The data sources considered mainly include the error rates of different types of work tickets, the error rates of station meetings, and the implementation of manufacturer education cards. The scoring rules for these data sources are shown in the following two ticket execution correctness evaluation tables:

[0090]

[0091] Table 4 Two-ticket execution correctness evaluation table

[0092] Evaluation of maintenance material support status, including:

[0093] Through the work area maintenance management system, equipment ledger management system, and the problems encountered in the execution of recent maintenance and defect elimination work, we sorted out and counted them. Combined with the data of the State Grid PMS system and one-stop one-database defect management platforms, we assigned points and weighted factors such as tool selection, tool quality, spare parts selection, spare parts quality, and special equipment mastery to form an evaluation table for maintenance material support. See Table 5 below, and convert it into the defect status evaluation score MS through the formula.

[0094]

[0095]

[0096] Table 5 Evaluation table of maintenance material support (equipment support)

[0097] Data attributes are selected by building new attributes based on existing attribute sets to discover deeper knowledge. In the decision tree construction process, a selection of sample sets was used to design a model based on the maintenance personnel's troubleshooting techniques, gender, and process standardization. The PA (Personal Ability) attribute represents personnel competency assessment, PC (Paper Executive Capability) represents two-ticket execution capability assessment, DS (Defect Status) represents defect status assessment, PS (Primary Device Status) represents primary device status assessment, TS (Transformer Substation Status) represents substation status assessment, and MS (Material Support) represents maintenance material support status assessment. After data preprocessing, there are 20 records, as shown in the table below.

[0098] Serial number TS PA DS PS PC MS Defect elimination process Process specifications 1 C A C C A A N N 2 C B A C B B N N 3 C A B B A A Y N 4 C C A C B B N N 5 C B C B C C N N 6 C C A C C C Y N 7 C A B B B B Y Y ... ... ... ... ... ... ... ... ... 20 B C A C B A N Y

[0099] Table 6 Sample data statistics

[0100] The C4.5 algorithm is used to analyze the evaluated multi-source data to obtain a decision tree, including:

[0101] Calculate the information entropy of the categorical attribute, then calculate the expected information entropy of the non-categorical attribute, and obtain the information gain rate through information gain and segmentation information. The attribute with the maximum information gain rate is used as the node of the decision tree to construct the decision tree; for the type attribute D, divide T into sets T1, T2, ..., Tn according to its value. When all records in each set produce the same result, Info(D, T) is 0, and the gain Gain(D, T) takes the maximum value at this time; so the gain ratio is used instead, that is:

[0102]

[0103] Since T is split based on the value of the type attribute D, SplitInfo(D, T) is the amount of information, and the GainRatio function is used to calculate and compare to construct the corresponding decision tree, and each node is the attribute with the largest gain ratio among the attributes.

[0104] In this example, statistics were collected from 60 sets of maintenance and troubleshooting processes, with 20 representative data sets listed. The C4.5 algorithm can analyze both discrete and continuous values. To more accurately reflect the impact of each parameter on the forming results, only discrete value analysis is performed. Condition attributes are represented by A, B, and C to represent various condition levels, and decision attributes are represented by descriptive words.

[0105] Analyze the influence of various parameters on the defect elimination process and process specifications, and use the C4.5 algorithm to analyze the condition attributes and decision attributes to obtain the decision tree as follows Figure 2 shown.

[0106] Then, by forming and reviewing the rules of the decision tree, we can get the following rules:

[0107] Rule i:PA=X PC=Y TS=Z MS=W then class Y / NY / N

[0108] 40 test set data were collected, the test data were applied to the classification rules, and the classification results of the test were compared and analyzed with the actual conclusions. The results showed that 33 of them matched and 7 did not match, with an accuracy rate of 82.5%. The accuracy of the classification prediction met the expected target.

[0109] It can be clearly seen from the pruned decision tree and the generated rules:

[0110] (1) The ability of maintenance personnel has the greatest contribution to maintenance quality, plays a leading role, and directly affects the changes in maintenance technology and process specifications.

[0111] (2) Personnel capabilities and ticketing quality have a significant impact on the standardization of maintenance processes. In the case of insufficient personnel capabilities and weak ticketing support, maintenance quality abnormalities are more likely to occur. Moreover, when personnel capabilities are weak, ticketing quality directly affects the standardization of maintenance processes.

[0112] (3) Defect status assessment and equipment and material support have a great impact on maintenance efficiency. When these two are strong, they can make up for maintenance process problems caused by weak personnel capabilities.

[0113] Therefore, we can determine that the factors influencing maintenance energy efficiency are, in order: personnel capabilities, ticketing quality, defect assessment, and equipment and material support. This allows for targeted and comprehensive improvements, achieving multi-pronged optimization of maintenance and defect elimination efficiency. By comprehensively identifying the four key links in the production process—planning, dispatching, implementation, and termination—we establish a production collaborative management and control platform, connecting it to the human resources system, intelligent ticketing system, intelligent warehousing system, and secondary operation and maintenance system. We customize grid management, standardize the entire production process, and ultimately achieve improved defect elimination efficiency.

[0114] This embodiment has achieved: the elimination rate of major and above defects has always remained above 95%, the correct operation rate of relay protection is 100%, the total number of defects has decreased by more than 20% compared with the previous year, and the number of tripping events has been significantly reduced compared with the same period.

[0115] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the specific device can be divided into different functional modules to complete all or part of the functions described above.

[0116] If the embodiment of the present application is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0117] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for optimizing substation maintenance efficiency based on big data analysis, characterized in that: The following steps are involved: Perform data collection and data preprocessing; Use the evaluation form to convert and evaluate the processed data; The C4.5 algorithm is used to analyze the evaluated multi-source data to obtain a decision tree; Optimize maintenance efficiency based on the rules obtained from the decision tree; The data collection includes: collecting data from the defect management platform, PMS system and work area maintenance management system, the collected data includes work tickets, defect statistics and maintenance personnel information, and extracting key data from the collected work tickets, defect statistics and maintenance personnel information; The processed data is converted and evaluated, including defect status evaluation, main equipment status evaluation, substation status evaluation, personnel capability evaluation, two-ticket execution capability evaluation, and maintenance material support status evaluation; The master device status evaluation includes: Through the substation ledger management system and PMS system data, the main equipment status quantities are divided into four categories: (1) Intrinsic state quantity : Derived from the device's model selection, voltage level, manufacturer type, operating environment, and commissioning time; (2) Stable state quantity : The overall reliability, family defect rate and impact level of similar equipment and equipment of the same model; (3) Risk-type state quantity : Obtained based on the equipment's most recent regular inspection cycle test results and historical failure rates; (4) Lifting state quantity : Whether the equipment can be improved or restored to a better performance level after maintenance, modification, countermeasures, and upgrades; The status evaluation of the main equipment is scored according to the above four types of status quantities. The final score PS is obtained by adding different weights using the hierarchical analysis method. The specific calculation formula is as follows: ; in: 、 、 is the weighting factor.

2. The method for optimizing substation maintenance efficiency based on big data analysis according to claim 1 is characterized in that: The data preprocessing includes: discarding noise data and incomplete data samples appearing in the data, listing the cleaned data, and dividing the cleaned data into a training sample set and a test sample set, wherein the ratio of the number of samples in the training sample set to the number of samples in the test sample set is 2:

1.

3. The method for optimizing substation maintenance efficiency based on big data analysis according to claim 1 is characterized in that: The substation status evaluation includes: The substation inventory information management platform collects data from all substations within the jurisdiction. Core influencing factors, including substation type, commissioning date, regular inspection cycle, voltage level, and integrated system type, are extracted and assigned different weights to obtain the final substation status assessment score TS. The model is shown below: ; in is the weight factor of the parameter item, ak is the weight factor of the parameter item sub-item, determined by the hierarchical analysis method, n is the number of state parameter evaluation items, l is the number of evaluation sub-items, Score the status item.

4. The method for optimizing substation maintenance efficiency based on big data analysis according to claim 1 is characterized in that: The C4.5 algorithm is used to analyze the evaluated multi-source data to obtain a decision tree, including: Calculate the information entropy of the categorical attribute, then calculate the expected information entropy of the non-categorical attribute, and obtain the information gain rate through information gain and segmentation information. The attribute with the maximum information gain rate is used as the node of the decision tree to construct the decision tree; for the type attribute D, divide T into sets T1, T2, ..., Tn according to its value. When all records in each set produce the same result, Info (D, T) is 0, and the gain Gain (D, T) takes the maximum value at this time; so the gain ratio is used instead, that is: ; Since T is split based on the value of the type attribute D, SplitInfo(D, T) is the amount of information. The GainRatio function is used to calculate and compare to construct the corresponding decision tree. Each node is the attribute with the maximum gain ratio among the attributes.

Citation Information

Patent Citations

  • Power secondary equipment risk assessment method and system thereof

    CN102324068A

  • A method for assessing a power communication network overhaul scheme

    CN103618638A

  • Power supply service quality analysis method based on Bayesian pruning decision tree

    CN110942098A