Converter transformer fault analysis method and related system

By combining machine learning models and power industry domain training with decision trees, the lack of accuracy and timeliness in converter transformer fault analysis was solved, efficient and accurate fault diagnosis was achieved, and the safe and stable operation of the power system was ensured.

CN120671081APending Publication Date: 2025-09-19CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510777118.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing converter transformer fault analysis methods have difficulty accurately identifying fault types and locations when faced with complex DC system faults, and are inefficient when processing large-scale data, affecting the safe and stable operation of the power system.

Method used

By combining machine learning models with power industry domain-specific training, normalized data processing and decision tree models are used to identify fault types in combination with rule-based data sets, achieving efficient and accurate fault analysis.

Benefits of technology

It significantly improves the accuracy of fault location and type identification, achieves fault diagnosis speed of minutes or even seconds, reduces the probability of missed diagnosis and misdiagnosis, and ensures the safe and stable operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671081A_ABST
    Figure CN120671081A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of power system fault analysis, and discloses a converter transformer fault analysis method and a related system.Unified preprocessing and standardization are carried out on multi-source data such as technical standards, empirical rules and case reports, data format and quality differences are eliminated, and the fault analysis efficiency is improved. The input data in the subsequent modeling and inference processes are highly consistent, and the error risk caused by data heterogeneity is greatly reduced, so that the basic accuracy of fault analysis is ensured. According to the method, a domain training process of a power industry model is adopted, so that the model fully learns electrical characteristics and professional experience in the industry, it is ensured that typical power system signal characteristics can be captured in the face of converter transformer faults, and the accuracy of fault positioning and type recognition is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power system fault analysis, and in particular relates to a converter transformer fault analysis method and a related system. Background Art

[0002] In power systems, converter transformers are key equipment in high-voltage direct current (HVDC) transmission projects. Their stability and reliability are crucial to the safe operation of the entire power system. Ultra-high voltage direct current (UHVDC) transmission, with its long-distance, high-capacity, and low-loss characteristics, plays a vital role in power systems. However, existing converter transformers still have many shortcomings and defects in DC fault analysis.

[0003] Traditional fault analysis methods for converter transformers rely heavily on empirical judgment and simple fault detection techniques. These methods often struggle to accurately identify fault types and locations when faced with complex DC system faults. Converter transformers are affected by a variety of factors during operation, including the combined effects of AC and DC electric and magnetic fields, equipment aging, and environmental changes. These factors can lead to various defects, such as insulation degradation, and existing fault diagnosis techniques have limitations in identifying these complex defects. Furthermore, with the continuous expansion of power systems and the integration of renewable energy sources, the operating environment of converter transformers has become increasingly complex, placing higher demands on the accuracy and efficiency of fault analysis. However, existing fault analysis techniques often suffer from high computational complexity and low efficiency when processing large amounts of data and complex systems. This makes it difficult to fully ensure the accuracy and timeliness of fault analysis in practical applications, impacting the safe and stable operation of power systems. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned shortcomings that the accuracy and timeliness of fault analysis are difficult to be fully guaranteed, thereby affecting the safe and stable operation of the power system, and to provide a converter transformer fault analysis method and related system.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for analyzing converter transformer faults, comprising the following steps: Determining normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set; Dividing the case data set into a training set and a test set, performing domain-specific training on the power industry model, using the test set to test the power industry model after domain-specific training, and comparing the test results with the results of the feature data set to ensure that the power industry model meets domain-specific requirements; Using the training set to train the machine learning model to obtain a trained machine learning model; Obtain the current case report, extract features from the current case report, feed the extracted features into the domain-trained power industry model, combine the power industry model with the rule dataset to obtain a first probability distribution, and feed the extracted features into the trained machine learning model to obtain a second probability distribution; The first probability distribution is combined with the second probability distribution to obtain the fault type.

[0006] A further improvement of the present invention is that determining normalized data includes: Acquire all data from technical standards, empirical rules, and case reports, clean the data, and obtain text information related to power converter failure analysis, which is used as normalized data. Summarize the failure cases in the normalized data to form a case data set; Summarize the empirical rules in the normalized data to form a rule data set; The feature descriptions in the normalized data that are relevant to the fault type are summarized to form a feature data set.

[0007] A further improvement of the present invention is to divide the case data set into a training set and a test set, perform domain-specific training on the power industry model, use the test set to test the optimized power industry model, and compare the test results with the feature data set. The specific method for making the power industry model meet the domain requirements is as follows: Divide the case data set into training set and test set; The power industry model is trained in the domain by using full parameter fine-tuning, LoRA fine-tuning, and prompt word engineering methods to ensure that the power industry model meets the required domain requirements. The power industry model after domain training is tested using the test set, and the test results of the power industry model are compared with the results in the feature dataset; If the comparison result meets the requirements, the current power industry model will be used as the power industry model after domain-based training; if the comparison result does not meet the requirements, the power industry model will be domain-based trained and tested again until the power industry model meets the domain-based requirements.

[0008] A further improvement of the present invention is that a training set is used to train the machine learning model. The specific method for obtaining the trained machine learning model is as follows: Build a decision tree model based on the machine learning model; Use the training set to train the decision tree model; Use information gain or Gini index as feature selection criteria to construct the branch structure of the decision tree; Combine the branching structure with the decision tree model to obtain a trained machine learning model.

[0009] A further improvement of the present invention is to obtain a current case report, extract features from the current case report, feed the extracted features into a domain-trained electric power industry model, combine the electric power industry model with a rule dataset to obtain a first probability distribution, and feed the extracted features into a trained machine learning model to obtain a second probability distribution. The specific method is as follows: Obtain the current case report and use the domain-trained power industry model to extract features from the current case report; Convert the extracted features into a new format to obtain an extracted feature dataset; The extracted feature dataset is fed into the domain-trained power industry model and combined with the rule dataset to obtain the first probability distribution; The extracted feature data set is fed into the trained machine learning model to obtain the second probability distribution.

[0010] A further improvement of the present invention is that after the fault type is obtained, the fault type is verified, and if the verification fails, the features in the current case report are re-extracted.

[0011] In a second aspect, the present invention provides a converter transformer fault analysis system, comprising: A data processing module, configured to determine normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set; A first model training module is used to divide the case data set into a training set and a test set, perform domain-specific training on the power industry model, use the test set to test the power industry model after domain-specific training, and compare the test results with the results of the feature data set to ensure that the power industry model meets the domain-specific requirements; A second model training module is used to train the machine learning model using the training set to obtain a trained machine learning model; A probability distribution acquisition module is used to obtain the current case report, extract features from the current case report, feed the extracted features into the domain-trained power industry model, combine the power industry model with the rule data set, obtain a first probability distribution, and feed the extracted features into the trained machine learning model to obtain a second probability distribution; The fault type inference module is used to combine the first probability distribution with the second probability distribution to obtain probability distribution information of the fault type, process the probability distribution information of the fault type, and compare it with the rule data set to obtain the fault type.

[0012] A further improvement of the present invention is that the functions of the data processing module are implemented by the following method: Acquire all data from technical standards, empirical rules, and case reports, clean the data, and obtain text information related to power converter failure analysis, which is used as normalized data. Summarize the failure cases in the normalized data to form a case data set; Summarize the empirical rules in the normalized data to form a rule data set; The feature descriptions in the normalized data that are relevant to the fault type are summarized to form a feature data set.

[0013] A further improvement of the present invention is that the function of the first model training module is implemented by the following method: Divide the case data set into training set and test set; The power industry model is trained in the domain by using full parameter fine-tuning, LoRA fine-tuning, and prompt word engineering methods to ensure that the power industry model meets the required domain requirements. The power industry model after domain training is tested using the test set, and the test results of the power industry model are compared with the results in the feature dataset; If the comparison result meets the requirements, the current power industry model will be used as the power industry model after domain-based training; if the comparison result does not meet the requirements, the power industry model will be domain-based trained and tested again until the power industry model meets the domain-based requirements.

[0014] A further improvement of the present invention is that the function of the second model training module is implemented by the following method: Build a decision tree model based on the machine learning model; Use the training set to train the decision tree model; Use information gain or Gini index as feature selection criteria to construct the branch structure of the decision tree; Combine the branching structure with the decision tree model to obtain a trained machine learning model.

[0015] A further improvement of the present invention is that the function of the probability distribution acquisition module is implemented by the following method: Obtain the current case report and use the domain-trained power industry model to extract features from the current case report; Convert the extracted features into a new format to obtain an extracted feature dataset; The extracted feature dataset is fed into the domain-trained power industry model to obtain the first probability distribution; The extracted feature data set is fed into the trained machine learning model to obtain the second probability distribution.

[0016] In a third aspect, the present invention provides a converter transformer fault analysis system, comprising: A data processing module, configured to determine normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set; A model training module is used to divide the case data set into a training set and a test set, perform domain-specific training on the power industry model, use the test set to test the power industry model after domain-specific training, and compare the test results with the results of the feature data set to ensure that the power industry model meets the domain-specific requirements; use the training set to train the machine learning model to obtain a trained machine learning model; The model inference module is used to obtain the current case report, extract features from the current case report, feed the extracted features into the power industry model after domain training, combine the power industry model with the rule data set to obtain a first probability distribution, feed the extracted features into the trained machine learning model to obtain a second probability distribution, and combine the first probability distribution with the second probability distribution to obtain the fault type.

[0017] In a fourth aspect, the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements a converter transformer fault analysis method when executing the computer program.

[0018] In a fifth aspect, the present invention provides a storage medium having a computer program stored thereon, wherein the computer program implements a converter transformer fault analysis method when executed by a processor.

[0019] Compared with the prior art, the present invention has the following beneficial effects: The present invention performs unified preprocessing and normalization on multi-source data such as technical standards, empirical rules and case reports, eliminating data format and quality differences, making the input data highly consistent in the subsequent modeling and inference process, greatly reducing the risk of errors introduced by data heterogeneity, and thus ensuring the basic accuracy of fault analysis. The present invention adopts the domain-based training process of the power industry model, so that the model can fully learn the electrical characteristics and professional experience in the industry, ensuring that it can capture typical power system signal characteristics when facing converter transformer faults, and significantly improving the accuracy of fault location and type identification. By training the machine learning model on the training set, the present invention can mine the complex nonlinear relationships and subtle patterns hidden in historical cases, complement the explicit knowledge of the domain-based model, and further reduce the probability of missed diagnosis and misdiagnosis. The present invention performs a weighted fusion of the first probability distribution output by the domain-based model and the second probability distribution output by the machine learning model, which not only comprehensively considers expert experience and data-driven insights, but also can quantify the confidence of different fault types through probability information, providing operation and maintenance personnel with an explainable and traceable decision basis. The present invention verifies the final probability distribution in conjunction with the rule data set, so that the analysis results not only conform to the model inference, but also meet industry standards and safety specifications, facilitate subsequent audits and accountability, and improve the compliance and transparency of system operations. The entire process of the present invention, from feature extraction of case reports to probabilistic reasoning, can be automatically executed, reducing human intervention to a minimum, achieving a fault diagnosis speed of minutes or even seconds, significantly improving the timeliness of fault response, and reducing the risk of secondary damage caused by delayed judgment. In summary, the present invention has multiple guarantees for the accuracy of fault analysis in the links of basic data processing, model construction, decision fusion and rule verification, and has greatly improved the response speed through full-process automation, thereby effectively overcoming the shortcomings of traditional fault analysis in accuracy and timeliness, and ensuring the safe, stable and efficient operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flow chart of Example 1; Figure 2 This is a system diagram of Example 2; Figure 3 This is a system diagram of Example 3; Figure 4 This is a system diagram of Example 9. DETAILED DESCRIPTION

[0021] In order to further understand the content of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention and are not intended to limit it.

[0022] Example 1: See also Figure 1A converter transformer fault analysis method comprises the following steps: S1, determining normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set.

[0023] S2. Divide the case data set into a training set and a test set, perform domain-specific training on the power industry model, use the test set to test the power industry model after domain-specific training, and compare the test results with the feature data set to ensure that the power industry model meets the domain-specific requirements.

[0024] S3: Use the training set to train the machine learning model to obtain a trained machine learning model.

[0025] S4, obtain the current case report, extract features from the current case report, send the extracted features into the power industry model after domain training, combine the power industry model with the rule data set to obtain the first probability distribution, and send the extracted features into the trained machine learning model to obtain the second probability distribution.

[0026] S5. Combine the first probability distribution with the second probability distribution to obtain a fault type.

[0027] Example 2: See also Figure 2 , a converter transformer fault analysis system, comprising: A data processing module, configured to determine normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set; A first model training module is used to divide the case data set into a training set and a test set, perform domain-specific training on the power industry model, use the test set to test the power industry model after domain-specific training, and compare the test results with the results of the feature data set to ensure that the power industry model meets the domain-specific requirements; A second model training module is used to train the machine learning model using the training set to obtain a trained machine learning model; A probability distribution acquisition module is used to obtain the current case report, extract features from the current case report, feed the extracted features into the domain-trained power industry model, combine the power industry model with the rule data set to obtain a first probability distribution, and feed the extracted features into the trained machine learning model to obtain a second probability distribution; The fault type inference module is used to combine the first probability distribution with the second probability distribution to obtain the fault type.

[0028] Example 3: See also Figure 3 , a converter transformer fault analysis system, comprising: A data processing module is used to determine normalized data; wherein, the normalized data includes a case data set, a rule data set and a feature data set; the case data set can form a case library, the rule data set can form a rule library, and the feature data set can form a feature library.

[0029] The model training module is used to divide the case data set into a training set and a test set, perform domain-based training on the power industry model, use the test set to test the power industry model after domain-based training, compare the test results with the feature data set, make the power industry model meet the domain-based requirements, and obtain a large model for the power industry; use the training set to train the machine learning model to obtain the trained machine learning model as a small model.

[0030] The model inference module is used to obtain the current case report, extract features from the current case report, feed the extracted features into the power industry model after domain training, combine the power industry model with the rule data set to obtain a first probability distribution, feed the extracted features into the trained machine learning model to obtain a second probability distribution, and combine the first probability distribution with the second probability distribution to obtain the fault type.

[0031] Example 4: This embodiment further defines the steps of S1 and the functions of the data processing module on the basis of the above embodiment, as follows: S11, obtain all data in technical standards, empirical rules and case reports, clean the data, obtain text information related to power converter failure analysis, and use the text information as normalized data.

[0032] S12, summarizing the failure cases in the normalized data to form a case data set.

[0033] S13, summarizing the empirical rules in the normalized data to form a rule data set.

[0034] S14, summarizing the feature descriptions in the normalized data that are relevant to the fault type to form a feature data set.

[0035] Specifically: Technical standards, empirical rules, case reports, and other materials are entered into the system in PDF or text format. These materials cover the design, operation, maintenance, and fault diagnosis of converter transformers and serve as a crucial basis for fault analysis. Optical character recognition (OCR) is performed on the input PDF files to convert the text into editable text data. For materials already in text format, subsequent processing can be performed directly. OCR technology can recognize text, tables, formulas, and other content within PDFs and convert them into structured text data, facilitating subsequent data processing and analysis.

[0036] Collected text data often contains irrelevant characters, garbled characters, formatting marks, and other noisy information. By writing a specific cleaning program, we remove special symbols, extra spaces, headers, footers, and other irrelevant content from the text, retaining only the core text information directly related to power converter fault analysis, such as case data sets, rule data sets, and feature data sets.

[0037] Example 5: This embodiment further defines the steps of S2, the functions of the first model training module, and some functions of the model training module on the basis of the above embodiment, as follows: S21, divide the case data set into a training set and a test set.

[0038] S22 uses full parameter fine-tuning, LoRA fine-tuning, and prompt word engineering methods to conduct domain-specific training on the power industry model, so that the power industry model meets the required domain requirements.

[0039] S23, use the test set to test the power industry model after domain training, and compare the test results of the power industry model with the results in the feature data set.

[0040] S24, if the comparison result meets the requirements, the current power industry model is used as the power industry model after domain-based training; if the comparison result does not meet the requirements, the power industry model is domain-based trained and tested again until the power industry model meets the domain-based requirements.

[0041] Specifically: Step 1: Divide the case dataset into training and test sets, which can effectively evaluate the performance and generalization ability of the model.

[0042] Step 2: For domain-specific training of the power industry model, fine-tune or optimize the power industry model using full parameter fine-tuning, LoRA fine-tuning, and prompt word engineering methods, and select a better training optimization strategy based on the performance on the test set.

[0043] Example 6: This embodiment further defines the steps of S3 and the functions of the second model training module and some functions of the model training module on the basis of the above embodiment, as follows: S31, building a decision tree model based on the machine learning model.

[0044] S32, using the training set to train the decision tree model.

[0045] S33, using information gain or Gini index as the feature selection criterion to construct the branch structure of the decision tree.

[0046] S34, combining the branch with the decision tree model to obtain a trained machine learning model.

[0047] Specifically: For the training of machine learning models, a decision tree model is adopted. The training set data is used to train the decision tree model, and information gain or Gini index is used as the feature selection criterion to construct the branch structure of the decision tree. For example, when the oil temperature change rate is in the "high" range and the winding insulation resistance value is lower than a certain threshold, it is judged that the risk of insulation fault is high.

[0048] When necessary, the cross-validation method can be used to repeat the domain training of the power industry model and the training steps of the machine learning model to conduct an in-depth evaluation of the model's generalization ability.

[0049] Example 7: This embodiment further defines the steps of S4, the functions of the probability distribution acquisition module, and some functions of the model inference module on the basis of the above embodiment, as follows: S41, obtain the current case report, and use the power industry model trained in the domain to extract features from the current case report.

[0050] S42, converting the format of the extracted features to obtain an extracted feature data set.

[0051] S43: Send the extracted feature data set into the electric power industry model after domain training and combine it with the rule data set to obtain a first probability distribution.

[0052] S44: Send the extracted feature data set to the trained machine learning model to obtain a second probability distribution.

[0053] Specifically: Step 1: For the input in the form of case reports, first use the power industry model to extract its features and convert it into the feature form in the feature data set.

[0054] Step 2: Input the fault cases based on feature description into the domain-based power industry model and decision tree model to obtain the probability distribution of fault types output by both.

[0055] Example 8: This embodiment further defines the steps of S5, the functions of the fault type inference module, and some functions of the model inference module on the basis of the above embodiment, as follows: In step one, the obtained supplementary information on the probability distribution of fault types, along with the fault case, is fed into the power industry model. Simultaneously, the model searches for relevant rules in the rule dataset and selects the relevant diagnostic rules. Combined with these, the final fault type is deduced step by step.

[0056] In step 2, the output fault type is then verified by the power industry model. If the verification fails, feature extraction is repeated until the upper limit of the number of attempts is reached.

[0057] Step three: Based on human natural language processing logic, auxiliary technologies such as template filling are used to ensure that the final output is more suitable for power failure analysis scenarios. Furthermore, a lightweight markup language is used to format text, enabling interactive streaming output and ensuring high readability. Furthermore, this invention supports user satisfaction evaluation of the output results and allows for interactive optimization through multiple rounds of dialogue.

[0058] Based on the relevant descriptions in the fault cases, the names of the fault types involved are extracted. The power industry model is trained based on a massive amount of power-related data, contains a large amount of knowledge related to converters, and has strong understanding and reasoning capabilities. The cases are input into the large model one by one, and the large model extracts the fault type information. The extracted fault information is placed in the large model memory library. Whenever the power industry model receives a new case report input, the fault type of the case report is extracted, and then the extracted fault type is compared with the fault types already in the memory library. If there is no similar fault type, the fault type is updated to the memory library, otherwise it is not updated.

[0059] Example 9: See also Figure 4 The present invention also provides an electronic device 100 for a converter transformer fault analysis method; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0060] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the converter transformer fault analysis method described in Example 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 can mainly include a program storage area and a data storage area. The program storage area can store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data (such as audio data) created based on the use of the electronic device 100. In addition, the memory 101 can include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device.

[0061] The at least one processor 102 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor, etc. The processor 102 is the control center of the electronic device 100 and connects various parts of the entire electronic device 100 using various interfaces and lines.

[0062] The memory 101 in the electronic device 100 stores a plurality of instructions to implement a method for analyzing converter transformer faults. The processor 102 can execute the plurality of instructions to implement: Determining normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set; Dividing the case data set into a training set and a test set, performing domain-specific training on the power industry model, using the test set to test the power industry model after domain-specific training, and comparing the test results with the results of the feature data set to ensure that the power industry model meets domain-specific requirements; Using the training set to train the machine learning model to obtain a trained machine learning model; Obtain the current case report, extract features from the current case report, feed the extracted features into the domain-trained power industry model, combine the power industry model with the rule dataset to obtain a first probability distribution, and feed the extracted features into the trained machine learning model to obtain a second probability distribution; The first probability distribution is combined with the second probability distribution to obtain the fault type.

[0063] Example 10: If the module / unit integrated in the electronic device 100 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory and read-only memory (ROM, Read-Only Memory).

[0064] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0065] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0066] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for analyzing converter transformer faults, characterized in that: The following steps are involved: Determining normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set; Dividing the case data set into a training set and a test set, performing domain-specific training on the power industry model, using the test set to test the power industry model after domain-specific training, and comparing the test results with the results of the feature data set to ensure that the power industry model meets domain-specific requirements; Using the training set to train the machine learning model to obtain a trained machine learning model; Obtain the current case report, extract features from the current case report, feed the extracted features into the domain-trained power industry model, combine the power industry model with the rule dataset to obtain a first probability distribution, and feed the extracted features into the trained machine learning model to obtain a second probability distribution; The first probability distribution is combined with the second probability distribution to obtain the fault type.

2. A converter transformer fault analysis method according to claim 1, characterized in that: Determine normalized data, including: Acquire all data from technical standards, empirical rules, and case reports, clean the data, and obtain text information related to power converter failure analysis, which is used as normalized data. Summarize the failure cases in the normalized data to form a case data set; Summarize the empirical rules in the normalized data to form a rule data set; The feature descriptions in the normalized data that are relevant to the fault type are summarized to form a feature data set.

3. A converter transformer fault analysis method according to claim 1, characterized in that: The case data set is divided into a training set and a test set, and the power industry model is trained in a domain-specific manner. The optimized power industry model is tested using the test set, and the test results are compared with the feature data set to ensure that the power industry model meets the domain-specific requirements. The specific method is as follows: Divide the case data set into training set and test set; The power industry model is trained in the domain by using full parameter fine-tuning, LoRA fine-tuning, and prompt word engineering methods to ensure that the power industry model meets the required domain requirements. The power industry model after domain training is tested using the test set, and the test results of the power industry model are compared with the results in the feature dataset; If the comparison results meet the requirements, the current power industry model will be used as the power industry model after domain training; If the comparison results do not meet the requirements, the power industry model will be trained and tested again until the power industry model meets the domain requirements.

4. A converter transformer fault analysis method according to claim 1, characterized in that: The specific method of using the training set to train the machine learning model and obtain the trained machine learning model is as follows: Build a decision tree model based on the machine learning model; Use the training set to train the decision tree model; Use information gain or Gini index as feature selection criteria to construct the branch structure of the decision tree; Combine the branching structure with the decision tree model to obtain a trained machine learning model.

5. A converter transformer fault analysis method according to claim 1, characterized in that: The specific method for obtaining the current case report, extracting features from the current case report, feeding the extracted features into the domain-trained power industry model, combining the power industry model with the rule dataset to obtain the first probability distribution, and feeding the extracted features into the trained machine learning model to obtain the second probability distribution is as follows: Obtain the current case report and use the domain-trained power industry model to extract features from the current case report; Convert the extracted features into a new format to obtain an extracted feature dataset; The extracted feature dataset is fed into the domain-trained power industry model and combined with the rule dataset to obtain the first probability distribution; The extracted feature data set is fed into the trained machine learning model to obtain the second probability distribution.

6. A converter transformer fault analysis method according to claim 1, characterized in that: After obtaining the fault type, the fault type is verified. If the verification fails, the features in the current case report are re-extracted.

7. A converter transformer fault analysis system, characterized in that: include: A data processing module, configured to determine normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set; A first model training module is used to divide the case data set into a training set and a test set, perform domain-specific training on the power industry model, use the test set to test the power industry model after domain-specific training, and compare the test results with the results of the feature data set to ensure that the power industry model meets the domain-specific requirements; A second model training module is used to train the machine learning model using the training set to obtain a trained machine learning model; A probability distribution acquisition module is used to obtain the current case report, extract features from the current case report, feed the extracted features into the domain-trained power industry model, combine the power industry model with the rule data set to obtain a first probability distribution, and feed the extracted features into the trained machine learning model to obtain a second probability distribution; The fault type inference module is used to combine the first probability distribution with the second probability distribution to obtain the fault type.

8. A converter transformer fault analysis system according to claim 7, characterized in that: The functions of the data processing module are realized through the following methods: Acquire all data from technical standards, empirical rules, and case reports, clean the data, and obtain text information related to power converter failure analysis, which is used as normalized data. Summarize the failure cases in the normalized data to form a case data set; Summarize the empirical rules in the normalized data to form a rule data set; The feature descriptions in the normalized data that are relevant to the fault type are summarized to form a feature data set.

9. The converter transformer fault analysis system according to claim 7, characterized in that: The function of the first model training module is realized by the following method: Divide the case data set into training set and test set; The power industry model is trained in the domain by using full parameter fine-tuning, LoRA fine-tuning, and prompt word engineering methods to ensure that the power industry model meets the required domain requirements. The power industry model after domain training is tested using the test set, and the test results of the power industry model are compared with the results in the feature dataset; If the comparison results meet the requirements, the current power industry model will be used as the power industry model after domain training; If the comparison results do not meet the requirements, the power industry model will be trained and tested again until the power industry model meets the domain requirements.

10. The converter transformer fault analysis system according to claim 7, characterized in that: The function of the second model training module is realized by the following method: Build a decision tree model based on the machine learning model; Use the training set to train the decision tree model; Use information gain or Gini index as feature selection criteria to construct the branch structure of the decision tree; Combine the branching structure with the decision tree model to obtain a trained machine learning model.

11. The converter transformer fault analysis system according to claim 7, characterized in that: The functions of the probability distribution acquisition module are implemented through the following methods: Obtain the current case report and use the domain-trained power industry model to extract features from the current case report; Convert the extracted features into a new format to obtain an extracted feature dataset; The extracted feature dataset is fed into the domain-trained power industry model to obtain the first probability distribution; The extracted feature data set is fed into the trained machine learning model to obtain the second probability distribution.

12. A converter transformer fault analysis system, characterized in that: include: A data processing module, configured to determine normalized data; wherein the normalized data includes a case data set, a rule data set, and a feature data set; A model training module is used to divide the case data set into a training set and a test set, perform domain-specific training on the power industry model, use the test set to test the power industry model after domain-specific training, and compare the test results with the results of the feature data set to ensure that the power industry model meets the domain-specific requirements; use the training set to train the machine learning model to obtain a trained machine learning model; The model inference module is used to obtain the current case report, extract features from the current case report, feed the extracted features into the power industry model after domain training, combine the power industry model with the rule data set to obtain a first probability distribution, feed the extracted features into the trained machine learning model to obtain a second probability distribution, combine the first probability distribution with the second probability distribution to obtain probability distribution information of the fault type, process the probability distribution information of the fault type, and compare it with the rule data set to obtain the fault type.

13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, a converter transformer fault analysis method according to any one of claims 1 to 6 is implemented.

14. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a converter transformer fault analysis method according to any one of claims 1 to 6 is implemented.