Training data generation method and system for power grid fault diagnosis model

Generating simulation training data through clustering algorithms and circuit simulation models solves the problem of insufficient training data of the power grid fault diagnosis model and improves the fault recognition effect.

CN120336845APending Publication Date: 2025-07-18STATE GRID ANHUI ELECTRIC POWER CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510367780.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-10
Filing Date
2025-03-26
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing power grid fault diagnosis model training data accounts for a relatively small proportion of fault data, which leads to the inability to effectively train the model and affects the fault identification effect.

Method used

The clustering algorithm is used to identify cluster centers with a relatively small proportion of fault data, and use the circuit simulation model to generate simulation training data for data supplementation, increasing the number and diversity of fault data, and using Simulink software to build a circuit simulation model.

Benefits of technology

The proportion and quality of fault data in training data is improved, the fault recognition ability of the power grid fault diagnosis model is enhanced, and the fault recognition effect of the model in practical applications is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336845A_ABST
    Figure CN120336845A_ABST
Patent Text Reader

Abstract

The invention discloses a training data generation method and system for a power grid fault diagnosis model, mainly relates to the technical field of training data generation, and is used for solving the problem that the power grid fault diagnosis model cannot be effectively trained through fault data because the fault data in the existing training data is relatively small. And the fault identification effect of the power grid fault diagnosis model is influenced. Comprising the steps that the clustering number of initial training data corresponding to each clustering center is obtained, and then the ratio of the clustering number corresponding to each clustering center to the total number is obtained; obtaining a plurality of simulation training data by using the initial training data set and the simulation model; inputting all the simulation training data into a clustering algorithm to obtain a clustering center of the simulation training data and a clustering distance between the simulation training data and the clustering center; and through the clustering center and the clustering distance of the simulation training data, obtaining qualified training data with a data supplement quantity from all the simulation training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of training data generation, and in particular, to a method and system for generating training data for a power grid fault diagnosis model. Background Art

[0002] With the continuous improvement of the informatization level of power grid enterprises, information systems have become the core components of power grid enterprises. However, due to the large scale of power grid enterprise information systems and the variety of equipment and systems involved, fault diagnosis has become increasingly complex and difficult. Therefore, how to obtain sufficient training data to train a power grid fault diagnosis model and then obtain a power grid fault diagnosis model with comprehensive prediction effects has become increasingly important.

[0003] At present, the solutions for obtaining a power grid fault diagnosis model with comprehensive prediction effects mainly include: integrating fault data from different equipment and systems, including historical data and real-time data. Cleaning and annotating the collected data to ensure data quality and accuracy. Adopting a parallel training method to improve training efficiency. Synchronously updating the current model parameters of the generator and discriminator based on the loss function and model parameters. Jointly training the diagnosis models corresponding to multiple fault types to improve the generalization ability of the model.

[0004] However, in the actual operation process, there are various types of power grid faults and complex fault causes, including improper equipment selection, influence of the surrounding environment and external objects, etc. This makes it difficult to collect and sort out fault data. Therefore, the proportion of fault data in the training data is relatively small, and thus it is impossible to effectively train a power grid fault diagnosis model through fault data, which in turn affects the fault recognition effect of the power grid fault diagnosis model. Summary of the Invention

[0005] In view of the above deficiencies of the prior art, the present application provides a method and system for generating training data for a power grid fault diagnosis model to solve the problem that the proportion of fault data in the existing training data is relatively small, and thus it is impossible to effectively train a power grid fault diagnosis model through fault data, which in turn affects the fault recognition effect of the power grid fault diagnosis model.

[0006] In a first aspect, the present application provides a method for generating training data for a power grid fault diagnosis model, the method comprising: Obtain the circuit simulation models corresponding to each fault type; obtain the initial training data, input the training data into the clustering algorithm, obtain the clustering numbers of the initial training data corresponding to each clustering center, and further obtain the ratio of the clustering number corresponding to each clustering center to the total number; obtain the set of initial training data corresponding to the clustering centers with ratios less than the preset ratio threshold; determine the simulation models corresponding to each set of initial training data, and determine the data supplement quantity corresponding to each set of initial training data; use the set of initial training data and the simulation models to obtain a number of simulation training data; input all the simulation training data into the clustering algorithm to obtain the clustering centers of the simulation training data and the clustering distances between the simulation training data and the clustering centers; through the clustering centers and clustering distances of the simulation training data, obtain the data supplement quantity of qualified training data from all the simulation training data.

[0007] The training data generation method provided by the embodiments of the present application processes the initial training data through the clustering algorithm, identifies the clustering centers with a relatively small proportion of fault data, and then supplements data for these clustering centers. It can specifically increase the quantity of fault data and improve the proportion of fault data in the overall training data, thereby more effectively training the power grid fault diagnosis model. Using the circuit simulation model to generate simulation training data can simulate various fault types and fault scenarios, increasing the diversity and complexity of the training data. Through the clustering algorithm and the preset ratio threshold, it can accurately identify the clustering centers that need data supplementation and determine the corresponding data supplement quantity. It can ensure that the supplemented data is targeted and of high quality, thus avoiding the interference of invalid data on model training. Since the proportion and quality of fault data in the training data are improved, the power grid fault diagnosis model can more fully learn and understand fault characteristics during the training process, and further improve its fault recognition effect in practical applications.

[0008] In an implementation manner of the present application, obtaining the simulation models corresponding to each fault type specifically includes: obtaining the circuit simulation models corresponding to each fault type uploaded through a preset interface, and then using the Simulink software model to complete the construction of the circuit simulation models.

[0009] In an implementation manner of the present application, determining the circuit simulation models corresponding to each set of initial training data, and determining the data supplement quantity corresponding to each set of initial training data specifically includes: Obtain the preset number of initial training data closest to the cluster center in the current initial training data set and the names of each simulation model, send the preset number of initial training data and the names of each simulation model to the preset user terminal, and obtain the simulation model name returned as the simulation model corresponding to the current initial training data set; obtain the total number of initial training data and the number of cluster centers, and obtain the qualified number of the initial training data set by dividing the total number by the number of cluster centers; further, obtain the data supplement number by subtracting the number corresponding to the current initial training data set from the qualified number.

[0010] In an implementation manner of the present application, the initial training data at least includes: circuit parameters, characteristics of input signals, actual test data; using the initial training data set and the circuit simulation model, obtain a number of simulation training data, specifically including: when the data supplement number is less than the preset first quantity threshold, input the circuit parameters and the characteristics of the input signals in the initial training data into the circuit simulation model to obtain simulation test data; when the data supplement number is greater than the preset second quantity threshold, input the circuit parameters and the characteristics of the input signals in the initial training data into the circuit simulation model to obtain simulation test data; at the same time, based on the preset adjustment amplitude, adjust the specific values corresponding to the circuit parameters and the characteristics of the input signals in the initial training data, and input the adjusted circuit parameters and the characteristics of the input signals into the circuit simulation model to obtain simulation test data.

[0011] In an implementation manner of the present application, through the cluster center and the cluster distance of the simulation training data, obtain the data supplement number of qualified training data from all the simulation training data, specifically including: screen out a number of simulation training data consistent with the current initial training data set from all the simulation training data; when the number of the consistent simulation training data is greater than the data supplement number, screen out the data supplement number of simulation training data with the smallest cluster distance from the consistent simulation training data as the qualified training data.

[0012] In a second aspect, the present application provides a training data generation system for a power grid fault diagnosis model, and the system includes: An acquisition module for acquiring circuit simulation models corresponding to various fault types; an obtaining module for obtaining initial training data, inputting the training data into a clustering algorithm to obtain the number of clusters of the initial training data corresponding to each cluster center, and further obtaining the ratio of the number of clusters corresponding to each cluster center to the total number; obtaining a set of initial training data corresponding to the cluster centers with ratios less than a preset ratio threshold; determining the simulation models corresponding to each set of initial training data and determining the data supplement quantity corresponding to each set of initial training data; using the set of initial training data and the simulation models to obtain a number of simulation training data; inputting all the simulation training data into the clustering algorithm to obtain the cluster centers of the simulation training data and the cluster distances between the simulation training data and the cluster centers; and obtaining a number of qualified training data equal to the data supplement quantity from all the simulation training data through the cluster centers and cluster distances of the simulation training data.

[0013] In an implementation manner of the present application, the acquisition module includes an acquisition unit for obtaining, through a preset interface, the circuit simulation models corresponding to various fault types uploaded, and then using the Simulink software model to complete the construction of the circuit simulation models.

[0014] In an implementation manner of the present application, the obtaining module includes a first obtaining unit for obtaining a preset number of initial training data closest to the cluster center in the current set of initial training data and the names of each simulation model, sending the preset number of initial training data and the names of each simulation model to a preset user terminal, and obtaining the simulation model name returned as the simulation model corresponding to the current set of initial training data; obtaining the total number of initial training data and the number of cluster centers, and obtaining the qualified number of the set of initial training data by dividing the total number by the number of cluster centers; and further obtaining the data supplement quantity by subtracting the number corresponding to the current set of initial training data from the qualified number.

[0015] In an implementation manner of the present application, the initial training data at least includes: circuit parameters, characteristics of input signals, and actual test data; the obtaining module includes a second obtaining unit for, when the data supplement quantity is less than a preset first quantity threshold, inputting the circuit parameters and characteristics of input signals in the initial training data into the circuit simulation model to obtain simulation test data; when the data supplement quantity is greater than a preset second quantity threshold, inputting the circuit parameters and characteristics of input signals in the initial training data into the circuit simulation model to obtain simulation test data; and at the same time, adjusting the specific values corresponding to the circuit parameters and characteristics of input signals in the initial training data based on a preset adjustment amplitude, and inputting the adjusted circuit parameters and characteristics of input signals into the circuit simulation model to obtain simulation test data.

[0016] In an implementation manner of the present application, the obtaining module includes a third obtaining unit, which is configured to screen out a number of simulation training data that are consistent with the current initial training data set from all the simulation training data; when the number of the consistent simulation training data is greater than the data supplement number, screen out the data supplement number of simulation training data with the smallest clustering distance from the consistent simulation training data as qualified training data.

[0017] Those skilled in the art can understand that the present application has at least the following beneficial effects: The present application discloses a method and system for generating training data of a power grid fault diagnosis model. By processing the initial training data through a clustering algorithm, the clustering centers with a relatively small proportion of fault data are identified, and then data supplementation is performed for these clustering centers. It can specifically increase the number of fault data and improve the proportion of fault data in the overall training data, thereby more effectively training the power grid fault diagnosis model. Using a circuit simulation model to generate simulation training data can simulate various fault types and fault scenarios, increasing the diversity and complexity of the training data. Through the clustering algorithm and a preset ratio threshold, the clustering centers that need data supplementation can be accurately identified, and the corresponding data supplement number can be determined. It can ensure that the supplemented data is targeted and of high quality, thereby avoiding the interference of invalid data on model training. Since the proportion and quality of fault data in the training data are improved, the power grid fault diagnosis model can more fully learn and understand fault characteristics during the training process, and thus improve its fault recognition effect in practical applications.

[0018] It solves the problem that the proportion of fault data in the existing training data is relatively small, and thus it is impossible to effectively train the power grid fault diagnosis model through the fault data, which in turn affects the fault recognition effect of the power grid fault diagnosis model. Description of the Drawings

[0019] The following describes some embodiments of the present disclosure with reference to the accompanying drawings, in which: Figure 1 is a flowchart of a method for generating training data of a power grid fault diagnosis model provided by an embodiment of the present application.

[0020] Figure 2 is a schematic internal structure diagram of a system for generating training data of a power grid fault diagnosis model provided by an embodiment of the present application. Detailed Embodiments

[0021] Those skilled in the art should understand that the embodiments described below are only the preferred embodiments of the present disclosure, and do not mean that the present disclosure can only be implemented through these preferred embodiments. These preferred embodiments are only used to explain the technical principles of the present disclosure, rather than to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts should still fall within the protection scope of the present disclosure.

[0022] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0023] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0024] The embodiment provides a method for generating training data of a power grid fault diagnosis model, as Figure 1 shown, the method provided by the embodiment of the present application mainly includes the following steps: Step 110: Obtain circuit simulation models corresponding to various fault types.

[0025] This step can be specifically: obtain the uploaded circuit simulation models corresponding to various fault types through a preset interface, and then use the Simulink software model to complete the construction of the circuit simulation models.

[0026] The specific process can be: Through a preset user-friendly preset interface, allow users or experts to upload circuit simulation model files corresponding to various fault types. These files may contain information such as the structure, parameters, and fault settings of the circuit. Verify the uploaded circuit simulation models to ensure their accuracy and integrity. If there are problems or the models do not meet the requirements, communicate with the uploader in a timely manner and make corrections. Sort out the verified circuit simulation models, classify and store them according to the fault types for subsequent use. Select Simulink as the circuit simulation tool and utilize its circuit simulation and analysis functions to complete the construction of the circuit simulation models.

[0027] Step 120: Obtain initial training data, input the training data into a clustering algorithm, obtain the clustering numbers of the initial training data corresponding to each clustering center, and then obtain the ratio of the clustering number corresponding to each clustering center to the total number.

[0028] It should be noted that the clustering algorithm can be the K-means algorithm. The value of K here is the preset number of fault categories + 1 (normal situation).

[0029] Among them, determining the circuit simulation model corresponding to each initial training data set and determining the data supplement quantity corresponding to each initial training data set can be specifically as follows: Obtain the preset number of initial training data closest to the clustering center in the current initial training data set and the names of each simulation model, send the preset number of initial training data and the names of each simulation model to the preset user terminal, and obtain the returned simulation model name as the simulation model corresponding to the current initial training data set; obtain the total number of initial training data and the number of clustering centers, and obtain the qualified quantity of the initial training data set (i.e., the expected number of data points) by dividing the total number by the number of clustering centers; furthermore, obtain the data supplement quantity by subtracting the quantity corresponding to the current initial training data set from the qualified quantity.

[0030] Those skilled in the art can understand that in this step, the preset number of initial training data closest to the clustering center in the current initial training data set is obtained. These data points can be considered as the representative data of this cluster. Send these representative data and the simulation model name to the preset user terminal (such as an expert system or an artificial review platform) to obtain the simulation model name confirmed by the user. The returned simulation model name is the simulation model corresponding to the current initial training data set.

[0031] Step 130: Obtain the initial training data sets corresponding to the clustering centers with ratios less than the preset ratio threshold; determine the simulation models corresponding to each initial training data set and determine the data supplement quantity corresponding to each initial training data set; use the initial training data sets and the simulation models to obtain a number of simulation training data.

[0032] It should be noted that the initial training data at least includes: the parameters of the circuit, the characteristics of the input signal, and the actual test data; In the step, using the initial training data set and the circuit simulation model to obtain a number of simulation training data can be specifically as follows: When the data supplement quantity is less than the preset first quantity threshold, input the parameters of the circuit and the characteristics of the input signal in the initial training data into the circuit simulation model to obtain simulation test data.

[0033] When the data supplement quantity is greater than a preset second quantity threshold, input the parameters of the circuit and the characteristics of the input signal in the initial training data into the circuit simulation model to obtain simulation test data; meanwhile, based on a preset adjustment amplitude, adjust the specific values corresponding to the parameters of the circuit and the characteristics of the input signal in the initial training data, and input the adjusted parameters of the circuit and the characteristics of the input signal into the circuit simulation model to obtain simulation test data.

[0034] It should be noted that the preset adjustment amplitude can be randomly increased or decreased by [1%, 5%] on the basis of the original value.

[0035] Those skilled in the art can understand that by judging whether the data supplement quantity is less than a preset first quantity threshold. This threshold is usually a relatively small value, used to determine whether to adopt a simple simulation strategy when the data gap is not large. If the data supplement quantity is greater than a preset second quantity threshold (this threshold is usually greater than the first quantity threshold, indicating a large data gap), then a more complex simulation strategy needs to be adopted to generate more diverse data.

[0036] Step 140: Input all the simulation training data into the clustering algorithm to obtain the clustering centers of the simulation training data and the clustering distances between the simulation training data and the clustering centers; through the clustering centers and clustering distances of the simulation training data, obtain the number of qualified training data equal to the data supplement quantity from all the simulation training data.

[0037] In some embodiments, obtaining the number of qualified training data equal to the data supplement quantity from all the simulation training data through the clustering centers and clustering distances of the simulation training data can specifically be: Screen out several simulation training data that are consistent with the current initial training data set from all the simulation training data; when the number of the consistent several simulation training data is greater than the data supplement quantity, screen out the number of simulation training data with the smallest clustering distance equal to the data supplement quantity from the consistent several simulation training data as the qualified training data.

[0038] In addition, this application Figure 2 is a training data generation system for a power grid fault diagnosis model provided by an embodiment of this application. As Figure 2 shown, the system provided by the embodiment of this application mainly includes: An acquisition module 210, configured to acquire circuit simulation models corresponding to each fault type.

[0039] The acquisition module 210 includes an acquisition unit, configured to obtain the circuit simulation models corresponding to each fault type uploaded through a preset interface, and then use the Simulink software model to complete the construction of the circuit simulation models.

[0040] An obtaining module 220 is configured to obtain initial training data, input the training data into a clustering algorithm, obtain the number of clusters of the initial training data corresponding to each cluster center, and further obtain the ratio of the number of clusters corresponding to each cluster center to the total number; obtain a set of initial training data corresponding to the cluster centers whose ratios are less than a preset ratio threshold; determine the simulation models corresponding to each set of initial training data, and determine the data supplement quantity corresponding to each set of initial training data; use the set of initial training data and the simulation models to obtain a number of simulated training data; input all the simulated training data into the clustering algorithm, obtain the cluster centers of the simulated training data and the cluster distances between the simulated training data and the cluster centers; and obtain, from all the simulated training data, the data supplement quantity of qualified training data through the cluster centers and the cluster distances of the simulated training data.

[0041] The obtaining module 220 includes a first obtaining unit configured to obtain a preset number of initial training data closest to the cluster center and the names of each simulation model in the current set of initial training data, send the preset number of initial training data and the names of each simulation model to a preset user terminal, and obtain the simulation model name returned as the simulation model corresponding to the current set of initial training data; obtain the total number of the initial training data and the number of cluster centers, and obtain the qualified number of the set of initial training data through total number / number of cluster centers; and further obtain the data supplement quantity through the qualified number - the quantity corresponding to the current set of initial training data.

[0042] It should be noted that the initial training data includes at least: the parameters of the circuit, the characteristics of the input signal, and the actual test data; the obtaining module 220 includes a second obtaining unit configured to, when the data supplement quantity is less than a preset first quantity threshold, input the parameters of the circuit and the characteristics of the input signal in the initial training data into a circuit simulation model to obtain simulated test data; when the data supplement quantity is greater than a preset second quantity threshold, input the parameters of the circuit and the characteristics of the input signal in the initial training data into a circuit simulation model to obtain simulated test data; and at the same time, based on a preset adjustment range, adjust the specific values corresponding to the parameters of the circuit and the characteristics of the input signal in the initial training data, and input the adjusted parameters of the circuit and the characteristics of the input signal into the circuit simulation model to obtain simulated test data.

[0043] The obtaining module 220 includes a third obtaining unit configured to screen out a number of simulated training data consistent with the current set of initial training data from all the simulated training data; and when the number of the consistent simulated training data is greater than the data supplement quantity, screen out the data supplement quantity of simulated training data with the smallest cluster distance from the consistent simulated training data as the qualified training data.

[0044] So far, the technical solutions of the present disclosure have been described in combination with multiple embodiments in the foregoing. However, it is easy for those skilled in the art to understand that the protection scope of the present disclosure is not limited to these specific embodiments. Without departing from the technical principle of the present disclosure, those skilled in the art can split and combine the technical solutions in the above-mentioned various embodiments, and can also make equivalent changes or replacements to the relevant technical features. Any changes, equivalent replacements, improvements, etc. made within the technical concept and / or technical principle of the present disclosure will fall within the protection scope of the present disclosure.

Claims

1. A method for generating training data of a power grid fault diagnosis model, characterized in that The method includes: Obtaining circuit simulation models corresponding to respective fault types; Obtaining initial training data, inputting the training data into a clustering algorithm, obtaining the number of clusters of the initial training data corresponding to respective cluster centers, and further obtaining the ratio of the number of clusters corresponding to respective cluster centers to the total number; Obtaining a set of initial training data corresponding to a cluster center with a ratio less than a preset ratio threshold; determining the simulation models corresponding to respective sets of initial training data, and determining the data supplement quantity corresponding to respective sets of initial training data; using the sets of initial training data and the simulation models to obtain a number of simulation training data; Inputting all the simulation training data into a clustering algorithm, obtaining the cluster centers of the simulation training data and the cluster distances between the simulation training data and the cluster centers; obtaining, from all the simulation training data, the quantity of qualified training data equal to the data supplement quantity through the cluster centers and cluster distances of the simulation training data.

2. The method for generating training data of the power grid fault diagnosis model according to claim 1, characterized in that Obtaining the simulation models corresponding to respective fault types specifically includes: Obtaining, through a preset interface, the circuit simulation models corresponding to respective fault types uploaded, and further using the Simulink software model to complete the construction of the circuit simulation models.

3. The method for generating training data of the power grid fault diagnosis model according to claim 1, wherein Determining the circuit simulation models corresponding to respective sets of initial training data, and determining the data supplement quantity corresponding to respective sets of initial training data specifically includes: Obtaining a preset number of initial training data closest to the cluster center in the current set of initial training data and the names of respective simulation models, sending the preset number of initial training data and the names of respective simulation models to a preset user terminal, and obtaining the name of the simulation model returned as the simulation model corresponding to the current set of initial training data; Obtaining the total number of the initial training data and the number of cluster centers, and obtaining the qualified quantity of the set of initial training data through total number / number of cluster centers; and further obtaining the data supplement quantity through qualified quantity - the quantity corresponding to the current set of initial training data.

4. The method for generating training data of the power grid fault diagnosis model according to claim 1, characterized in that The initial training data includes at least: parameters of the circuit, characteristics of the input signal, actual test data; Using the sets of initial training data and the circuit simulation models to obtain a number of simulation training data specifically includes: When the data supplement quantity is less than a preset first quantity threshold, Inputting the parameters of the circuit and the characteristics of the input signal in the initial training data into the circuit simulation model to obtain simulation test data; When the data supplement quantity is greater than a preset second quantity threshold, Inputting the parameters of the circuit and the characteristics of the input signal in the initial training data into the circuit simulation model to obtain simulation test data; Meanwhile, based on a preset adjustment range, adjusting the specific values corresponding to the parameters of the circuit and the characteristics of the input signal in the initial training data, and inputting the adjusted parameters of the circuit and the characteristics of the input signal into the circuit simulation model to obtain simulation test data.

5. The method for generating training data of the power grid fault diagnosis model according to claim 1, characterized in that Obtaining, from all the simulation training data, the quantity of qualified training data equal to the data supplement quantity through the cluster centers and cluster distances of the simulation training data specifically includes: Screening out a number of simulation training data consistent with the current set of initial training data from all the simulation training data; When the number of several consistent simulation training data is greater than the data supplement number, screen out the data supplement number of simulation training data with the smallest clustering distance from the several consistent simulation training data as qualified training data.

6. A training data generation system for a power grid fault diagnosis model, characterized in that, The system includes: An acquisition module, configured to acquire circuit simulation models corresponding to respective fault types; An obtaining module, configured to acquire initial training data, input the training data into a clustering algorithm to obtain the clustering numbers of the initial training data corresponding to respective clustering centers, and further obtain the ratios of the clustering numbers corresponding to respective clustering centers to the total number; acquire a set of initial training data corresponding to a clustering center whose ratio is less than a preset ratio threshold; determine the simulation models corresponding to respective sets of initial training data, and determine the data supplement numbers corresponding to respective sets of initial training data; use the sets of initial training data and the simulation models to obtain several simulation training data; input all the simulation training data into the clustering algorithm to obtain the clustering centers of the simulation training data and the clustering distances between the simulation training data and the clustering centers; through the clustering centers and the clustering distances of the simulation training data, obtain the data supplement number of qualified training data from all the simulation training data.

7. The training data generation system for the power grid fault diagnosis model according to claim 6, characterized in that The acquisition module includes an acquisition unit, configured to obtain, through a preset interface, circuit simulation models corresponding to respective fault types uploaded, and then use a Simulink software model to complete the construction of the circuit simulation models.

8. The training data generation system for the power grid fault diagnosis model according to claim 6, wherein The obtaining module includes a first obtaining unit, configured to obtain a preset number of initial training data closest to the clustering center in the current set of initial training data and the names of respective simulation models, send the preset number of initial training data and the names of respective simulation models to a preset user terminal, and obtain the simulation model name returned as the simulation model corresponding to the current set of initial training data; acquire the total number of the initial training data and the number of clustering centers, and obtain the qualified number of the set of initial training data through total number / number of clustering centers; and further obtain the data supplement number through qualified number - the number corresponding to the current set of initial training data.

9. The training data generation system for the power grid fault diagnosis model according to claim 6, wherein The initial training data includes at least: circuit parameters, characteristics of input signals, and actual test data; the obtaining module includes a second obtaining unit, configured to when the data supplement number is less than a preset first quantity threshold, input the circuit parameters and the characteristics of the input signals in the initial training data into the circuit simulation model to obtain simulation test data; when the data supplement number is greater than a preset second quantity threshold, input the circuit parameters and the characteristics of the input signals in the initial training data into the circuit simulation model to obtain simulation test data; meanwhile, based on a preset adjustment range, adjust the specific values corresponding to the circuit parameters and the characteristics of the input signals in the initial training data, and input the adjusted circuit parameters and the characteristics of the input signals into the circuit simulation model to obtain simulation test data.

10. The training data generation system for the power grid fault diagnosis model according to claim 6, characterized in that The obtaining module includes a third obtaining unit, Used to screen out a number of simulation training data that are consistent with the current initial training data set from all simulation training data; when the number of the consistent simulation training data is greater than the data supplement quantity, screen out the data supplement quantity of simulation training data with the smallest clustering distance from the consistent simulation training data as qualified training data.