Three-phase inverter open-circuit fault diagnosis method and device based on deep transfer learning
Through deep transfer learning and pseudo-label iterative optimization methods, the problems of insufficient samples and low accuracy under variable operating conditions in open circuit fault diagnosis of three-phase inverter are solved, and efficient fault diagnosis effect is achieved, improving the generalization ability and accuracy of the model.
Patent Information
- Application Number
- CN202510293697.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
In the open circuit fault diagnosis of three-phase inverters, traditional methods are difficult to effectively utilize limited sample data, especially in variable operating conditions, and the performance of deep learning models in the target domain is degraded, so deep features cannot be fully extracted.
Using a method based on deep transfer learning, a convolutional neural network is built, combining continuous wavelet transformation and pseudo-label iterative optimization, a pre-trained model is trained using source domain data, and transfer and fine-tune iteratively in the target domain to generate pseudo-labels for model optimization, enhancing the generalization ability of the model.
It improves the accuracy and robustness of open circuit fault diagnosis of three-phase inverters, can effectively utilize a small amount of labeled data in the target domain, reduce the risk of overfitting, and improves the degree of automation and accuracy of diagnosis.
Smart Images

Figure CN120234690A_ABST
Abstract
Description
Background Art
[0002] New energy, especially solar photovoltaic power generation, has occupied a core position in the future power system. Three-phase inverters play a crucial role in solar power generation systems by converting solar energy into grid-connected electricity. It consists of power semiconductor devices such as insulated gate bipolar transistors (IGBTs) and related control systems. However, harsh operating conditions, high temperatures, and high electrical loads can cause unpredictable failures in IGBTs, leading to catastrophic damage. Common fault modes of inverters mainly include open-circuit faults and short-circuit faults. Short-circuit faults can cause overcurrents, usually interrupted by fuses, resulting in open-circuit faults. Although open-circuit faults can operate for a short time, they can cause non-fault devices to overheat and overcurrent, ultimately leading to system paralysis and huge economic losses. Therefore, in the application of power inverters, implementing an accurate and fast IGBT open-circuit fault diagnosis method is crucial.
[0003] Traditional power inverter open-circuit fault diagnosis methods are mainly divided into model-based methods and data-driven methods. Model-based methods rely on accurate mathematical models and detect faults by comparing the differences between predictions and measurements. However, due to modeling complexity and parameter uncertainty, it is difficult to establish an accurate model. With the development of artificial intelligence (AI) technology, AI-based data-driven methods have received attention because they do not require accurate physical modeling. Although traditional machine learning methods (such as support vector machines, SVMs) have achieved some success in IGBT fault diagnosis, they have the following disadvantages: 1) Shallow network structures are difficult to capture complex non-linear features; 2) Feature extraction depends on expert experience; 3) The degree of automation is low, and features need to be manually selected. Compared with traditional machine learning methods, deep learning can automatically extract features through multiple non-linear layers and end-to-end network structures, thus improving the accuracy of fault diagnosis.
[0004] Traditional AI methods (including deep learning) rely on a large amount of labeled data and assume that the data in the source domain and the target domain follow the same distribution. However, this assumption often does not hold in practical applications, resulting in a decline in the performance of the model in the target domain. In practical applications, collecting a large amount of labeled data is costly and time-consuming, especially for occasional fault data. In addition, although the two data sets have the same topological structure, their distributions will also be different under different operating conditions, power levels, and operating environments. This makes it difficult for the model trained in the source domain to be effectively transferred to the target domain, resulting in low data utilization. Therefore, a method that can effectively extract features from limited samples and solve the problem of insufficient samples under variable working conditions, resulting in low diagnostic accuracy, is needed.
[0005] As a branch of machine learning, transfer learning can utilize existing knowledge to solve problems in related fields. Due to the reusability of deep learning models, combining them with transfer learning can achieve classification in new fields by leveraging the features of existing data. Transfer learning has been less applied in the open - circuit fault diagnosis of inverters. Existing methods mainly input the original one - dimensional time - series data into a convolutional neural network and adopt a simple feature extraction fine - tuning strategy, without considering the differences in feature distributions between domains, and the model is too simple to fully extract deep features. Summary of the Invention
[0006] The purpose of the present invention is to provide a three - phase inverter open - circuit fault diagnosis method based on deep transfer learning to solve the problems proposed in the above background technology.
[0007] The first - aspect embodiment of this application proposes a three - phase inverter open - circuit fault diagnosis method based on deep transfer learning, which is characterized by including the following steps:
[0008] Step 100, collect the labeled sample set 1 of the three - phase inverter in the normal state and N fault states in the source domain; collect the labeled sample set 2 of the normal state and N fault states in the target domain, and the unlabeled sample set 3 in the target domain;
[0009] Step 200, construct a convolutional neural network model for fault diagnosis, and use the sample set 1 in the source domain to train the model to obtain the pre - trained model α;
[0010] Step 300, construct a fault diagnosis model based on transfer learning on the basis of the pre - trained model α, and use the sample set 2 in the target domain to train the above - mentioned model to obtain the transfer model β;
[0011] Step 400, use the transfer model β to predict the unlabeled samples in the unlabeled sample set 3 to generate pseudo - labels, and adopt the consistency regularization technique to smooth the pseudo - label prediction results for the subsequent training of the transfer model β to obtain the final optimized model γ;
[0012] Step 500, input the data of the three - phase inverter to be detected into the model γ, and output the predicted health status.
[0013] Preferably, the method for obtaining samples in each sample set in step 100 is: collect the three - phase current waveform data of the three - phase inverter, and perform wavelet transform on it to generate a time - frequency diagram as a sample.
[0014] Preferably, the continuous wavelet transform is performed on the three - phase current waveform signal using the cmor3 - 3 wavelet basis function to generate a time - frequency diagram.
[0015] Preferably, the convolutional neural network model includes 5 convolutional layers, 3 pooling layers, 2 fully connected layers, and 1 Dropout layer. The Softmax is used as the activation function for the final output layer, and the LeakyReLU is used as the activation function for the remaining convolutional layers and fully connected layers.
[0016] Preferably, constructing a fault diagnosis model based on transfer learning on the basis of the pre-trained model α includes: migrating the network structure and parameters of the pre-trained model α to the sample set 2 in the target domain, freezing the first three convolutional layers of the network, fixing their parameters unchanged, and then fine-tuning the parameters of the deep network structure. Using the training set and test set in the target domain sample set 2 to train the model, and using the validation set in the target domain sample set 2 to verify the accuracy of the trained model, to obtain a transfer model β that meets the accuracy requirements.
[0017] Preferably, an early stopping algorithm is introduced during the subsequent training process of the transfer model β, and whether to abort the training early is determined according to the performance of the validation set to avoid overfitting.
[0018] The second aspect embodiment of this application proposes a three-phase inverter open-circuit fault diagnosis device based on deep transfer learning, including:
[0019] Module 100 is used to collect the labeled sample set 1 in the normal state and N fault states of the three-phase inverter in the source domain; collect the labeled sample set 2 in the normal state and N fault states in the target domain, and the unlabeled sample set 3 in the target domain;
[0020] Module 200 constructs a convolutional neural network model for fault diagnosis, and uses the sample set 1 in the source domain to train the model to obtain the pre-trained model α;
[0021] Module 300 constructs a fault diagnosis model based on transfer learning on the basis of the pre-trained model α, and uses the sample set 2 in the target domain to train the above model to obtain the transfer model β;
[0022] Module 400 uses the transfer model β to predict the unlabeled samples in the unlabeled sample set 3, generates pseudo-labels, and uses the consistency regularization technology to smooth the pseudo-label prediction results for the subsequent training of the transfer model β to obtain the final optimized model γ;
[0023] Module 500 is used to input the data of the three-phase inverter to be detected into the model γ and output the predicted health status.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] (1) Using continuous wavelet transform (CWT) to process the current signal and generate a time-frequency diagram to fully extract the local features of the original signal;
[0026] (2) Extract depth information from the time-frequency image using an optimized deep two-dimensional convolutional neural network, and introduce a locally frozen transfer learning method to enhance the generalization ability of the model to other working conditions;
[0027] (3) To prevent overfitting, this patent adds a Dropout layer to the convolutional neural network and introduces a pseudo-label iterative optimization method. Combining a small amount of labeled data and a large amount of unlabeled data, pseudo-labels are screened through a confidence threshold, and the model is iteratively optimized to solve the problem of insufficient samples in the target domain. Brief Description of the Drawings
[0028] Figure 1 It is a flowchart of a three-phase inverter open-circuit fault diagnosis method based on deep transfer learning of the present invention.
[0029] Figure 2 It is the network structure of the convolutional neural network model of the present invention. Detailed Embodiments
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] As Figure 1 shown, a three-phase inverter open-circuit fault diagnosis method based on deep transfer learning according to an embodiment of the present application includes the following steps:
[0032] Step 100, collect the labeled sample set 1 of the three-phase inverter in the normal state and N fault states in the source domain; collect the labeled sample set 2 of the normal state and N fault states in the target domain, and the unlabeled sample set 3 in the target domain;
[0033] Step 200, construct a convolutional neural network model for fault diagnosis, and use the sample set 1 in the source domain to train the model to obtain a pre-trained model α;
[0034] Step 300, construct a fault diagnosis model based on transfer learning on the basis of the pre-trained model α, and use the sample set 2 in the target domain to train the above model to obtain a transfer model β;
[0035] Step 400, use the transfer model β to predict the unlabeled samples in the unlabeled sample set 3, generate pseudo-labels, and use the consistency regularization technology to smooth the pseudo-label prediction results for the subsequent training of the transfer model β to obtain the final optimized model γ;
[0036] Step 500: Input the three-phase inverter data to be detected into model γ, and output the predicted health status.
[0037] According to an embodiment of the present application, the method for obtaining samples in the sample sets in step 100 is as follows: collect the three-phase current waveform data of the three-phase inverter, and perform wavelet transform on it to generate a time-frequency diagram as a sample. The cmor3-3 wavelet basis function is used to perform continuous wavelet transform on the three-phase current waveform signal to generate a time-frequency diagram.
[0038] Specifically, select operating condition A as the source domain, and simulate and obtain the fault data under the normal operating state and six open-circuit fault states based on the MATLAB / Simulink platform to train the model. The sampling frequency of the system is 5 kHz, and the rated frequency is 50 Hz, ensuring that the sampling frequency is high enough to capture the high-frequency characteristics of the signal. The fault diagnosis task needs to be completed within a fundamental wave period of 0.02 s. Therefore, the three-phase current waveform data within a fundamental wave period constitutes a 1×300 data sample. For each fault and normal operating state, 200 data samples are collected, and a total of 1400 source domain data samples are obtained. Then, collect the labeled data samples under operating condition B as the target domain data based on the hardware-in-the-loop platform. 40 data samples are collected for each of the seven operating states. In addition, 700 unlabeled data samples under operating condition B are also collected, including all seven operating states and with different quantities, simulating the unlabeled signals obtained in industrial practice for the training process of pseudo-label generation and iterative optimization.
[0039] Wavelet analysis is a time-frequency analysis method. It is based on the same principle as the windowed Fourier transform with an adaptively changing window width. By changing the scale factor and translation factor, the wavelet basis function is used to identify and separate the low-frequency and high-frequency components of the signal. Compared with the short-time Fourier transform, which is also a time-frequency method, wavelet transform can better display the time-frequency resolution performance. Compared with discrete wavelets, the coefficients of continuous wavelet transform (CWT) can more completely characterize the signal features, which is beneficial to the identification and analysis of similar signals. Since the fault signal features of each power IGBT in the inverter are similar, continuous wavelet transform is used for its identification and analysis. Considering that the function waveform of the continuous wavelet should be similar to the fault signal features of the power IGBT in the inverter, this patent selects the cmor wavelet basis. By setting the scale factor range from 1 to 100, the cmor3-3 wavelet basis function is used to perform continuous wavelet transform on the three-phase current signal to generate a time-frequency image to retain the local features of the signal.
[0040] According to an embodiment of the present application, the convolutional neural network model includes 5 convolutional layers, 3 pooling layers, 2 fully connected layers, and 1 Dropout layer. The Softmax is used as the activation function for the final output layer, and the Leaky ReLU is used as the activation function for the remaining convolutional layers and fully connected layers.
[0041] Specifically, in a convolutional neural network, the selection of the number of network layers needs to balance the feature extraction ability and the model complexity: too few layers are likely to cause feature loss, and too many layers will introduce redundant parameters. The deep convolutional neural network constructed in this patent increases the network capacity, greatly improving the network's ability to extract deep features from complex images. At the same time, by adding a Dropout layer, the risk of overfitting is reduced; also, the use of pseudo-label iteration optimization further improves the generalization ability of the model. Through comparative experiments, the optimal network structure under pseudo-label iteration optimization as shown in Figure 1 is selected, including 5 convolutional layers for feature extraction, 3 pooling layers for dimensionality reduction and increasing the spatial invariance of features, 2 fully connected layers for high-level feature integration, and 1 Dropout layer for avoiding overfitting. The Softmax is used as the activation function for the final output layer, and the Leaky ReLU is used as the activation function for the remaining convolutional layers and fully connected layers. The sample batch size is 20, and the Adam optimizer is used. The source domain data is input into the above convolutional neural network model to obtain the pre-trained model α.
[0042] According to an embodiment of the present application, constructing a fault diagnosis model based on transfer learning on the basis of the pre-trained model α includes: constructing a fault diagnosis model based on transfer learning on the basis of the pre-trained model α includes: migrating the network structure and parameters of the pre-trained model α to the sample set 2 of the target domain, freezing the first three convolutional layers of the network, keeping their parameters unchanged, and then fine-tuning the parameters of the deep network structure. The training set and test set in the target domain sample set 2 are used to train the model, and the validation set in the target domain sample set 2 is used to verify the accuracy of the trained model, obtaining the transfer model β that meets the accuracy requirements.
[0043] Specifically, the goal of transfer learning is to build a model based on data with sufficient samples and labels (i.e., source domain data) to predict and identify the actual engineering data with insufficient samples and labels under different working conditions (i.e., target domain data). Due to different equipment working conditions, the probability distributions of data sets under other working conditions are also different, and the function established in the source domain cannot be directly used for fault classification in the target domain. Therefore, transfer learning helps the target domain model adapt to the data distribution under different working conditions by leveraging the existing knowledge in the source domain.
[0044] This application first uses the fault dataset under operating condition A with sufficient samples to train the network, determine the structure and parameters of the network, so that the network can learn standard features. Then, the general network structure and parameters are migrated to the small-sample dataset under operating condition B. Subsequently, the first three convolutional layers of the network are frozen, and their parameters are fixed. Then, the parameters of the deep network structure are fine-tuned, and a part of the small-sample labeled data under operating condition B is used for retraining, so that most of the network parameters remain unchanged, and only the deep network structure learns the features under different operating conditions. Finally, the transfer model β is obtained.
[0045] According to an embodiment of the present application, the transfer model β is used to predict the unlabeled samples in the unlabeled sample set 3, generate pseudo-labels, and the consistency regularization technique is adopted to smooth the pseudo-label prediction results for the subsequent training of the transfer model β to obtain the final optimized model γ; the early stopping algorithm is introduced during the subsequent training process of the transfer model β, and it is determined whether to terminate the training in advance according to the performance of the validation set to avoid overfitting.
[0046] Specifically, in industrial operation, the annotation cost of the operation data of electrical equipment is very high. Under such a premise, using pseudo-label iterative optimization can utilize the information in a large amount of unlabeled data to expand the training data volume. By continuously iteratively optimizing the pseudo-labels, the pseudo-label iterative optimization can expand the training dataset and help the model obtain better prediction results when the labeled data is limited. During the iterative optimization process, the model needs to process pseudo-labels of different qualities, which helps the model learn more stable and robust feature representations, enhance the model's adaptability to noise and data changes, and at the same time reduce the risk of overfitting.
[0047] The fine-tuned transfer model β is used to predict the unlabeled data under operating condition B, and the prediction results are assigned as pseudo-labels to the unlabeled data. The model is retrained using the new training data containing pseudo-labels. To ensure the quality of the pseudo-labels, the consistency regularization method is adopted to optimize the reliability of the labels by smoothing the pseudo-label prediction results. To ensure that the iterative process is not interfered by samples with too low confidence, the pseudo-labels with high uncertainty are discarded, and only the pseudo-labels with low uncertainty and high credibility are retained for subsequent training, and the confidence threshold is set to 70%. In addition, the early stopping algorithm is introduced during the iterative process, and it is determined whether to terminate the training in advance according to the improvement of the validation set performance to avoid the risk of overfitting. A total of three iterations are performed to obtain the final optimized model γ.
[0048] Corresponding to the three - phase inverter open - circuit fault diagnosis method based on deep transfer learning provided by the above - mentioned several embodiments, an embodiment of the present application also provides a three - phase inverter open - circuit fault diagnosis device based on deep transfer learning. Since the three - phase inverter open - circuit fault diagnosis device based on deep transfer learning provided by the embodiments of the present application corresponds to the three - phase inverter open - circuit fault diagnosis method based on deep transfer learning provided by the above - mentioned several embodiments, the implementation manners of the above - mentioned three - phase inverter open - circuit fault diagnosis method based on deep transfer learning are also applicable to the three - phase inverter open - circuit fault diagnosis device provided by this embodiment, and will not be described in detail in this embodiment.
[0049] A three - phase inverter open - circuit fault diagnosis device provided by an embodiment of the present invention includes:
[0050] Module 100 is used to collect the labeled sample set 1 of the three - phase inverter in the normal state and N fault states in the source domain; collect the labeled sample set 2 of the normal state and N fault states in the target domain, and the unlabeled sample set 3 in the target domain;
[0051] Module 200 constructs a convolutional neural network model for fault diagnosis, and uses the sample set 1 in the source domain to train the model to obtain a pre - trained model α;
[0052] Module 300 constructs a fault diagnosis model based on transfer learning on the basis of the pre - trained model α, and uses the sample set 2 in the target domain to train the above - mentioned model to obtain a transfer model β;
[0053] Module 400 uses the transfer model β to predict the unlabeled samples in the unlabeled sample set 3, generates pseudo - labels, and adopts a consistency regularization technique to smooth the pseudo - label prediction results for subsequent training of the transfer model β to obtain a final optimized model γ;
[0054] Module 500 is used to input the data of the three - phase inverter to be detected into the model γ and output the predicted health status.
[0055] It should be understood that various forms of processes shown above can be used, re - ordering, adding or deleting steps. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0056] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0057] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined. Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0058] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0059] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or combinations thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0060] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0061] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0062] The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A three-phase inverter open circuit fault diagnosis method based on deep transfer learning, characterized in that: The steps include: Step 100 is used to collect a labeled sample set 1 of a three-phase inverter in a normal state and N fault states in a source domain, collect a labeled sample set 2 of a three-phase inverter in a normal state and N fault states in a target domain, and an unlabeled sample set 3 in the target domain; Step 200, constructing a convolutional neural network model for fault diagnosis, and training the model using a sample set 1 of the source domain to obtain a pre-trained model α; Step 300, constructing a fault diagnosis model based on transfer learning based on the pre-trained model α, and training the model using the sample set 2 of the target domain to obtain a transfer model β; Step 400, using the transfer model β to predict the unlabeled samples in the unlabeled sample set 3, generating pseudo labels, and using the consistency regularization technology to smooth the pseudo label prediction results for subsequent training of the transfer model β to obtain the final optimized model γ; Step 500: input the data of the three-phase inverter to be detected into the model γ, and output the predicted health status.
2. The three-phase inverter open circuit fault diagnosis method based on deep transfer learning according to claim 1, characterized in that: The method of obtaining samples in each sample set in step 100 is: collecting three-phase current waveform data of the three-phase inverter, performing wavelet transform on it and generating a time-frequency diagram as the sample.
3. The three-phase inverter open circuit fault diagnosis method based on deep transfer learning as claimed in claim 2, characterized in that: The cmor3-3 wavelet basis function is used to perform continuous wavelet transform on the three-phase current waveform signal to generate a time-frequency diagram.
4. The three-phase inverter open circuit fault diagnosis method based on deep transfer learning according to claim 1, characterized in that: The convolutional neural network model includes 5 convolutional layers, 3 pooling layers, 2 fully connected layers and 1 Dropout layer. The final output layer uses Softmax as the activation function, and the remaining convolutional layers and fully connected layers use Leaky ReLU as the activation function.
5. The three-phase inverter open circuit fault diagnosis method based on deep transfer learning according to claim 1, characterized in that: The method of constructing a fault diagnosis model based on transfer learning on the basis of the pre-trained model α includes: migrating the network structure and parameters of the pre-trained model α to the sample set 2 of the target domain, freezing the first three convolutional layers of the network and fixing their parameters, then fine-tuning the parameters of the deep network structure, training the model using the training set and test set in the target domain sample set 2, and verifying the accuracy of the trained model using the verification set in the target domain sample set 2, so as to obtain a migration model β that meets the accuracy requirements.
6. The three-phase inverter open circuit fault diagnosis method based on deep transfer learning according to claim 1, characterized in that: An early stopping algorithm is introduced in the subsequent training process of the migration model β, and whether to terminate the training early is determined based on the performance of the validation set to avoid overfitting.
7. A three-phase inverter open circuit fault diagnosis device based on deep transfer learning, characterized in that: include: Module 100, for collecting labeled sample sets 1 of three-phase inverters in normal states and N fault states in a source domain; Collect labeled sample set 2 in the normal state and N fault states in the target domain, and unlabeled sample set 3 in the target domain; Module 200, constructing a convolutional neural network model for fault diagnosis, using a sample set 1 of the source domain to train the model, and obtaining a pre-trained model α; Module 300, building a fault diagnosis model based on transfer learning based on the pre-trained model α, using the sample set 2 of the target domain to train the above model to obtain a transfer model β; Module 400, using the transfer model β to predict the unlabeled samples in the unlabeled sample set 3, generating pseudo labels, and using the consistency regularization technology to smooth the pseudo label prediction results for subsequent training of the transfer model β to obtain the final optimized model γ; Module 500 is used to input the data of the three-phase inverter to be detected into the model γ and output the predicted health status.
Citation Information
Cited By
Gearbox composite fault diagnosis method and storage medium
CN120805068A