Thermal power generator fault prediction and maintenance method based on deep learning model and large language model

By combining deep learning models and large language models, automatic diagnosis of thermal power unit faults and generation of intelligent maintenance suggestions are achieved, solving the problems of low efficiency and reliance on manual experience in existing technologies, and improving the accuracy of fault diagnosis and the intelligence of maintenance strategies.

CN120672313APending Publication Date: 2025-09-19BEIJING HELI INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510763764.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies for fault diagnosis and maintenance strategy generation of thermal power units are inefficient and rely heavily on manual experience, leading to unplanned downtime and maintenance losses.

Method used

A comprehensive approach based on deep learning models and large language models is adopted. By constructing a deep learning model with Transformer encoder and cross-attention mechanism, combined with a large language model training set, automatic diagnosis of thermal power unit faults and generation of intelligent maintenance suggestions are achieved.

Benefits of technology

It improves the accuracy of fault diagnosis of thermal power units and the intelligence of maintenance strategies, reduces unplanned downtime and maintenance losses, and improves the operating stability and economic benefits of thermal power units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672313A_ABST
    Figure CN120672313A_ABST
Patent Text Reader

Abstract

The invention discloses a thermal power generator fault prediction and maintenance method based on a deep learning model and a large language model, belongs to the field of machinery, and particularly relates to a thermal power generator fault prediction and maintenance method. The objective of the invention is to solve the problems of overlong time of application of a large language model in thermal power machinery fault and low diagnosis and maintenance efficiency in a recovery process in the prior art. The method comprises the following steps: acquiring vibration data and temperature data of a gearbox under normal working conditions and different fault types of a thermal power generator; obtaining a trained deep learning model; constructing a question and answer pair composed of the fault type and the maintenance mode; obtaining a trained large language model; inputting gear box vibration data and temperature data of the to-be-tested thermal power generator into the trained deep learning model, outputting whether the thermal power generator has a fault, and if no fault exists, outputting a result; if so, outputting a fault type; and inputting the output fault type into the trained large language model, and outputting a maintenance mode corresponding to the fault type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machinery, and in particular relates to a method for predicting and maintaining a thermal generator fault. Background Art

[0002] It's estimated that approximately 60% to 70% of global electricity supply relies on thermal power generation. The efficient and stable operation of thermal power units is crucial to energy security. However, thermal power units are subject to the harsh operating conditions of high temperature, high pressure, and high dust levels for extended periods of time. Hidden failures in key equipment (such as boiler pipes, turbine blades, and generator bearings) can easily lead to unplanned downtime. Statistics show that after a serious failure in a thermal power unit, the average power generation efficiency drops by more than 12%, and the recovery period can take over 72 hours, significantly impacting the grid's peak-shaving capacity and power supply continuity.

[0003] Manual maintenance of key components in thermal power units faces multiple challenges. First, internal boiler inspections must be conducted in extremely high temperatures (>500°C) and confined spaces, posing significant risks to personnel safety. Second, fault diagnosis and maintenance decisions rely heavily on expert experience, and complex faults such as abnormal turbine vibration and generator bearing wear are difficult to quickly locate. The average time from fault occurrence (T1) to accurate diagnosis (T2) accounts for 65%-75% of total downtime, further exacerbating power generation losses. For example, a 660MW supercritical unit failed to promptly identify microcracks in the boiler water-wall tubes, resulting in a 48-hour delay in the T2 phase and direct economic losses exceeding 8 million yuan.

[0004] Thermal power units use coal-fired boilers to generate high-temperature steam, which drives the turbine rotor, which in turn drives the generator to output electrical energy. The reliability of the turbine-generator shafting directly determines the unit's operational stability. Failures such as rotor imbalance and the loss of the babbitt alloy coating on bearings can trigger a chain reaction, leading to shafting failures in severe cases. Therefore, intelligent fault diagnosis and predictive maintenance technologies for turbines, generators, and auxiliary systems have become core issues for ensuring the safe and economical operation of thermal power plants.

[0005] Current thermal power plant fault diagnosis primarily relies on multimodal methods such as vibration spectrum analysis, infrared thermal imaging monitoring, and oil wear particle detection. For example, an improved spatiotemporal convolutional neural network can provide early warning of tube wall creep damage by fusing boiler tube wall temperature fields with steam pressure fluctuation data. A dynamic graph attention network, combined with turbine bearing vibration signals and lubricant metal content data, can accurately identify the extent of bearing wear. However, existing methods primarily focus on fault detection and lack the ability to intelligently generate maintenance strategies. For example, when a boiler tube burst or turbine blade crack is detected, manual experience is still required to develop a shutdown and maintenance plan, and the rationality of the plan is limited by the technical level of the on-site engineers.

[0006] Thermal power unit maintenance faces a dual dilemma: First, coupled failures in complex equipment systems (such as desulfurization and denitrification units and feedwater pumps) require cross-disciplinary coordination, making single-domain maintenance solutions prone to secondary risks. Second, maintenance operations in high-temperature, high-pressure environments have a very low tolerance for error. Heavy operations, such as turbine openings for overhaul, can lead to secondary damage such as seal failure and rotor dynamic balance loss if spare parts mismatches or torque calibration errors occur. In one case, a 1000MW unit experienced a 3.5% torque error during generator bearing replacement, resulting in a shutdown and adjustment, resulting in an additional loss of 12 million kWh of power generation. Summary of the Invention

[0007] The purpose of the present invention is to solve the problem that the large language model in the existing technology is applied to thermal power machinery failure for a long time and the diagnosis and maintenance efficiency of the recovery process is low, and to propose a thermal power generator fault prediction and maintenance method based on deep learning model and large language model.

[0008] The specific process of the thermal generator fault prediction and maintenance method based on deep learning model and large language model is as follows:

[0009] Step 1: Collect the gearbox vibration data and temperature data under normal operating conditions of the thermal power generator, as well as the gearbox vibration data and temperature data under different fault types of the thermal power generator, as a deep learning model training set;

[0010] Step 2: Build a deep learning model, which includes a Transformer encoder and a cross-attention mechanism;

[0011] Train the deep learning model based on the deep learning model training set in step 1 to obtain a trained deep learning model;

[0012] Step 3: Construct question-answer pairs consisting of fault type and repair method, and build a large language model training set based on the question-answer pairs;

[0013] Step 4: Based on the large language model training set, obtain a trained large language model;

[0014] Step 5: Input the gearbox vibration data and temperature data of the thermal generator to be tested into the trained deep learning model. The trained deep learning model outputs whether the thermal generator is faulty. If there is no fault, the result is output; if there is a fault, the fault type is output.

[0015] Step 6: Input the fault type output in step 5 into the trained large language model, and the trained large language model outputs the maintenance method corresponding to the fault type.

[0016] Preferably, in step 1, the gearbox vibration data and temperature data under normal working conditions of the thermal power generator, as well as the gearbox vibration data and temperature data under different fault types of the thermal power generator are collected as a deep learning model training set; the specific process is:

[0017] The gearbox vibration and temperature data of thermal generators under normal working conditions, as well as the gearbox vibration and temperature data of thermal generators under different fault types, were collected from the Gearbox Fault Dataset released by EPRI as training sets for the deep learning model.

[0018] Preferably, in step 2, a deep learning model is constructed, and the deep learning model includes a Transformer encoder and a cross attention mechanism;

[0019] Train the deep learning model based on the deep learning model training set in step 1 to obtain a trained deep learning model;

[0020] The specific process is:

[0021] Step 21: Build a deep learning model, which includes a Transformer encoder and a cross-attention mechanism;

[0022] Step 22:

[0023] The gearbox vibration data of the thermal generator under normal working conditions and the gearbox vibration data of the thermal generator under different fault types in the deep learning model training set are input into the Transformer encoder, and the Transformer encoder outputs the feature F s ;

[0024] The gearbox temperature data of the thermal generator under normal working conditions and the gearbox temperature data of the thermal generator under different fault types in the deep learning model training set are input into the Transformer encoder, and the Transformer encoder outputs the feature F l ;

[0025] Steps 2 and 3

[0026] The feature F output by the Transformer encoder s Input linear layer, linear layer output query Q;

[0027] The feature F output by the Transformer encoder l Input linear layer, linear layer output key K;

[0028] The feature F output by the Transformer encoder l Input linear layer, linear layer output value V;

[0029] Based on the query Q, key K and value V, calculate the cross attention O; expressed as:

[0030]

[0031] Among them, the superscript T means to find the transpose, represents the dimension of the key vector K;

[0032] Step 24

[0033] The feature F output by the Transformer encoder l Input linear layer, linear layer output query Q′;

[0034] The feature F output by the Transformer encoder s Input linear layer, linear layer output key K′;

[0035] The feature F output by the Transformer encoder s Input linear layer, linear layer output value V′;

[0036] Based on the query Q′, key K′ and value V′, calculate the cross attention O′; expressed as:

[0037]

[0038] Among them, the superscript T means to find the transpose, represents the dimension of the key vector K′;

[0039] Step 25

[0040] The cross attention O passes through the linear layer and the softmax layer in sequence, and the softmax layer outputs the feature

[0041] The cross attention O′ passes through the linear layer and the softmax layer in sequence, and the softmax layer outputs the feature

[0042] Step 26: Features and features Perform element-by-element summation, and the summed features are used as the output features of the deep learning model;

[0043] Step 27: Train the deep learning model until the loss function converges to obtain a trained deep learning model.

[0044] Preferably, the loss function in step 27 is a cross entropy loss function.

[0045] Preferably, in step 3, question-answer pairs consisting of fault type and maintenance method are constructed, and a large language model training set is constructed based on the question-answer pairs;

[0046] The specific process is:

[0047] 1) Obtain the maintenance log of the thermal power plant from the open source dataset, use the pymupdf library to extract the first page of text from the maintenance log of the wind farm, and save it in variable A;

[0048] 2) Input the text stored in variable A into the large language model. The large language model outputs a question-answer pair consisting of the fault type and the repair method. The question-answer pair consisting of the fault type and the repair method output by the large language model is saved in variable B.

[0049] 3) Save variable B as a json file;

[0050] 4) Repeat 1) to 3) until all pages of the wind farm maintenance log are saved as JSON files as the large language model training set.

[0051] Preferably, in step 4, a trained large language model is obtained based on the large language model training set; the specific process is:

[0052] Input the large language model training set into the large language model, and use the Lora optimization method to train the large language model until the large language model loss function converges to obtain a trained large language model;

[0053] The large language model loss function is L(θ); it is expressed as:

[0054]

[0055] in,

[0056] L(θ) is the large language model loss function;

[0057] θ is the large language model parameter;

[0058] D is the large language model training set;

[0059] (x i ,y i ) is the i-th sample in the large language model training set, x i For input text, y i is the corresponding output text; α, β, γ are weight coefficients;

[0060] f(x i ; θ) is the predicted distribution;

[0061] L CE is the cross entropy loss;

[0062] L FL For loss of fluency;

[0063] L KL is the knowledge distillation loss.

[0064] Preferably, the cross entropy loss L CE Expressed as:

[0065]

[0066] Among them, p i is the probability distribution of the i-th sample predicted by the model, q i is the distribution of the true label of the i-th sample. Preferably, the fluency loss L FL Expressed as:

[0067]

[0068] in,

[0069] P data is the data distribution;

[0070] Indicates the variable y in the real data P data Calculation of expected value under distribution;

[0071] T′ is the length of the output sequence;

[0072] P(y t |y <t ; θ) is the probability of the t-th word given the previous t-1 words.

[0073] Preferably, the knowledge distillation loss L KL Expressed as:

[0074] L KL (θ)=D KL (P student (y|x;θ)||P teacher (y|x))

[0075] in,

[0076] P student (y|x; θ) is the output probability distribution of the student model;

[0077] P teacher (y|x) is the output probability distribution of the teacher model;

[0078] D KL is the KL divergence;

[0079] DKL (P student (y|x;θ)||P teacher (y|x)) is a measure of the probability distribution P student (y|x;θ) and the probability distribution P teacher The difference between (y|x).

[0080] The beneficial effects of the present invention are:

[0081] This method is an innovative fault diagnosis and maintenance method. The overall structure is as follows Figure 2 As shown in the figure, the system of this method includes two core modules: a fault diagnosis module and a large model maintenance suggestion module. It integrates the diagnostic ability of neural networks in vibration signal analysis with the reasoning and text generation capabilities of large language models, providing a new method for fault diagnosis and efficient maintenance of thermal power generation equipment.

[0082] The innovation of the present invention is:

[0083] The cross-attention mechanism is integrated into the Transformer and applied to thermal power plant fault diagnosis. The improved Transformer network can capture the dependencies between long and short time windows of multi-source signals in thermal power plants, enhance the fusion of contextual information, and improve the accuracy of fault diagnosis tasks.

[0084] An innovative large language model fine-tuning method is proposed, which combines cross-entropy loss, fluency loss, and knowledge distillation into a new loss function to improve the accuracy, fluency, and generalization ability of the fine-tuned model.

[0085] An innovative dataset generation method for large language model training has been developed to convert professional knowledge into question-answer data format, and to convert professional knowledge and work experience suitable for human learning into a data format that can be learned by large language models. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 Schematic diagram of the overall structure of the model of the present invention;

[0087] Figure 2 This is a fault diagnosis module diagram of the present invention;

[0088] Figure 3 This is a data set conversion flow chart of the present invention;

[0089] Figure 4 This is the flow chart of fine-tuning the large model of the present invention, A represents the low-rank projection matrix, N(0,σ 2 ) means the mean is 0 and the variance is σ 2 Gaussian distribution, B represents the low-rank reconstruction matrix;

[0090] Figure 5 Experimental diagram for performance evaluation of the method;

[0091] Figure 6a For maintenance experts, trainee maintenance personnel, basic large model, the total equal division comparison diagram of the method of the present invention;

[0092] Figure 6b Comparison diagram of hazard assessment distribution for maintenance experts, trainee maintenance personnel, basic large model, and the method of the present invention;

[0093] Figure 6c Comparative graphs of evaluation distributions are provided for maintenance experts, trainee maintenance personnel, basic large-scale models, and the method of the present invention;

[0094] Figure 6d Comparative diagram of distribution of evaluation for maintenance expert, trainee maintenance worker, basic large model, and steps of the method of the present invention;

[0095] Figure 7a This is the confusion matrix diagram of the EhCNN diagnosis method accuracy;

[0096] Figure 7b This is the confusion matrix diagram of the Resnet50 diagnosis method accuracy;

[0097] Figure 7c This is the confusion matrix diagram of the accuracy of the Uniformer diagnosis method;

[0098] Figure 7d This is a confusion matrix diagram of the accuracy of the diagnostic method of the present invention;

[0099] Figure 8a This is the classification result diagram of the EhCNN method. t-SNE component 1 and t-SNE component 2 are clustering outputs of mapping high-dimensional data to low-dimensional space.

[0100] Figure 8b This is the classification result diagram of the Resnet50 method;

[0101] Figure 8c This is the classification result diagram of the Uniformer method;

[0102] Figure 8d This is the classification result diagram of the method of the present invention. DETAILED DESCRIPTION

[0103] Specific implementation method 1: Combination Figure 1 This embodiment describes the method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model. The specific process is as follows:

[0104] Step 1: Collect the gearbox vibration data and temperature data under normal operating conditions of the thermal power generator, as well as the gearbox vibration data and temperature data under different fault types of the thermal power generator, as a deep learning model training set;

[0105] Step 2: Build a deep learning model, which includes a Transformer encoder and a cross-attention mechanism;

[0106] Train the deep learning model based on the deep learning model training set in step 1 to obtain a trained deep learning model;

[0107] Step 3: Construct question-answer pairs consisting of fault type and repair method, and build a large language model training set based on the question-answer pairs;

[0108] Step 4: Based on the large language model training set, obtain a trained large language model;

[0109] Step 5: Input the gearbox vibration data and temperature data of the thermal generator to be tested into the trained deep learning model. The trained deep learning model outputs whether the thermal generator is faulty. If there is no fault, the result is output; if there is a fault, the fault type is output.

[0110] Step 6: Input the fault type output in step 5 into the trained large language model, and the trained large language model outputs the maintenance method corresponding to the fault type.

[0111] Specific embodiment 2: This embodiment differs from specific embodiment 1 in that: in step 1, the gearbox vibration data and temperature data under normal working conditions of the thermal generator, as well as the gearbox vibration data and temperature data under different fault types of the thermal generator are collected as a deep learning model training set; the specific process is:

[0112] The gearbox vibration and temperature data of thermal generators under normal working conditions, as well as the gearbox vibration and temperature data of thermal generators under different fault types, were collected from the Gearbox Fault Dataset released by EPRI as training sets for the deep learning model.

[0113] Other steps and parameters are the same as those in the first embodiment.

[0114] Specific implementation method three: The difference between this implementation method and specific implementation method one or two is that in step two, a deep learning model is constructed, and the deep learning model includes a Transformer encoder and a cross attention mechanism, such as Figure 2 ;

[0115] Train the deep learning model based on the deep learning model training set in step 1 to obtain a trained deep learning model;

[0116] The specific process is:

[0117] Step 21: Build a deep learning model, which includes a Transformer encoder and a cross-attention mechanism;

[0118] Step 22:

[0119] The gearbox vibration data of the thermal generator under normal working conditions and the gearbox vibration data of the thermal generator under different fault types in the deep learning model training set are input into the Transformer encoder, and the Transformer encoder outputs the feature F s ;

[0120] The gearbox temperature data of the thermal generator under normal working conditions and the gearbox temperature data of the thermal generator under different fault types in the deep learning model training set are input into the Transformer encoder, and the Transformer encoder outputs the feature F l ;

[0121] Steps 2 and 3

[0122] The feature F output by the Transformer encoder s Input linear layer, linear layer output query Q;

[0123] The feature F output by the Transformer encoder l Input linear layer, linear layer output key K;

[0124] The feature F output by the Transformer encoder l Input linear layer, linear layer output value V;

[0125] Based on the query Q, key K and value V, calculate the cross attention O; expressed as:

[0126]

[0127] Among them, the superscript T means to find the transpose, represents the dimension of the key vector K;

[0128] Step 24

[0129] The feature F output by the Transformer encoder l Input linear layer, linear layer output query Q′;

[0130] The feature F output by the Transformer encoder s Input linear layer, linear layer output key K′;

[0131] The feature F output by the Transformer encoder s Input linear layer, linear layer output value V′;

[0132] Based on the query Q′, key K′ and value V′, calculate the cross attention O′; expressed as:

[0133]

[0134] Among them, the superscript T means to find the transpose, represents the dimension of the key vector K′;

[0135] Step 25

[0136] The cross attention O passes through the linear layer and the softmax layer in sequence, and the softmax layer outputs the feature

[0137] The cross attention O′ passes through the linear layer and the softmax layer in sequence, and the softmax layer outputs the feature

[0138] Step 26: Features and features Perform element-by-element summation, and the summed features are used as the output features of the deep learning model;

[0139] Step 27: Train the deep learning model until the loss function converges to obtain a trained deep learning model.

[0140] Other steps and parameters are the same as those in the first or second embodiment.

[0141] Specific embodiment 4: This embodiment differs from any one of specific embodiments 1 to 3 in that the loss function in step 27 is a cross entropy loss function.

[0142] The other steps and parameters are the same as those in the first to third embodiments.

[0143] Specific implementation method 5: This implementation method is different from one of the specific implementation methods 1 to 4 in that: in step 3, a question-answer pair consisting of the fault type and the maintenance method is constructed, and a large language model training set is constructed based on the question-answer pair, such as Figure 3 ;

[0144] The specific process is:

[0145] 1) Obtain the maintenance log of the thermal power plant from the open source dataset (https: / / wenku.baidu.com / view / c57ab523f9d6195f312b3169a45177232f60e49d.html?_wkts_=1745739141265), use the pymupdf library to extract the first page of text from the maintenance log of the wind farm, and save it in variable A;

[0146] 2) Input the text stored in variable A into the large language model. The large language model outputs a question-answer pair consisting of the fault type and the repair method. The question-answer pair consisting of the fault type and the repair method output by the large language model is saved in variable B.

[0147] 3) Save variable B as a json file;

[0148] 4) Repeat 1) to 3) until all pages of the wind farm maintenance log are saved as JSON files as the large language model training set.

[0149] Question and answer format:

[0150] {"instruction":"What are the two main domains in which vibration data is typically presented?","input":"","output":"Is vibration data typically presented in the time domain or the frequency domain?"},{"instruction":"How are the amplitude and frequency of a vibration signal typically described in vibration analysis?","input":"","output":"The terms used to describe the amplitude and frequency of a vibration signal are typically defined by their appearance in the time domain signal, where amplitude refers to the height of the waveform and frequency refers to the number of cycles per unit time.},{"instruction":"What are discrete frequencies in vibration analysis?","input":"","output":"A discrete frequency in vibration analysis refers to a specific frequency that can be individually identified based on its amplitude and frequency."},{"instruction":"How does random noise differ from a discrete signal in vibration analysis?","input":"","output":"Random noise in vibration analysis is characterized by a mixture of many frequencies where the amplitude of each frequency varies randomly. In contrast, a discrete signal consists of identifiable frequencies with consistent amplitudes."}

[0151] The overall construction method of the dataset conversion process is as follows Figure 5 As shown in the figure, firstly, text data is extracted from the filtered relevant information, and then it is input into the large language model with prompt words, and the large language model is supervised to convert it into the question-answer pair format.

[0152] Other steps and parameters are the same as those in Specific Embodiments 1 to 4-1.

[0153] Specific implementation method 6: This implementation method is different from any one of the specific implementation methods 1 to 5 in that: in step 4, a trained large language model is obtained based on a large language model training set, such as Figure 4 ;

[0154] The specific process is:

[0155] Input the large language model training set into the large language model, and use the Lora optimization method to train the large language model until the large language model loss function converges to obtain a trained large language model;

[0156] The large language model loss function is L(θ); it is expressed as:

[0157]

[0158] in,

[0159] L(θ) is the large language model loss function, which is used to fine-tune the model parameters θ;

[0160] θ is the large language model parameter;

[0161] D is the large language model training set;

[0162] (x i ,y i ) is the i-th sample in the large language model training set, x i For input text (fault type), y i is the corresponding output text (maintenance method);

[0163] α, β, and γ are weight coefficients used to balance the importance of different loss terms;

[0164] f(x i ; θ) is the predicted distribution;

[0165] L CE is the cross entropy loss, which measures the predicted distribution f(x i ;θ) and the true label y i the differences between;

[0166] L FL For loss of fluency;

[0167] L KL is the knowledge distillation loss.

[0168] Other steps and parameters are the same as those in Specific Implementations 1 to 5-1.

[0169] Specific embodiment seven: This embodiment differs from any one of specific embodiments one to six in that: the cross entropy loss L CE Expressed as:

[0170]

[0171] Among them, p i is the probability distribution of the i-th sample predicted by the model, q i is the distribution of the true label of the i-th sample.

[0172] The other steps and parameters are the same as those in the first to sixth embodiments.

[0173] Specific embodiment eight: This embodiment differs from any one of specific embodiments one to seven in that: the fluency loss L FL Expressed as:

[0174]

[0175] in,

[0176] P data is the data distribution;

[0177] Indicates the variable y in the real data P data Calculation of expected value under distribution;

[0178] T′ is the length of the output sequence;

[0179] P(y t |y <t ; θ) is the probability of the t-th word given the previous t-1 words;

[0180] Fluency loss enhances the coherence and naturalness of generated text by introducing a probability scoring mechanism of the language model.

[0181] Other steps and parameters are the same as those in Specific Embodiments 1 to 7-1.

[0182] Specific embodiment 9: This embodiment differs from any one of specific embodiments 1 to 8 in that: the knowledge distillation loss L KL Expressed as:

[0183] L KL (θ)=D KL (P student (y|x;θ)||P teacher (y|x))

[0184] in,

[0185] P student (y|x; θ) is the output probability distribution of the student model (the fine-tuned model);

[0186] P teacher(y|x) is the output probability distribution of the teacher model (original model);

[0187] D KL is the KL divergence, which measures the difference between two probability distributions;

[0188] D KL (P student (y|x;θ)||P teacher (y|x)) is a measure of the probability distribution P student (y|x;θ) and the probability distribution P teacher The difference between (y|x).

[0189] Knowledge distillation loss is used to maintain the consistency between the fine-tuned model and the original model, ensuring the effectiveness of transfer learning.

[0190] The other steps and parameters are the same as those in the specific implementation modes 1 to 8-1.

[0191] Example:

[0192] Experimental process:

[0193] Three experiments were designed to evaluate the system performance. Figure 6a 、 Figure 6b 、 Figure 6c 、 Figure 6d As shown in the figure, Experiment 1 aims to compare the accuracy of the LLM module of the system in providing maintenance steps and post-maintenance suggestions, and compare it with the performance of existing models and maintenance experts. Experiment 2 aims to evaluate the accuracy of the Signal-Transformer module of the system in gearbox fault diagnosis. Experiment 3 tests the impact of the system on maintenance personnel of different technical levels in a real environment to verify its effectiveness and economy in practical applications. The process is as follows Figure 5 As shown;

[0194] Experimental results:

[0195] Experiment 1 prediction model LLM module performance comparison:

[0196] The goal of Experiment 1 was to evaluate the accuracy of the LLM module in providing repair procedures and subsequent maintenance recommendations. To this end, we collected a set of standard cases, including common gearbox failure scenarios and their correct repair procedures and recommendations. These cases covered a variety of fault types, such as gear wear and bearing damage, ensuring comprehensive and representative testing.

[0197] First, the LLM module was used to process these cases and record the repair and maintenance steps it provided. At the same time, 6 maintenance experts with more than 5 years of work experience and 6 maintenance workers who were in internship were invited to independently handle the same cases and record their repair suggestions. In addition, the present invention also used the unadjusted LLaMA3-8B model as the base model for comparison. The comparison results are shown in Figures 6a and 6b. Figure 6b 、 Figure 6c 、 Figure 6d As shown;

[0198] Experiment 2: Signal-Transformer performance evaluation:

[0199] The purpose of Experiment 2 is to verify the accuracy of the proposed method in gearbox fault diagnosis. This study uses a wind turbine gearbox vibration signal dataset containing 494,819 data points, which covers various gear fault types such as gear fracture and bearing damage.

[0200] The Signal-Transformer model was trained using this data and its performance was evaluated on an independent test set. The test set contains 10% of the data to ensure data diversity and representativeness. The accuracy, recall, and F1 score of the model in fault identification were calculated and compared with existing fault diagnosis models. The experimental results are shown in Figure 1. Figure 7a 、 Figure 7b 、 Figure 7c 、 7d As shown, Figure 7a 、 Figure 7b 、 Figure 7c The figure shows the performance of the classic diagnostic model on the data set. Figure 7d The figure shows the performance of the present invention on the data set.

[0201] Figure 7a This is the confusion matrix diagram of the EhCNN diagnosis method accuracy. Figure 7b This is the confusion matrix diagram of the Resnet50 diagnosis method accuracy. Figure 7c This is the confusion matrix diagram of the Uniformer diagnosis method accuracy. Figure 7d This is a confusion matrix diagram of the accuracy of the diagnostic method of the present invention;

[0202] Figure 8a This is the classification result diagram of the EhCNN method. Figure 8b This is the classification result diagram of the Resnet50 method. Figure 8c This is the classification result diagram of the Uniformer method. Figure 8d This is the classification result diagram of the method of the present invention.

[0203] Experimental results show that the proposed method performs exceptionally well in gearbox fault diagnosis, achieving 95.2% accuracy, 94.7% recall, and a 94.9% F1 score. Compared to existing fault diagnosis models, the proposed method significantly improves performance, particularly in identifying complex fault patterns. This demonstrates that the proposed method can effectively identify fault patterns within gearboxes, providing reliable support for subsequent repairs.

[0204] The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for predicting and maintaining thermal power generator faults based on a deep learning model and a large language model, characterized by: The specific process of the method is: Step 1: Collect the gearbox vibration data and temperature data under normal operating conditions of the thermal power generator, as well as the gearbox vibration data and temperature data under different fault types of the thermal power generator, as a deep learning model training set; Step 2: Build a deep learning model, which includes a Transformer encoder and a cross-attention mechanism; Train the deep learning model based on the deep learning model training set in step 1 to obtain a trained deep learning model; Step 3: Construct question-answer pairs consisting of fault type and repair method, and build a large language model training set based on the question-answer pairs; Step 4: Based on the large language model training set, obtain a trained large language model; Step 5: Input the gearbox vibration data and temperature data of the thermal generator to be tested into the trained deep learning model. The trained deep learning model outputs whether the thermal generator is faulty. If there is no fault, the result is output; if there is a fault, the fault type is output. Step 6: Input the fault type output in step 5 into the trained large language model, and the trained large language model outputs the maintenance method corresponding to the fault type.

2. The method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model according to claim 1, characterized in that: In the step 1, the gearbox vibration data and temperature data under normal working conditions of the thermal generator, as well as the gearbox vibration data and temperature data under different fault types of the thermal generator are collected as a deep learning model training set; the specific process is: The gearbox vibration and temperature data of thermal generators under normal working conditions, as well as the gearbox vibration and temperature data of thermal generators under different fault types, were collected from the Gearbox Fault Dataset released by EPRI as training sets for the deep learning model.

3. The method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model according to claim 2, characterized in that: In step 2, a deep learning model is constructed, which includes a Transformer encoder and a cross-attention mechanism; Train the deep learning model based on the deep learning model training set in step 1 to obtain a trained deep learning model; The specific process is: Step 21: Build a deep learning model, which includes a Transformer encoder and a cross-attention mechanism; Step 22: The gearbox vibration data of the thermal generator under normal working conditions and the gearbox vibration data of the thermal generator under different fault types in the deep learning model training set are input into the Transformer encoder, and the Transformer encoder outputs the feature F s ; The gearbox temperature data of the thermal generator under normal working conditions and the gearbox temperature data of the thermal generator under different fault types in the deep learning model training set are input into the Transformer encoder, and the Transformer encoder outputs the feature F l ; Steps 2 and 3 The feature F output by the Transformer encoder s Input linear layer, linear layer output query Q; The feature F output by the Transformer encoder l Input linear layer, linear layer output key K; The feature F output by the Transformer encoder l Input linear layer, linear layer output value V; Based on the query Q, key K and value V, calculate the cross attention O; expressed as: Among them, the superscript T means to find the transpose, represents the dimension of the key vector K; Step 24 The feature F output by the Transformer encoder l Input linear layer, linear layer output query Q′; The feature F output by the Transformer encoder s Input linear layer, linear layer output key K′; The feature F output by the Transformer encoder s Input linear layer, linear layer output value V′; Based on the query Q′, key K′ and value V′, calculate the cross attention O′; expressed as: Among them, the superscript T means to find the transpose, represents the dimension of the key vector K′; Step 25 The cross attention O passes through the linear layer and the softmax layer in sequence, and the softmax layer outputs the feature The cross attention O′ passes through the linear layer and the softmax layer in sequence, and the softmax layer outputs the feature Step 26: Features and features Perform element-by-element summation, and the summed features are used as the output features of the deep learning model; Step 27: Train the deep learning model until the loss function converges to obtain a trained deep learning model.

4. The method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model according to claim 3 is characterized in that: The loss function in step 27 is a cross entropy loss function.

5. The method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model according to claim 4 is characterized in that: In step 3, question-answer pairs consisting of fault type and maintenance method are constructed, and a large language model training set is constructed based on the question-answer pairs; The specific process is: 1) Obtain the maintenance log of the thermal power plant from the open source dataset, use the pymupdf library to extract the first page of text from the maintenance log of the wind farm, and save it in variable A; 2) Input the text stored in variable A into the large language model. The large language model outputs a question-answer pair consisting of the fault type and the repair method. The question-answer pair consisting of the fault type and the repair method output by the large language model is saved in variable B. 3) Save variable B as a json file; 4) Repeat 1) to 3) until all pages of the wind farm maintenance log are saved as JSON files as the large language model training set.

6. The method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model according to claim 5, characterized in that: In step 4, a trained large language model is obtained based on the large language model training set; the specific process is: Input the large language model training set into the large language model, and use the Lora optimization method to train the large language model until the large language model loss function converges to obtain a trained large language model; The large language model loss function is L(θ); it is expressed as: in, L(θ) is the large language model loss function; θ is the large language model parameter; D is the large language model training set; (x i ,y i ) is the i-th sample in the large language model training set, x i For input text, y i is the corresponding output text; α, β, and γ are weight coefficients; f(x i ; θ) is the predicted distribution; L CE is the cross entropy loss; L FL For loss of fluency; L KL is the knowledge distillation loss.

7. The method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model according to claim 6, characterized in that: The cross entropy loss L CE Expressed as: Among them, p i is the probability distribution of the i-th sample predicted by the model, q i is the distribution of the true label of the i-th sample.

8. The method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model according to claim 7, characterized in that: The fluency loss L FL Expressed as: in, P data is the data distribution; Indicates the variable y in the real data P data Calculation of expected value under distribution; T′ is the length of the output sequence; P(y t |y <t ; θ) is the probability of the t-th word given the previous t-1 words.

9. The method for predicting and maintaining faults of thermal power generators based on a deep learning model and a large language model according to claim 8, characterized in that: The knowledge distillation loss L KL Expressed as: L KL (θ)=D KL (P student (y|x:θ)||P teacher (y|x)) in, P student (y|x; θ) is the output probability distribution of the student model; P teacher (y|x) is the output probability distribution of the teacher model; D KL is the KL divergence; D KL (P student (y|x;θ)||P teacher (y|x)) is a measure of the probability distribution P student (y|x;θ) and the probability distribution P teacher The difference between (y|x).

Citation Information

Patent Citations

  • Power transmission line insulator fault diagnosis method based on deep learning

    CN116739996A

  • Underwater target detection and identification method based on bidirectional matching migration

    CN117457022A

  • Automobile fault analysis method, system and device

    CN117668038A

  • Language modal depolarization visual question answering method based on knowledge distillation

    CN118885586A

  • Large and small model cooperative training method and device for multi-modal large language model

    CN119514645A