Method and system for detecting generator components based on AI hybrid expert model
Through AI hybrid expert models and multi-view image acquisition technology, the problems of low accuracy and efficiency in generator component detection have been solved, efficient and automated detection and evaluation have been achieved, and the operating reliability and safety of the generator have been improved.
Patent Information
- Application Number
- CN202510590223.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies in generator component inspection suffer from low accuracy, low efficiency, non-real-time and incompleteness, especially poor results in identifying tiny or fuzzy defects, and the model's generalization ability and adaptability are insufficient.
A detection method based on an AI hybrid expert model is adopted. Through multi-view image acquisition, data preprocessing, hybrid expert model training and deployment fine-tuning, combined with data enhancement and pseudo-labeling technology, an efficient generator component detection system is constructed, including a perception unit, a data processing unit, and a decision output unit.
It improves the accuracy and efficiency of generator component status detection, realizes the automated detection of minor defects and damage, reduces maintenance costs, and improves the reliability and safety of generator operation.
Smart Images

Figure CN120672657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of generator component detection, and in particular to a method and system for detecting generator components based on an AI hybrid expert model. Background Art
[0002] Traditional generator maintenance relies primarily on manual visual inspection and periodic physical measurements for component inspection and condition assessment. These methods are not only time-consuming and labor-intensive but also susceptible to limitations in operator experience and judgment, leading to inconsistent and inaccurate inspection results. Advances in computer vision and deep learning technologies have enabled automated image recognition methods to improve inspection accuracy and efficiency. However, existing technologies still face challenges in processing the highly complex and variable images of generator components, particularly in identifying small or ambiguous defects.
[0003] The Chinese patent application "A Wind Turbine Image Recognition Fault Detection Method" (Application Number: 202311836704.8) discloses a wind turbine image recognition fault detection method. This method utilizes image recognition technology to detect wind turbine faults, eliminating the need for additional sensors. This improves detection efficiency and accuracy, reduces system complexity, and facilitates wind turbine condition monitoring and fault prediction. This approach focuses on utilizing image recognition technology and convolutional neural networks (CNNs) to detect wind turbine faults. However, these approaches suffer from the following drawbacks: 1) While CNNs are used for feature extraction and classification, the performance of CNN models relies on a large amount of labeled training data. Insufficient or unrepresentative training data can lead to model overfitting or poor generalization. This can make the model inadequately adaptable to unseen fault types or complex scenarios in practical applications, impacting the accuracy and robustness of fault detection. 2) This method integrates the operating principles and fault mechanisms of wind turbines for fault diagnosis, but the construction and updating of a knowledge base requires extensive expertise and experience. If the knowledge base is incomplete or not updated promptly, this can lead to inaccurate diagnosis of certain faults or the inability to diagnose new faults.
[0004] 3) Use fuzzy logic or neural networks to evaluate faults and determine their severity and urgency. However, the rule setting of fuzzy logic and the training results of neural networks may be subjective. Different experts or models may give different evaluation results, affecting the consistency of fault handling decisions.
[0005] The Chinese patent application "Control System and Method for Internal Inspection of Stator and Rotor of Generator Sets" (Application No. 202411770404.9) discloses a control system and method for internal inspection of stator and rotor of generator sets. In this solution, the main control system is designed so that after the internal module collects data, the processing module can compare and analyze the sounds collected by the internal module through spectrum comparison. Simultaneously, the processing module can compare and analyze the images collected by the internal module through similarity calculation. This process does not require extensive computation, only data comparison and processing, resulting in high processing speed. This allows for rapid analysis of stator and rotor problems, improving inspection efficiency and facilitating rapid repair of problematic stator and rotors. The core of this solution is to quickly determine whether there are problems with the stator and rotor through spectrum comparison and image similarity calculation. However, it still suffers from the following technical deficiencies: ①. This method uses spectrum comparison and image comparison to determine whether there are problems with the stator and rotor, but its effectiveness is highly dependent on the integrity and accuracy of the stored non-destructive data. If the stored data is biased, incomplete, or mismatched with the actual inspection environment, misjudgment may occur. ②. Although this solution describes a high processing speed, the complexity of data processing in actual applications may be underestimated. For example, steps like preprocessing, feature extraction, and comparative analysis of sound and image data require significant computing resources and time, especially when processing large amounts of data, which may not meet real-time requirements. (3) The system is designed for a specific type of generator set. For generators of different models and structures, extensive customization may be required to adapt it. This limits the system's versatility and potential for widespread adoption.
[0006] In summary, the existing technology for detecting generator components still has technical problems such as low accuracy, low efficiency, and non-real-time and incomplete detection. Summary of the Invention
[0007] To address the aforementioned technical issues, this paper provides a method and system for inspecting generator components based on an AI hybrid expert model. By integrating multiple deep learning algorithms and image processing techniques, this system achieves high-precision assessment of the status of generator stators and rotors. The system includes data acquisition, preprocessing, model training, deployment, and fine-tuning, aiming to improve inspection efficiency and accuracy through automated methods.
[0008] The technical solution adopted by the present invention is:
[0009] The method for detecting generator components based on an AI hybrid expert model includes the following steps:
[0010] Step 1: Use multiple real-time cameras to capture images of generator components;
[0011] Step 2: Preprocess the collected generator component image data, including but not limited to denoising, grayscale conversion, edge enhancement, etc. Step 3: Manually or semi-automatically label the preprocessed generator component image data to form a training dataset Step 4: Use a computer vision model with a hybrid expert model architecture to train the labeled training dataset;
[0012] Step 5: Expand and diversify the training dataset through label aggregation methods such as data augmentation and pseudo-labeling techniques;
[0013] Step 6: Adjust the hyperparameters of the mixture of experts model and perform feature engineering to improve the accuracy and generalization ability of the model;
[0014] Step 7: Retrain the adjusted hybrid expert model until it meets the preset performance indicators;
[0015] Step 8: Deploy the retrained hybrid expert model into the production environment for real-time generator component detection.
[0016] Step 9: Regularly fine-tune the hybrid expert model of the production environment to adapt to changes in the production environment and emerging generator component status characteristics.
[0017] In step 2, the preprocessing uses a deep learning model to perform image super-resolution to improve image quality so that the model can more accurately identify minor damage and defects.
[0018] Preprocessing includes:
[0019] 1) Denoising:
[0020] The adaptive filter dynamically adjusts its filtering parameters according to the local characteristics of the image, as follows:
[0021] Assume z xy is the pixel value of the image at position (x, y), and the output zx′y of the adaptive median filter is calculated by the following steps: xy Select a window W around it and calculate the minimum value z in W min , maximum value z max and median z med If z med Not equal to z min and z max , then go to the next step; otherwise, increase the window W and repeat the previous steps. xy Not equal to z min or z max , then z′ xy =z xy Otherwise, z′ xy =zmed .
[0022] 2) Contrast adjustment:
[0023] Let p r (r k ) is the histogram of the original image, where r k is the kth gray level,
[0024]
[0025] Among them, n k is the gray level r k The number of pixels, N is the total number of pixels in the image; the gray level after equalization s k Calculated by the cumulative distribution function CDF;
[0026]
[0027] Where L is the total number of possible gray levels.
[0028] 3) Edge detection:
[0029] Let's define the image's gradient strength G and its direction θ. The Canny edge detection algorithm consists of the following steps: ① Smoothing: Using a Gaussian filter to reduce image noise; ② Gradient Calculation: Applying the Sobel operator to the smoothed image to calculate the gradient strength G and direction θ; ③ Non-Maximum Suppression (NMS): Retaining only local maxima in the gradient direction to refine potential edges; and ④ Dual Threshold Detection: Determining true edges by setting high and low thresholds.
[0030] In the step 3: manually or semi-automatically label the pre-processed generator component image data to form a training data set; Figure 2 As shown in the present invention, after undergoing preprocessing operations such as edge enhancement, noise suppression, and resolution normalization, images of generator components are fed into the data annotation module. This module supports both manual annotation and semi-automatic annotation using prior models. Targets for annotation include, but are not limited to, key areas such as stator slots, rotor surface cracks, and carbon deposits. After review and correction by technical personnel, the annotated data is exported in a unified COCO or custom JSON format for training deep learning models. This process not only improves annotation efficiency but also enhances the consistency and accuracy of sample labels.
[0031] In step 4, the hybrid expert model includes multiple expert sub-models and a gating mechanism. The expert sub-models are responsible for processing data subsets or features of different generator components, and the gating mechanism is responsible for dynamically selecting the most appropriate expert sub-model based on the input data to improve the accuracy and efficiency of the overall model.
[0032] The hybrid expert model is a composite neural network architecture consisting of multiple expert networks (ExpertNetworks) and a gating network (GatingNetwork); each expert network is a VisionTransformer (ViT) designed to recognize and process specific types of features in images of generator components, including cracks, corrosion, wear, etc.
[0033] Assume there are N expert networks and one gating network. For a given input image x, the output of the i-th expert network is represented as y i (x), the weighted output of the gating network for the i-th expert is represented as g i (x), where i = 1, 2, ..., N. The final output Y(x) of the hybrid expert model can be obtained by weighted summation:
[0034]
[0035] Among them, the weight g i (x) is calculated through the gating network and satisfies This is usually achieved by applying the softmax function:
[0036]
[0037] Among them, z i (x) represents the original output of the gating network to the i-th expert network.
[0038] In generator component detection, the hybrid expert model can dynamically adjust the weights of each expert network and select the most appropriate expert network for processing based on the characteristics of the input data, that is, the image of the generator stator or rotor component.
[0039] In step 4, the hybrid expert model is trained as follows:
[0040] In supervised learning, each input example has a corresponding label or output value. For the image recognition task of the generator component, this means that each image is labeled as belonging to a specific category, and the goal of the mixture of experts model is to predict the label of unseen images by learning the relationship between the input image and the label.
[0041] 1): The cross entropy loss function measures the difference between the model prediction value and the true label. The mathematical expression of the cross entropy loss function is:
[0042]
[0043] Where N is the number of samples, y i is the true label of the i-th sample, is the model's predicted probability for the i-th sample. For multi-classification problems, the cross entropy loss can be expanded to:
[0044]
[0045] Where C is the number of categories; y ic Is a one-hot encoding vector, indicating that the i-th sample belongs to category c; is the probability that the model predicts that the i-th sample belongs to category c.
[0046] 2): Adopt Adam optimization algorithm to adaptively adjust the learning rate of each parameter:
[0047] The update rule of the Adam algorithm is:
[0048] m t =β1m t- 1+(1-β1g t )
[0049] v t =β2v t- 1+(1-β2)g t 2
[0050]
[0051] Where: g t is the gradient at time step t; m t and v t are estimates of the first and second moments of the gradient, respectively; β1 and β2 are decay rate parameters; η is the learning rate; ∈ is a small constant to prevent division by zero.
[0052] In step 5, label aggregation methods, such as data augmentation and pseudo-labeling techniques, are used to expand and diversify the training dataset. Data augmentation methods perform various transformations on labeled image samples, including image rotation, mirror flipping, random cropping, brightness adjustment, and noise superposition, thereby generating diverse image inputs while maintaining the labels. This method effectively improves the model's robustness to generator components under different operating and imaging conditions. Pseudo-labeling techniques are applied to unlabeled image data. The specific steps are as follows: first, a trained hybrid expert model is used to predict the unlabeled images, obtaining highly confident prediction results as "pseudo-labels"; then, the pseudo-labeled data is merged with the manually labeled data and retrained for the model. This method can significantly improve data utilization and optimize model performance in scenarios where sample labeling costs are high and real data is insufficient. By combining the two label aggregation methods mentioned above, the quantity and diversity of training data can be significantly increased while ensuring label quality, thereby enhancing the model's discriminative and generalization capabilities.
[0053] In step 6, the hyperparameters of the mixture of experts model are adjusted and feature engineering is performed as follows:
[0054] An autoencoder consists of two parts: an encoder and a decoder. The role of the encoder is to compress the input data into a lower-dimensional encoding, and the decoder reconstructs the original data from this encoding. By training to make the reconstructed output as close as possible to the input, the autoencoder can learn an effective low-dimensional representation of the data;
[0055] Let the input image be represented as x ∈ Rn, where n is the input dimension. The encoder maps x to a hidden representation z ∈ Rm, where m < n, and the mapping function can be expressed as z = f(x). Then, the decoder maps z back to the reconstructed image That is f(·) and g(·) represent the non-linear transformations of the encoder and decoder respectively, which are implemented through neural networks.
[0056] The goal of the autoencoder is to minimize the difference between the input x and the reconstructed output This difference is usually quantified by a loss function. Here, the mean squared error is adopted, as follows:
[0057]
[0058] In step 7, a probability model between the hyperparameters and the performance of the mixture of experts model is constructed through Bayesian optimization, and this probability model is used to predict the performance of untested hyperparameter combinations. At each step, Bayesian optimization will select a hyperparameter combination that is most likely to improve performance according to the current model for testing, and then update the mixture of experts model according to the test results, gradually narrowing the search scope until the optimal solution is found.
[0059] In step 8, the deployment process involves integrating the trained model into the production environment so that it can process real-time data and provide instant predictions. It includes the following steps:
[0060] a1: Model compression and quantization to reduce the model size and improve the inference speed, making the model more suitable for running on online or edge devices. a2: Package the model as an API or service that can be called by the production environment;
[0061] a3: Integrate the model service into the data processing flow to ensure that the model can receive real-time data and output prediction results;
[0062] a4: After deployment, it is necessary to continuously monitor the model performance and system health status in order to detect and solve problems in a timely manner.
[0063] In step 9, fine-tuning is the process of retraining the model after deployment based on its performance in real applications and newly collected data. The specific steps include:
[0064] b1: Regularly collect and annotate new data that the model processes in production, especially samples where the model makes incorrect predictions.
[0065] b1: Retrain or fine-tune the model using newly collected data to correct incorrect predictions and improve the model's accuracy and robustness. b3: Apply a progressive learning strategy that enables the model to gradually adapt to new data and new tasks without forgetting old knowledge.
[0066] A system for detecting generator components based on an AI hybrid expert model, the system comprising:
[0067] A sensing unit, comprising at least one real-time camera, for collecting image data of generator components in real time;
[0068] The data processing unit is used to perform tasks such as preprocessing of generator component image data, feature extraction, and hybrid expert model training. The data processing unit uses a high-performance computing platform that supports large-scale parallel processing and real-time data analysis to ensure efficient model training and evaluation.
[0069] The decision output unit outputs the current status of the generator components based on the evaluation results of the hybrid expert model, including but not limited to levels such as normal, damaged, and worn, providing decision support for maintenance and repair work.
[0070] One or more model fine-tuning units regularly evaluate the performance of the hybrid expert model deployed in the production environment and automatically adjust model parameters or retrain based on the evaluation results to continuously optimize model performance.
[0071] The decision output unit includes a user interface that can display the detection results in graphical or textual form and provide detailed information on the status of generator components and maintenance suggestions.
[0072] The present invention provides a method and system for detecting generator components based on an AI hybrid expert model, and the technical effects are as follows:
[0073] 1) The present invention improves the accuracy and efficiency of status detection of generator stator and rotor components, and uses advanced image processing and deep learning methods to automatically detect and evaluate minor defects, wear or damage in generator components.
[0074] 2) This invention combines advanced hybrid expert models, data preprocessing, model training strategies, and effective deployment and fine-tuning mechanisms to provide a novel solution for generator component detection and condition assessment. This solution not only improves detection accuracy and efficiency, but also significantly reduces maintenance costs through automated processing, improving the reliability and safety of generator operations.
[0075] 3) The present invention not only realizes efficient and automated detection of generator components, but also improves the accuracy and reliability of detection, significantly enhancing its applicability and convenience in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The present invention will be further described below with reference to the accompanying drawings and examples:
[0077] Figure 1 This is a flow chart of the method for detecting generator components based on the AI hybrid expert model of the present invention.
[0078] Figure 2 Flowchart for labeling image data of generator components.
[0079] Figure 3 This is a schematic diagram of an image acquisition device located between a stator and a rotor performing image acquisition operations;
[0080] Figure 3 In the figure, 1-generator rotor, 2-stator-rotor gap, 3-generator stator, 4-stator slot wedge detection device.
[0081] Figure 4 This is a schematic diagram of the position relationship between the image acquisition device and the stator after removing the rotor;
[0082] Figure 4 In the figure, 5 is an image acquisition device, and 6 is a stator slot wedge.
[0083] Figure 5 Schematic diagram of the structure of the image acquisition device;
[0084] Figure 5 In the figure, 5-01 is the camera, 5-02 is the depth camera, 5-03 is the driving electromagnet, and 5-04 is the guide rail slider assembly.
[0085] Figure 6 This is a scan of the generator stator surface.
[0086] Figure 7 Schematic diagram of the Bayesian optimization process. DETAILED DESCRIPTION
[0087] The present invention discloses a computer vision technology using a Mixture of Experts (MoE) model architecture, aiming to achieve a method and system for high-precision detection and status evaluation of generator stator and rotor components, such as Figure 1 This technology uses a combination of image recognition technology and is implemented through the following detailed steps:
[0088] 1) Image capture: Use high-resolution real-time camera equipment to capture images of target generator components.
[0089] 2) Data preprocessing: Perform necessary preprocessing operations on the captured image data, such as denoising, grayscale conversion, edge enhancement, etc., as well as data labeling to meet subsequent processing requirements.
[0090] 3) Model Training: The preprocessed and labeled data is fed into a hybrid expert model for training to learn the characteristics and potential defects of stator and rotor components. A label aggregation strategy is used to expand the training sample set to enhance the model's generalization ability.
[0091] 4) Model optimization: Refine model hyperparameter settings and perform feature engineering to optimize model performance, and then retrain the optimized model to ensure its accuracy and stability.
[0092] 5) Model deployment and fine-tuning: Deploy the tuned model in the production environment and perform regular fine-tuning to adapt to environmental changes and component wear.
[0093] The system architecture of the present invention consists of three major units: a perception unit, a data processing unit, and a decision-making output unit. The perception unit is responsible for real-time acquisition of image data from generator components; the data processing unit is responsible for performing data preprocessing, feature extraction, and real-time model adjustment; and the decision-making output unit outputs the current status of the generator components based on the model's evaluation results, providing a scientific basis for decision-making for subsequent maintenance and repair work. Through the above-mentioned technical approach, the present invention not only achieves efficient and automated detection of generator components, but also improves the accuracy and reliability of detection, significantly enhancing its applicability and convenience in practical applications.
[0094] (1) Data collection:
[0095] High-resolution real-time cameras shoot generator components from multiple angles and distances under different lighting conditions. Figure 3 、 Figure 4The images show the camera's location and installation method during the inspection, ensuring clear and comprehensive image data in all environments. This approach, based on the theory of multi-view geometry in computer vision, posits that observing the same object from different perspectives and distances can provide more information about its structure and surface properties, helping the model more accurately identify and assess the component's condition.
[0096] x=K[R|t]X
[0097] Among them, K is the internal parameter matrix, which contains information such as focal length and optical center; R and t are the rotation matrix and translation vector of the camera relative to the object, respectively, representing the external parameters.
[0098] (2) Pretreatment:
[0099] 1. Adaptive filter:
[0100] Image denoising is an important preprocessing step aimed at reducing noise introduced during image capture while preserving important image details. Adaptive filters dynamically adjust their filtering parameters based on the local characteristics of an image. For example, an adaptive median filter can more effectively remove salt and pepper noise while maintaining edge clarity.
[0101] Assume z xy is the pixel value of the image at position (x, y), the output z′ of the adaptive median filter xy It can be calculated by the following steps: xy Select a window W around it. Compute the minimum value z in W min , maximum value z max and median z med If z med Not equal to z min and z max , then go to the next step; otherwise, increase the window W and repeat step 2. xy Not equal to z min or z max , then z′ xy =z xy Otherwise, z′ xy =z med .
[0102] 2. Contrast adjustment:
[0103] Histogram equalization is a method for increasing image contrast by adjusting the image's histogram to make it more evenly distributed, thereby enhancing the overall image contrast. This is particularly useful for improving the visibility of generator components, such as the stator and rotor, in images, especially in conditions with uneven lighting or low contrast.
[0104] Let pr (r k ) is the histogram of the original image, where r k is the kth gray level:
[0105]
[0106] n k is the gray level r k The number of pixels, N is the total number of pixels in the image. The gray level s after equalization k It can be calculated using the cumulative distribution function CDF.
[0107]
[0108] Where L is the total number of possible gray levels.
[0109] 3. Edge detection:
[0110] The Canny edge detection algorithm is a popular edge detection method that aims to mark edges in an image as accurately as possible. It uses a multi-step process, including smoothing the image to remove noise, calculating the image gradient strength and direction, applying non-maximum suppression (NMS), and using a dual threshold algorithm to detect and connect edges. Let the gradient strength of the image be G and the gradient direction be θ. The key steps of the Canny algorithm can be summarized as follows: 1. Smoothing: Use a Gaussian filter to reduce image noise. 2. Gradient calculation: Apply the Sobel operator to the smoothed image to calculate the gradient strength G and direction θ. 3. Non-maximum suppression (NMS): Only local maxima are retained in the gradient direction to refine potential edges. 4. Dual threshold detection: Determine the true edge by setting high and low thresholds.
[0111] These preprocessing steps together improve the quality of image data and lay a solid foundation for subsequent model training and analysis. Denoising and contrast adjustment ensure image clarity, helping the model to more accurately identify and understand component features; while edge detection emphasizes the structural boundaries of components, which is crucial for identifying the shape and locating defects of generator components. Figure 5As shown, in the present invention, in order to more accurately identify geometric anomalies and positional offsets of generator components, a depth camera is introduced to collect three-dimensional data of the target component. The depth camera can record the distance information between each pixel and the camera, thereby generating a depth map or point cloud data containing spatial structural information. Compared with traditional RGB images, the depth map can reveal the changes in the shape characteristics of the object surface in three-dimensional space, such as tiny undulations, depressions, protrusions, cracks, etc. This enables the system to detect and evaluate the following defects: wear and deformation of stator slots and rotor edges, component misalignment caused by loose bolts, and structural deformation caused by thermal expansion and contraction. In addition, the geometric registration algorithm combining the depth map with multi-view images can achieve precise positioning of the component in space. If a significant deviation is detected between the component position and its standard working condition reference position, it can be determined as a "positioning defect." Therefore, the introduction of the depth camera not only improves the model's perception of surface details, but also significantly enhances the system's ability to detect component shape integrity and assembly accuracy.
[0112] Through such preprocessing, the accuracy and efficiency of subsequent model recognition can be greatly improved, thereby achieving more reliable generator component detection.
[0113] (3) Hybrid Expert Model:
[0114] The Mixture of Experts (MoE) model is a composite neural network architecture consisting of multiple expert networks (Expert Networks) and a gating network (Gating Network). This model aims to solve complex problems by combining the knowledge of multiple experts, where each expert focuses on learning a specific aspect of the input data, and the gating network is responsible for determining the weight of each expert based on the input data, and ultimately deciding which expert (or combination of experts) is most suitable for the current task. In the present invention, each expert network is a Vision Transformer (ViT), which is designed to identify and process specific types of features in images of generator parts, such as cracks, corrosion, wear, etc. Figure 6 shown.
[0115] Assume there are N expert networks and one gating network. For a given input image x, the output of the i-th expert network is represented as y i (x), the weighted output of the gating network for the i-th expert is represented as g i (x), where i = 1, 2, ..., N. Then the final output Y(x) of the hybrid expert model can be obtained by weighted summation:
[0116]
[0117] Among them, the weight g i (x) is calculated through the gating network and satisfies This is usually achieved by applying the softmax function:
[0118]
[0119] Here, z i (x) is the original output of the gating network for the i-th expert.
[0120] In generator component inspection, the MoE model can dynamically adjust the weights of each expert, selecting the most appropriate expert for processing based on the characteristics of the input data (such as an image of a generator stator or rotor component). For example, when an image shows crack characteristics, the model can rely more on an expert network specialized in crack detection; when an image shows corrosion or wear characteristics, the model relies more on the corresponding expert network. This improves the model's adaptability to different types of parts and damage. As new types of parts or damage emerge, more expert networks (such as a new ViT expert network to detect new damage types) can be added without having to train the entire model from scratch, thereby improving the model's scalability. By combining the knowledge of multiple experts, the MoE model can reduce the potential bias of a single model and improve overall prediction accuracy. Each expert network can process data in parallel, improving the model's computational efficiency.
[0121] Through this structural innovation, the hybrid expert model provides an efficient, accurate and scalable solution for handling complex image recognition tasks, especially in the field of generator component inspection (such as detecting cracks, corrosion and wear on stators and rotors).
[0122] (4) Training process:
[0123] In supervised learning, each input example has a corresponding label, or output value. For the generator component image recognition task, this means each image is labeled as belonging to a specific category (such as different types of generator component damage). The goal of the model is to predict the label for unseen images by learning the relationship between the input image and the label.
[0124] The cross entropy loss function is a common method to measure the difference between the model's predicted value and the true label, especially in classification tasks. For a binary classification problem, the mathematical expression of the cross entropy loss function is:
[0125]
[0126] Where N is the number of samples, y i is the true label of the i-th sample, Is the model's predicted probability for the i-th sample. For multi-classification problems, the cross entropy loss can be expanded to:
[0127]
[0128] Where C is the number of categories, y ic is a one-hot encoding vector, indicating that the i-th sample belongs to category c, is the probability that the model predicts that the i-th sample belongs to category c.
[0129] Adam optimization algorithm (AdaptiveMomentEstimation) is an adaptive learning rate optimization algorithm that combines the advantages of Momentum and RMSprop algorithms. Adam adjusts the learning rate of each parameter by calculating the first-order moment estimate (i.e., mean) and second-order moment estimate (i.e., uncentered variance) of the gradient;
[0130] The update rule of the Adam algorithm is:
[0131] m t =β1m t- 1+(1-β1g t )
[0132] v t =β2v t- 1+(1-β2)g t 2
[0133]
[0134] Among them, g t is the gradient at time step t, m t and v t are estimates of the first and second moments of the gradient, β1 and β2 are decay rate parameters, with recommended values of 0.9 and 0.999 respectively, η is the learning rate, and ∈ is a small constant to prevent division by zero, usually 1e-8.
[0135] The cross-entropy loss function directly measures the difference between the model's output probability distribution and the true distribution, making it particularly effective for classification problems. It provides explicit guidance on how to adjust model parameters to improve performance. The Adam optimization algorithm adaptively adjusts the learning rate of each parameter, accelerating training convergence while reducing parameter tuning effort. It combines the acceleration of Momentum with the adaptive learning rate feature of RMSprop, making it suitable for most deep learning tasks. Combining the cross-entropy loss function with the Adam optimization algorithm can effectively improve the training efficiency and accuracy of hybrid expert models in supervised learning environments, making it particularly well-suited for complex image recognition tasks such as generator component detection.
[0136] (V) Feature Engineering:
[0137] Autoencoder is an unsupervised learning algorithm whose main purpose is to learn the compressed representation of input data. In image processing, autoencoder can achieve image compression and abstract representation by learning the low-dimensional representation of the input image, while trying to retain the important information of the original image as much as possible. This mechanism makes autoencoder particularly suitable for tasks such as image denoising, feature extraction, and dimensionality reduction. Autoencoder consists of two parts: an encoder and a decoder. The role of the encoder is to compress the input data into a lower-dimensional encoding (abstract representation), and the decoder then attempts to reconstruct the original data from this encoding. By training to make the reconstructed output as close as possible to the input, autoencoder can learn an effective low-dimensional representation of the data.
[0138] Let the input image be represented as x ∈ Rn, where n is the input dimension. The encoder maps x to a hidden representation z ∈ Rm, where m < n, and the mapping function can be expressed as z = f(x). Then, the decoder maps z back to the reconstructed image That is f(·) and g(·) represent the non-linear transformations of the encoder and decoder respectively, usually implemented through neural networks.
[0139] The goal of Autoencoder is to minimize the difference between the input x and the reconstructed output This difference is usually quantified by a loss function, such as the mean squared error (MSE)
[0140]
[0141] Autoencoder can learn the compressed representation of input data. For images, this means that the same information can be stored with fewer bytes. By training autoencoder to reconstruct clean images from noisy inputs, it can learn the noise-removing representation, which is very effective for image denoising. The low-dimensional representation learned by the encoder can be regarded as the features of the input data, and these features are very useful for many machine learning and image processing tasks. By abstracting the representation of images, autoencoder can not only help reduce storage and computational costs, but also play an important role in multiple aspects such as image reconstruction, feature extraction, and denoising. Especially when dealing with large-scale image data, autoencoder provides an effective data compression and information extraction strategy.
[0142] (VI) Hyperparameter Tuning:
[0143] Bayesian Optimization is an optimization method based on Bayesian theory. Bayesian optimization constructs a probabilistic model between hyperparameters and model performance, usually a Gaussian process, and uses this model to predict the performance of untested hyperparameter combinations, thereby guiding the search process in a more promising direction. At each step, Bayesian optimization selects a hyperparameter combination that is most likely to improve performance based on the current model for testing, and then updates the model based on the test results, gradually narrowing the search range until the optimal solution is found. Through reasonable hyperparameter adjustment, the performance and efficiency of the model can be significantly improved. Figure 7 As shown in Figure 2. Bayesian optimization uses a Gaussian process to build a probabilistic model for the current observations. It then uses an acquisition function (expected improvement) to select the next optimal hyperparameter point for trial. The new trial results are then incorporated into the model, and the posterior distribution is updated. Through continuous iteration, the global optimal hyperparameter combination is gradually found. This method is more efficient than random search or grid search.
[0144] (7) Deployment and fine-tuning:
[0145] The deployment process involves integrating the trained model into the production environment so that it can process real-time data and provide instant predictions. This involves the following steps:
[0146] 1) Model compression and quantization to reduce model size and improve inference speed, making the model more suitable for running on online or edge devices;
[0147] 2) Encapsulate the model as an API or service that can be called by the production environment.
[0148] 3) Integrate model services into the data processing process to ensure that the model can receive real-time data and output prediction results.
[0149] 4) After deployment, model performance and system health need to be continuously monitored to detect and resolve problems in a timely manner.
[0150] Fine-tuning is the process of retraining a model after deployment based on its performance in real-world applications and newly collected data. This step is crucial for adapting to changes in the environment and emerging patterns. Specific steps include:
[0151] 1) Regularly collect and annotate new data that the model processes in production, especially samples where the model makes incorrect predictions.
[0152] 2) Retrain or fine-tune the model using newly collected data to correct incorrect predictions and improve the model’s accuracy and robustness.
[0153] 3) Apply progressive learning strategies to enable the model to gradually adapt to new data and new tasks without forgetting old knowledge.
[0154] This invention combines advanced hybrid expert models, data preprocessing, model training strategies, and effective deployment and fine-tuning mechanisms to provide a novel solution for the detection and condition assessment of generator components. This technology not only improves detection accuracy and efficiency, but also significantly reduces maintenance costs through automated processing, improving the reliability and safety of generator operation. In the future, with the continuous advancement of deep learning technology and the accumulation of experimental data, the detection capabilities of this system will be further enhanced. It can be widely applied to equipment maintenance and fault prediction in the power generation industry and other fields, demonstrating significant commercial value and social benefits.
Claims
1. A method for detecting generator components based on an AI hybrid expert model, characterized in that The following steps are involved: Step 1: Collect images of generator components; Step 2: Preprocessing the collected image data of generator components; Step 3: Annotate the preprocessed generator component image data to form a training data set; Step 4: Use the hybrid expert model to train the labeled training dataset; Step 5: Expand and diversify the training dataset through label aggregation methods; Step 6: Adjust the hyperparameters of the mixture of experts model and perform feature engineering; Step 7: Retrain the adjusted hybrid expert model until it meets the preset performance indicators.
2. The method for detecting generator components based on an AI hybrid expert model according to claim 1, characterized in that: The method also includes step 8: deploying the retrained hybrid expert model into the production environment for real-time generator component detection.
3. The method for detecting generator components based on an AI hybrid expert model according to claim 2, characterized in that: It also includes step 9: regularly fine-tuning the hybrid expert model of the production environment to adapt to changes in the production environment and newly emerging status characteristics of generator components.
4. The method for detecting generator components based on an AI hybrid expert model according to claim 1, characterized in that: In step 2, preprocessing uses a deep learning model to perform image super-resolution, and the preprocessing includes: 1) Denoising: The adaptive filter dynamically adjusts its filtering parameters according to the local characteristics of the image, as follows: Assume z xy is the pixel value of the image at position (x, y), the output of the adaptive median filter Calculated by the following steps: xy Select a window W around it and calculate the minimum value z in W min , maximum value z max and median z med ; if z med Not equal to z min and z max , then go to the next step; otherwise, increase the window W and repeat the previous steps; if z xy Not equal to z min or z max ,but otherwise, 2) Contrast adjustment: Let p r (r k ) is the histogram of the original image, where r k is the kth gray level, Among them, n k is the gray level r k The number of pixels, N is the total number of pixels in the image; the gray level after equalization s k Calculated by the cumulative distribution function CDF; Where L is the total number of possible gray levels; 3) Edge detection: Suppose the gradient strength of the image is G and the gradient direction is θ. The Canny edge detection algorithm includes the following steps: ①. Smoothing: Use a Gaussian filter to reduce image noise; ②. Gradient calculation: Apply the Sobel operator to the smoothed image to calculate the gradient strength G and direction θ; ③. Non-maximum suppression NMS: Only retain the local maximum in the gradient direction to refine the potential edge; ④. Dual threshold detection: Determine the true edge by setting high and low thresholds.
5. The method for detecting generator components based on an AI hybrid expert model according to claim 1, characterized in that: In step 4, the hybrid expert model includes multiple expert sub-models and a gating mechanism, the expert sub-models are responsible for processing data subsets or features of different generator components, and the gating mechanism is responsible for dynamically selecting the most appropriate expert sub-model based on the input data; The hybrid expert model is a composite neural network architecture consisting of multiple expert networks (ExpertNetworks) and a gating network (GatingNetwork); each expert network is a VisionTransformer (ViT) designed to recognize and process specific types of features in images of generator components, including cracks, corrosion, and wear; Suppose there are N expert networks and a gating network. For a given input image x, the output of the i-th expert network is represented as y i (x), the weighted output of the gating network for the i-th expert is represented as g i (x), where i = 1, 2, ..., N; the final output Y(x) of the hybrid expert model can be obtained by weighted summation: Among them, the weight g i (x) is calculated through the gating network and satisfies This is achieved by applying the softmax function: Among them, z i (x) represents the original output of the gating network to the i-th expert network; In generator component detection, the hybrid expert model can dynamically adjust the weights of each expert network and select the most appropriate expert network for processing based on the characteristics of the input data, that is, the image of the generator stator or rotor component.
6. The method for detecting generator components based on an AI hybrid expert model according to claim 5, characterized in that: In step 4, the hybrid expert model is trained as follows: In supervised learning, each input example has a corresponding label or output value; for the image recognition task of the generator component, this means that each image is labeled as belonging to a specific category, and the goal of the mixture of experts model is to predict the label of unseen images by learning the relationship between the input image and the label; 1): The cross entropy loss function measures the difference between the model prediction value and the true label. The mathematical expression of the cross entropy loss function is: Where N is the number of samples, y i is the true label of the i-th sample, is the model's predicted probability for the i-th sample; for multi-classification problems, the cross entropy loss can be expanded to: Where C is the number of categories; y ic Is a one-hot encoding vector, indicating that the i-th sample belongs to category c; is the probability that the model predicts that the i-th sample belongs to category c; 2): Adopt Adam optimization algorithm to adaptively adjust the learning rate of each parameter: The update rule of the Adam algorithm is: m t =β1m t-1 +(1-β1g t ) v t =β2v t-1 +(1-β2)g t 2 Where: g t is the gradient at time step t; m t and v t are estimates of the first and second moments of the gradient, respectively; β1 and β2 are decay rate parameters; η is the learning rate; ∈ is a small constant to prevent division by zero.
7. The method for detecting generator components based on an AI hybrid expert model according to claim 6, characterized in that: In step 6, the hyperparameters of the hybrid expert model are adjusted and feature engineering is performed, as follows: An autoencoder consists of two parts: an encoder and a decoder. The encoder compresses the input data into a lower-dimensional code, while the decoder reconstructs the original data from this code. By training the reconstructed output to be as close as possible to the input, the autoencoder can learn an effective low-dimensional representation of the data. Let the input image be represented as \(x\in\mathbb{R}^n\), where \(n\) is the input dimension; the encoder maps \(x\) to a hidden representation \(z\in\mathbb{R}^m\), where \(m < n\), and the mapping function is expressed as \(z = f(x)\); then, the decoder maps \(z\) back to the reconstructed image \(\hat{x}=g(z)\), where \(f(\cdot)\) and \(g(\cdot)\) respectively represent the non - linear transformations of the encoder and decoder, implemented through neural networks; n , where \(n\) is the input dimension; the encoder maps \(x\) to a hidden representation \(z\in\mathbb{R}^m\) m , where \(m < n\), and the mapping function is expressed as \(z = f(x)\); then, the decoder maps \(z\) back to the reconstructed image i.e., \(\hat{x}=g(z)\), where \(f(\cdot)\) and \(g(\cdot)\) respectively represent the non - linear transformations of the encoder and decoder, implemented through neural networks; The goal of the autoencoder is to minimize the input x and reconstruct the output This difference is usually quantified by the loss function, here the mean square error is used, as shown below:
8. The method for detecting generator components based on an AI hybrid expert model according to claim 2, characterized in that: In step 8, the deployment process involves integrating the trained model into the production environment so that it can process real-time data and provide immediate predictions. This includes the following steps: a1: Model compression and quantization to reduce model size and improve inference speed, making the model more suitable for running online or on edge devices a2: Encapsulate the model into an API or service that can be called by the production environment; a3: Integrate the model service into the data processing process to ensure that the model can receive real-time data and output prediction results; a4: After deployment, you need to continuously monitor model performance and system health to detect and resolve problems in a timely manner.
9. The method for detecting generator components based on an AI hybrid expert model according to claim 3, characterized in that: In step 9, fine-tuning is the process of retraining the model after deployment based on its performance in actual applications and newly collected data. The specific steps include: b1: Regularly collect and annotate new data processed by the model in the production environment, especially samples with incorrect model predictions; b1: Retrain or fine-tune the model using newly collected data to correct incorrect predictions and improve the model's accuracy and robustness; b3: Apply a progressive learning strategy to enable the model to gradually adapt to new data and new tasks without forgetting old knowledge.
10. A system for detecting generator components based on AI hybrid expert model, characterized by The system includes: A sensing unit, comprising at least one real-time camera, for collecting image data of generator components in real time; A data processing unit, used to perform preprocessing, feature extraction, and hybrid expert model training tasks on generator component image data; A decision output unit outputs the current status of the generator components according to the evaluation results of the hybrid expert model; One or more model fine-tuning units regularly evaluate the performance of the hybrid expert model deployed in the production environment and automatically adjust model parameters or retrain based on the evaluation results to continuously optimize model performance; The decision output unit includes a user interface that can display the detection results in graphical or textual form and provide detailed information on the status of generator components and maintenance suggestions.
Citation Information
Patent Citations
Wind generating set image identification fault detection method
CN118053111A
Generator set stator and rotor internal inspection device control system and control method
CN119556132A