Offshore wind turbine generator gearbox fault diagnosis method based on image recognition
Through multi-band combination light source and deep learning method, combined with polarization filtering and frequency domain filtering, the adversarial generation network and deformable convolutional layer are used to solve the problem of salt spray crystallization artifact interference in the gearbox of offshore wind turbine units, and efficient crack feature extraction and fault diagnosis are achieved.
Patent Information
- Application Number
- CN202510666744.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-22
AI Technical Summary
Existing image recognition methods are difficult to effectively distinguish between salt spray crystal artifacts and real cracks in the gearbox of offshore wind turbines, resulting in the loss of fault characteristic information, especially when the thickness of the salt crystal layer increases after the typhoon season, the diagnosis effect is poor.
The multi-band combination light source is used to combine polarization filtering and frequency domain filtering to generate networks and deformable convolutional layers. Through multi-physics coupled imaging optimization and deep learning architecture, crack features are extracted and salt crystal layer interference is suppressed, and cross-scale recognition is achieved by combining frequency domain and airspace feature analysis.
It significantly improves the reliability of crack diagnosis of offshore wind power gearboxes, reduces dependence on hardware transformation, can cope with complex working conditions of different salt spray concentrations and crystalline forms, and improves the accuracy and robustness of fault diagnosis.
Smart Images

Figure CN120375093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gearbox fault diagnosis, and particularly to a fault diagnosis method for a gearbox of an offshore wind turbine based on image recognition. Background Art
[0002] The gearbox of an offshore wind turbine is long-term exposed to a marine environment with high salt fog and high humidity. Salt fog particles continuously deposit on the gear surface and form a crystal layer, which has an impact on vision-based fault diagnosis. In recent years, with the breakthrough of deep learning technology, image recognition solutions based on high-resolution industrial cameras and convolutional neural networks have gradually become the mainstream, which achieve crack detection by capturing subtle changes in tooth surface textures. However, the complex artifacts such as dendritic or layered structures formed by salt fog crystallization on the tooth surface are highly similar to the morphological features of real cracks, resulting in difficulty for existing algorithms to effectively distinguish between the two in the feature extraction stage. Especially after the typhoon season, the thickness of the salt crystal layer can reach the millimeter level, further obscuring the true topography of the tooth surface and causing the loss of fault feature information.
[0003] To cope with salt fog interference, some technical routes focus on multi-modal data fusion and physical reinforcement learning directions. For example, multi-spectral imaging technology is adopted to analyze the reflectivity differences between salt crystals and metal substrates in different bands. Another team introduces polarized light imaging to suppress specular reflection noise by using the modulation characteristics of the crystal layer on the polarization state. However, in such solutions, multi-spectral imaging requires special optical hardware and is difficult to deploy inside the narrow gearbox cabin. Polarized imaging is sensitive to the thickness of the crystal layer, and the failure risk increases sharply when the crystal morphology is irregular. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] The present invention provides a fault diagnosis method for a gearbox of an offshore wind turbine based on image recognition to solve the problems that the crystallization artifacts on the gear surface caused by the offshore salt fog environment interfere with the existing image diagnosis methods, and there are limitations in artifact separation, cross-scale feature extraction, etc.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] An embodiment of the present invention provides a fault diagnosis method for a gearbox of an offshore wind turbine based on image recognition, which includes:
[0008] Step S1, collecting image data of the gear meshing surface through a multi-band combined light source and an industrial camera, where the multi-band combined light source includes a visible light band and a short-wave infrared band;
[0009] Step S2, preprocessing the image data, including dynamic interference removal based on polarization filtering and noise suppression based on frequency domain filtering;
[0010] Step S3: Input the preprocessed image into the crystallization artifact separation module, which includes an adversarial generation network and a deformable convolutional layer, to extract crack features from the salt spray crystallization covered area;
[0011] Step S4: Perform multi-scale fusion on the separated crack features through a spatio-temporal feature extraction network, which includes a frequency domain analysis branch and a spatial domain attention branch;
[0012] Step S5: Input the fused features into a classification model to output the fault type and severity;
[0013] Step S6: Generate a maintenance decision instruction according to the classification result and synchronously update the gearbox health status database.
[0014] As a preferred solution of the method for diagnosing faults of a gearbox of an offshore wind turbine based on image recognition according to the present invention, wherein: the visible light band is 450 - 680 nm, the short-wave infrared band is 900 - 1700 nm, the light source adopts an alternating pulse triggering mode, and the irradiation time ratio of visible light to short-wave infrared is 1:3.
[0015] As a preferred solution of the method for diagnosing faults of a gearbox of an offshore wind turbine based on image recognition according to the present invention, wherein: the preprocessing in step S2 specifically includes:
[0016] Obtain the linearly polarized reflection image of the gear surface through a polarization filtering device to suppress the specular reflection interference of the salt crystal layer;
[0017] Perform frequency domain filtering on the linearly polarized reflection image, adopt a band-pass filter with an adaptive bandwidth, and retain the frequency band related to crack features;
[0018] In step S2, the center frequency of the band-pass filter is dynamically adjusted according to the thickness of the salt crystal layer, and the thickness detection is realized by analyzing the change of the absorption rate of the short-wave infrared band image. The adjustment rule is: the center frequency of the band-pass filter shifts towards the high-frequency direction as the thickness of the salt crystal layer increases.
[0019] As a preferred solution of the method for diagnosing faults of a gearbox of an offshore wind turbine based on image recognition according to the present invention, wherein: the adversarial generation network of the crystallization artifact separation module adopts a dual-generator architecture:
[0020] The first generator is used to generate a synthetic image with salt spray crystallization artifacts, and the second generator is used to generate a tooth surface image without artifacts;
[0021] The deformable convolutional layer dynamically adjusts the shape of the convolutional kernel in the decoding stage to match the morphological changes of the crystallization artifacts;
[0022] In the double - generator architecture, the input of the first generator is the tooth surface image without crystallization artifacts and a random noise vector, and the output is a synthetic image with superimposed crystallization artifacts; the input of the second generator is the real image with crystallization artifacts, and the output is the tooth surface image after artifact separation; the two generators share a feature extraction layer and achieve adversarial training through a gradient reversal layer.
[0023] As a preferred solution of the method for diagnosing faults in a gearbox of an offshore wind turbine based on image recognition according to the present invention, in step S3, the pre - processed image is input into a crystallization artifact separation module, which is composed of an adversarial generation network and a deformable convolutional layer, and is used to extract crack features from the area covered by salt - fog crystallization, and is divided into the following two parts:
[0024] a) Dual - generator image mapping:
[0025] The first generator fuses the artifact - free tooth surface image with random noise to generate a synthetic image with salt - fog crystallization artifacts:
[0026]
[0027] Among them, represents the synthetic image with artifacts, G1(·,·;θ G1 ) represents the network mapping of the first generator, x represents the input artifact - free real tooth surface image, represents a random noise vector subject to a standard multivariate normal distribution, I is a k×k identity matrix, which is consistent with the dimension of the noise vector, θ G1 represents the network weights of the first generator, Enc(·) represents the shared encoder mapping, W z ∈R d×k represents a matrix that projects the k - dimensional noise z into a d - dimensional encoded feature space, k represents the dimension of the random noise vector, d represents the number of encoder output feature channels, R represents the real number field, and Dec1(·) represents the decoder mapping of the first generator;
[0028] In the formula,
[0029]
[0030] Among them, w i,j represents the weight that maps the j - th noise component to the i - th encoded feature dimension;
[0031] The second generator removes artifacts and restores the true texture of the tooth surface on the basis of the shared encoder features:
[0032]
[0033] Among them, represents the tooth surface image after artifact removal, G2(·;θG2 ) represents the second generator network mapping, y represents the input real artifact image, θ G2 represents the network weights of the second generator, and Dec2(·) represents the decoder mapping of the second generator;
[0034] b) Deformable convolution adaptive offset:
[0035] Introduce a deformable convolution layer in the decoding stage to adaptively learn the sampling offset according to the intermediate features to match the morphology of the crystallization artifacts:
[0036] Δp(p n ) = H off (F mid (p0 + p n )); θ off ),
[0037] F out (p0) = ∑ pn∈R w(p n )·F in (p0 + p n + Δp(p n ))),
[0038] where, Δp(p n ) represents the learned offset of the relative position p n , H off (·; θ off ) represents the offset generation network mapping, F mid represents the intermediate feature map of the decoder, p0 represents the center position of the current output pixel, p n ∈ R represents the set of relative positions in the standard convolution sampling grid, θ off represents the weights of the offset generation network, F out (p0) represents the value of the output feature map at the position p0, w(p n ) represents the weight coefficient of the convolution kernel at the relative position p n , F in represents the input feature map;
[0039] In the formula,
[0040] H off (F; θ off ) = σ(W off F + b off ),
[0041] where, here F is the intermediate feature vector extracted at F mid (p0 + p n ), represents the fully connected weight matrix of the offset network, Denote the bias vector of the offset network, σ denote the activation function, ReLU and LeakyReLU are selected, C denote the number of feature channels, N denote the number of sampling points, R is the standard sampling grid, and |R| = N.
[0042] As a preferred solution of the method for diagnosing faults in a gearbox of an offshore wind turbine based on image recognition according to the present invention, wherein: the frequency-domain analysis branch of the spatio-temporal feature extraction network performs the following operations:
[0043] Perform wavelet transform on the input image to extract frequency-domain features in multiple directions and at multiple scales;
[0044] Dynamically enhance the frequency-domain components corresponding to crack features through a learnable frequency band selection matrix.
[0045] As a preferred solution of the method for diagnosing faults in a gearbox of an offshore wind turbine based on image recognition according to the present invention, wherein: the spatial attention branch of the spatio-temporal feature extraction network includes:
[0046] An attention weight generation module based on the prior of gear stress distribution, which dynamically adjusts the weights of the feature map according to the position of the gear meshing point;
[0047] Cascaded Transformer encoders for capturing long-range spatial dependencies of crack features;
[0048] In the attention weight generation module, the stress distribution prior obtains the maximum principal stress distribution map during the gear meshing process through finite element simulation and encodes it into a spatial weight template; the module performs Hadamard product operation on the template and the real-time image feature map to generate attention weights.
[0049] As a preferred solution of the method for diagnosing faults in a gearbox of an offshore wind turbine based on image recognition according to the present invention, wherein: in the frequency-domain analysis branch of step S4, the frequency-domain coefficients after Fourier transform are weighted by a learnable frequency band selection matrix to dynamically highlight the frequency-domain components corresponding to cracks, including:
[0050] Perform two-dimensional discrete Fourier transform on the preprocessed artifact-free image X to obtain frequency-domain coefficients:
[0051]
[0052] wherein, X(p,q) represents the pixel value at the spatial coordinates (p,q), H and W respectively represent the height and width of the image, u = 0,..., H-1, v = 0,..., W-1 represent the frequency-domain coordinates, j represents the imaginary unit, and F(u,b) represents the frequency-domain coefficients;
[0053] Divide the frequency domain plane into N b non-overlapping frequency band sets And calculate the average amplitude of each frequency band:
[0054]
[0055] Among them, B i represents the coordinate set of the i-th frequency band, |B i | represents the set size, |F(u, v)| represents the amplitude of the frequency domain coefficient, and f i represents the feature quantity of the i-th frequency band;
[0056] Using the frequency band feature vector as the input, through the learnable matrix W m and the bias b m generate the frequency band weight vector:
[0057] w = softmax(W m f + b m ),
[0058] Among them, represents the learnable frequency band selection matrix, represents the learnable bias vector, N b represents the preset number of frequency band divisions, softmax(·) represents the component-wise normalization function, represents the weights of each frequency band;
[0059] The elements of the matrix W m are expressed as:
[0060]
[0061] Among them, w ij represents the mapping coefficient of the j-th frequency band feature to the i-th output weight;
[0062] For each frequency band, apply the weight w i to the corresponding frequency domain coefficient to enhance the crack signal:
[0063]
[0064] Among them, F en (u, v) represents the weighted frequency domain coefficient, and w i represents the weighting factor of the i-th frequency band;
[0065] Restore the weighted frequency domain coefficient to the spatial domain image for subsequent spatial domain branch processing:
[0066]
[0067] Among them, X en(p, q) represents the enhanced spatial domain pixel value.
[0068] As a preferred solution of a fault diagnosis method for an offshore wind turbine gearbox based on image recognition according to the present invention, wherein: the classification model fuses image features and vibration time series features, and the specific implementation method is as follows:
[0069] Construct a dual-channel input structure, where the first channel receives image spatio-temporal features, and the second channel receives the spectral features of the vibration signal collected synchronously;
[0070] Implement cross-modal feature interaction through a cross-attention mechanism to generate a joint representation vector.
[0071] As a preferred solution of a fault diagnosis method for an offshore wind turbine gearbox based on image recognition according to the present invention, wherein: the generation of the maintenance decision instruction includes:
[0072] Match the preset maintenance priority rules according to the fault type;
[0073] Combine offshore meteorological data and the operating status of the unit to dynamically adjust the maintenance time window;
[0074] Output a three-dimensional visualization report including fault location markings.
[0075] The beneficial effects of the present invention are as follows: through the imaging optimization of multi-physical field coupling and the innovation of the deep learning architecture, the reliability of crack diagnosis of offshore wind turbine gearboxes is significantly improved; the cooperation of polarization filtering and dynamic frequency domain filtering effectively suppresses the optical artifact interference of the salt crystal layer; the dual-generator adversarial network realizes the decoupling of artifact features and avoids the loss of crack information; the frequency domain branch captures the spectral energy distribution of cracks, and the spatial domain branch combines the stress distribution prior to strengthen the detection of key areas, solving the cross-scale recognition problem of micropitting and macroscopic tooth breakage; the adaptive frequency band selection and deformable convolution design enable the algorithm to cope with complex working conditions of different salt fog concentrations and crystal morphologies, reducing the dependence on hardware transformation. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0077] Figure 1 It is a schematic flow chart of a fault diagnosis method for an offshore wind turbine gearbox based on image recognition in Embodiment 1. DETAILED DESCRIPTION OF THE INVENTION
[0078] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.
[0079] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0080] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.
[0081] Example 1, referring to Figure 1 , this example provides a method for fault diagnosis of a gearbox of an offshore wind turbine based on image recognition, including the following steps:
[0082] Step S1, collect image data of the gear meshing surface through a multi-band combined light source and an industrial camera. The multi-band combined light source includes a visible light band and a short-wave infrared band;
[0083] The visible light band is 450 - 680 nm, and the short-wave infrared band is 900 - 1700 nm. The light source adopts an alternating pulse trigger mode, and the irradiation time ratio of visible light to short-wave infrared is 1:3. This ratio can balance the detail capture of visible light and the penetration requirement of infrared, and the incident angle of the short-wave infrared light source is controlled within the range of 30° ± 5°. This angle range can suppress specular reflection; here, 450 - 680 nm (visible light) and 900 - 1700 nm (short-wave infrared) belong to the industry-standard spectral segments;
[0084] Step S2, preprocess the image data, including dynamic interference removal based on polarization filtering and noise suppression based on frequency-domain filtering;
[0085] The preprocessing in Step S2 specifically includes:
[0086] Obtain the linearly polarized reflection image of the gear surface through a polarization filtering device to suppress the specular reflection interference of the salt crystal layer;
[0087] Perform frequency-domain filtering on the linearly polarized reflection image, using a band-pass filter with an adaptive bandwidth to retain the frequency band related to crack features;
[0088] In step S2, the center frequency of the band - pass filter is dynamically adjusted according to the thickness of the salt crystal layer. The thickness detection is achieved by analyzing the change in the absorption rate of the short - wave infrared band image. The adjustment rule is that the center frequency of the band - pass filter shifts towards the high - frequency direction as the thickness of the salt crystal layer increases;
[0089] Step S3: Input the pre - processed image into the crystal artifact separation module. This module contains an adversarial generation network and a deformable convolutional layer, and is used to extract crack features from the area covered by salt fog crystals;
[0090] The adversarial generation network of the crystal artifact separation module adopts a dual - generator architecture:
[0091] The first generator is used to generate a synthetic image with salt fog crystal artifacts, and the second generator is used to generate a tooth surface image without artifacts;
[0092] The deformable convolutional layer dynamically adjusts the shape of the convolutional kernel in the decoding stage to match the morphological changes of the crystal artifacts;
[0093] In the dual - generator architecture, the input of the first generator is a tooth surface image without crystal artifacts and a random noise vector, and the output is a synthetic image with superimposed crystal artifacts; the input of the second generator is a real image with crystal artifacts, and the output is a tooth surface image after artifact separation; the two generators share a feature extraction layer and achieve adversarial training through a gradient reversal layer;
[0094] In step S3, input the pre - processed image into the crystal artifact separation module. This module consists of an adversarial generation network and a deformable convolutional layer, and is used to extract crack features from the area covered by salt fog crystals, which is divided into the following two parts:
[0095] a) Dual - generator image mapping:
[0096] The first generator fuses the artifact - free tooth surface image and random noise to generate a synthetic image with salt fog crystal artifacts:
[0097]
[0098] Among them, represents the synthetic artifact - bearing image, G1(·,·;θ G1 ) represents the first generator network mapping, x represents the input artifact - free real tooth surface image, represents a random noise vector subject to a standard multivariate normal distribution, I is a k×k identity matrix, which is consistent with the dimension of the noise vector, θ G1 represents the network weights of the first generator, Enc(·) represents the shared encoder mapping, W z ∈R d×kA matrix that projects the k-dimensional noise z into the d-dimensional encoded feature space, where k represents the dimension of the random noise vector, d represents the number of encoder output feature channels, R represents the real number field, and Dec1(·) represents the decoder mapping of the first generator;
[0099] In the formula,
[0100]
[0101] Among them, w i,j represents the weight that maps the j-th noise component to the i-th encoded feature dimension;
[0102] The second generator removes artifacts and restores the true tooth surface texture on the basis of sharing encoder features:
[0103]
[0104] Among them, represents the tooth surface image after removing artifacts, G2(·; θ G2 ) represents the second generator network mapping, y represents the input real artifact-containing image, θ G2 represents the network weight of the second generator, and Dec2(·) represents the decoder mapping of the second generator;
[0105] b) Deformable convolution adaptive offset:
[0106] In the decoding stage, a deformable convolution layer is introduced to adaptively learn the sampling offset according to the intermediate features to match the crystal artifact morphology:
[0107] Δp(p n ) = H off (F mid (p0 + p n ) ; θ off ),
[0108]
[0109] Among them, Δp(p n ) represents the learned offset of the relative position p n , H off (·; θ off ) represents the offset generation network mapping, F mid represents the intermediate feature map of the decoder, p0 represents the center position of the current output pixel, p n ∈R represents the set of relative positions in the standard convolution sampling grid, θ off represents the weight of the offset generation network, F out (p0) represents the value of the output feature map at the position p0, and w(p n ) represents the convolution kernel at the relative position pn weight coefficient, F in represents the input feature map;
[0110] In the formula,
[0111] H off (F; θ off ) = σ(W off F + b off ),
[0112] where F here is the intermediate feature vector extracted at F mid (p0 + p n ). The full connection weight matrix of the offset network is represented by, The bias vector of the offset network is represented by, σ represents the activation function, ReLU and LeakyReLU are selected, C represents the number of feature channels, N represents the number of sampling points, R is the standard sampling grid, and |R| = N; Specifically, in this step, through the organic combination of the adversarial generation network and the deformable convolutional layer, the accurate generation and removal of salt spray crystallization artifacts are realized. The dual generator structure is within the same encoder feature space. Through noise projection and decoding mapping, it not only enhances the diversity of artifact samples but also realistically simulates the visual effects of the crystallization layer under different thicknesses and morphologies. Furthermore, it provides rich positive and negative sample pairs for adversarial training. The de-artifact generator focuses on restoring the tooth surface texture to ensure the detail reducibility after artifact removal. The deformable convolutional layer dynamically generates sampling offsets through the offset network and adaptively adjusts the convolutional kernel position in the decoding stage, enabling the network to flexibly respond to the changes in artifact edges and morphologies, greatly improving the response ability to micro-crack features. The overall module not only retains the key information of the cracks but also effectively suppresses the crystallization interference, laying a high-quality foundation for subsequent spatio-temporal feature extraction and significantly improving the accuracy and robustness of fault diagnosis;
[0113] In step S4, the separated crack features are subjected to multi-scale fusion through the spatio-temporal feature extraction network, and the spatio-temporal feature extraction network includes a frequency domain analysis branch and a spatial domain attention branch;
[0114] The frequency domain analysis branch of the spatio-temporal feature extraction network performs the following operations:
[0115] Perform a curvelet transform on the input image to extract multi-directional and multi-scale frequency domain features;
[0116] Dynamically enhance the frequency domain components corresponding to the crack features through a learnable frequency band selection matrix;
[0117] The spatial domain attention branch of the spatio-temporal feature extraction network includes:
[0118]
[0119] An attention weight generation module based on the prior of gear stress distribution dynamically adjusts the weights of the feature map according to the position of the gear meshing point;
[0120] A cascaded Transformer encoder is used to capture the long-range spatial dependencies of crack features;
[0121] In the attention weight generation module, the stress distribution prior obtains the maximum principal stress distribution map during the gear meshing process through finite element simulation and encodes it into a spatial weight template; the module performs a Hadamard product operation on the template and the real-time image feature map to generate attention weights;
[0122] In the frequency domain analysis branch of step S4, the frequency domain coefficients after Fourier transform are weighted by a learnable frequency band selection matrix to dynamically highlight the frequency domain components corresponding to cracks, including:
[0123] Perform a two-dimensional discrete Fourier transform on the preprocessed artifact-free image X to obtain the frequency domain coefficients:
[0124]
[0125] where X(p,q) represents the pixel value at the spatial coordinates (p,q), H and W represent the height and width of the image respectively, u = 0,..., H-1, v = 0,..., W-1 represent the frequency domain coordinates, j represents the imaginary unit, and F(u,v) represents the frequency domain coefficients;
[0126] Divide the frequency domain plane into N b non-overlapping frequency band sets and calculate the average amplitude of each frequency band:
[0127]
[0128] where B i represents the coordinate set of the i-th frequency band, |B i | represents the set size, |F(u,v)| represents the amplitude of the frequency domain coefficient, and f i represents the feature quantity of the i-th frequency band;
[0129] Taking the frequency band feature vector as the input, through the learnable matrix W m and the bias b m generate the frequency band weight vector:
[0130] w = softmax(W m f + b m ),
[0131] where, represents the learnable frequency band selection matrix, represents the learnable bias vector, Nb represents the preset number of frequency band divisions, and softmax(·) represents the component-wise normalization function. represents the weights of each frequency band;
[0132] The elements of matrix W m are expressed as:
[0133]
[0134] where w ij represents the mapping coefficient of the j-th frequency band feature to the i-th output weight;
[0135] For each frequency band, the weight w i is applied to the corresponding frequency domain coefficient to enhance the crack signal:
[0136]
[0137] where F en (u, v) represents the weighted frequency domain coefficient, and w i represents the weighting factor of the i-th frequency band;
[0138] The weighted frequency domain coefficients are restored to the spatial domain image for subsequent processing by the spatial domain branch:
[0139]
[0140] where X en (p, q) represents the enhanced spatial domain pixel value;
[0141] Specifically, this frequency domain branch dynamically weights different frequency band features through a learnable matrix, can adaptively highlight the spectral energy distribution corresponding to cracks. The frequency band division and average amplitude extraction steps compress the two-dimensional Fourier coefficients into a lower-dimensional feature vector that is easier to learn. The learnable frequency band selection matrix W m and the bias b m capture the unique spectral patterns of cracks during the training process, enabling the weight vector w to be automatically adjusted for different working conditions. The weighting operation of the frequency domain coefficients further enhances the signal-to-noise ratio and highlights the crack texture in the spatial domain. Finally, the inverse transform ensures that the image transmitted to the spatial domain branch not only retains the original details but also has higher contrast in the crack area;
[0142] Step S5, input the fused features into the classification model to output the fault type and severity;
[0143] The classification model fuses the image features and vibration time series features. The specific implementation method is as follows:
[0144] Construct a dual-channel input structure, where the first channel receives the spatio-temporal features of the image and the second channel receives the spectral features of the vibration signal collected synchronously;
[0145] Implement cross-modal feature interaction through a cross-attention mechanism to generate a joint representation vector;
[0146] Step S6, generate a maintenance decision instruction according to the classification result and synchronously update the gearbox health status database;
[0147] The generation of the maintenance decision instruction includes:
[0148] Match the preset maintenance priority rules according to the fault type;
[0149] Combine the offshore meteorological data and the unit operation status to dynamically adjust the maintenance time window;
[0150] Output a three-dimensional visualization report containing the fault location mark.
[0151] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A fault diagnosis method for the gearbox of an offshore wind turbine based on image recognition, characterized in that, Including, Step S1: Collect gear meshing surface image data through a multi-band combined light source and an industrial camera. The multi-band combined light source includes a visible light band and a short-wave infrared band; Step S2: Preprocess the image data, including dynamic interference removal based on polarization filtering and noise suppression based on frequency-domain filtering; Step S3: Input the preprocessed image into a crystallization artifact separation module, which includes an adversarial generation network and a deformable convolutional layer, for extracting crack features from the salt spray crystallization covered area; Step S4: Perform multi-scale fusion on the separated crack features through a spatio-temporal feature extraction network, which includes a frequency-domain analysis branch and a spatial domain attention branch; Step S5: Input the fused features into a classification model to output the fault type and severity; Step S6: Generate a maintenance decision instruction according to the classification result and synchronously update the gearbox health status database.
2. The fault diagnosis method for the gearbox of an offshore wind turbine based on image recognition according to claim 1, wherein The visible light band is 450 - 680nm, the short-wave infrared band is 900 - 1700nm, the light source adopts an alternating pulse trigger mode, and the irradiation time ratio of visible light to short-wave infrared is 1:
3.
3. The method for diagnosing faults in a gearbox of an offshore wind turbine based on image recognition according to claim 1, wherein The preprocessing in Step S2 specifically includes: Obtain the linearly polarized reflection image of the gear surface through a polarization filtering device to suppress the specular reflection interference of the salt crystal layer; Perform frequency-domain filtering on the linearly polarized reflection image, using a band-pass filter with an adaptive bandwidth to retain the frequency band related to crack features; In Step S2, the center frequency of the band-pass filter is dynamically adjusted according to the thickness of the salt crystal layer. The thickness detection is achieved by analyzing the change in the absorption rate of the short-wave infrared band image. The adjustment rule is: the center frequency of the band-pass filter shifts towards the high-frequency direction as the thickness of the salt crystal layer increases.
4. The fault diagnosis method for the gearbox of an offshore wind turbine based on image recognition according to claim 1, characterized in that, The adversarial generation network of the crystallization artifact separation module adopts a dual-generator architecture: The first generator is used to generate a synthetic image with salt spray crystallization artifacts, and the second generator is used to generate a tooth surface image without artifacts; The deformable convolutional layer dynamically adjusts the convolutional kernel shape in the decoding stage to match the morphological changes of the crystallization artifacts; In the dual-generator architecture, the input of the first generator is a tooth surface image without crystallization artifacts and a random noise vector, and the output is a synthetic image with superimposed crystallization artifacts; the input of the second generator is a real image with crystallization artifacts, and the output is a tooth surface image after artifact separation; the two generators share a feature extraction layer and perform adversarial training through a gradient reversal layer.
5. The method for diagnosing faults of an offshore wind turbine gearbox based on image recognition according to claim 4, wherein In Step S3, the preprocessed image is input into a crystallization artifact separation module, which consists of an adversarial generation network and a deformable convolutional layer, for extracting crack features from the salt spray crystallization covered area, and is divided into the following two parts: a) Dual-generator image mapping: The first generator fuses the artifact-free tooth surface image and random noise to generate a synthetic image with salt spray crystallization artifacts: Among them, represents the synthesized artifact-containing image, G1(·,·; θ G1 ) represents the mapping of the first generator network, x represents the input artifact-free real tooth surface image, represents a random noise vector subject to a standard multivariate normal distribution, I is the k×k identity matrix, consistent with the dimension of the noise vector, θ G1 represents the network weights of the first generator, Enc(·) represents the shared encoder mapping, W z ∈R d×k represents the matrix that projects the k-dimensional noise z into the d-dimensional encoded feature space, k represents the dimension of the random noise vector, d represents the number of encoder output feature channels, R represents the real number field, and Dec1(·) represents the decoder mapping of the first generator; where where, w i,j represents the weight for mapping the j-th noise component to the i-th encoded feature dimension; The second generator removes the artifacts based on the shared encoder features and restores the true texture of the tooth surface: Among them, represents the tooth surface image after removing artifacts, G2(·; θ G2 ) represents the mapping of the second generator network, y represents the input real artifact-containing image, θ G2 represents the network weights of the second generator, and Dec2(·) represents the decoder mapping of the second generator; b) Deformable convolution adaptive offset: Introduce a deformable convolutional layer in the decoding stage to adaptively learn the sampling offset according to the intermediate features to match the morphology of the crystallization artifacts: Δp(p n ) = H off (F mid (p0 + p n )); θ off ) where, Δp(p n ) represents the learning offset of the relative position p n , H off (·; θ off ) represents the offset generation network mapping, F mid represents the intermediate feature map of the decoder, p0 represents the center position of the current output pixel, p n ∈ R represents the set of relative positions in the standard convolutional sampling grid, θ off represents the weights of the offset generation network, F out (p0) represents the value of the output feature map at the position p0, w(p n ) represents the weight coefficient of the convolutional kernel at the relative position p n , F in represents the input feature map; where H off (F; θ off ) = σ(W off F + b off ), Among them, here F is F mid (p0 + p n ) is the intermediate feature vector extracted, represents the fully connected weight matrix of the offset network, represents the bias vector of the offset network, σ represents the activation function, ReLU and LeakyReLU are selected, C represents the number of feature channels, N represents the number of sampling points, R is the standard sampling grid, and |R| = N.
6. The fault diagnosis method for an offshore wind turbine gearbox based on image recognition according to claim 1, characterized in that The frequency-domain analysis branch of the spatio-temporal feature extraction network performs the following operations: Perform a curvelet transform on the input image to extract frequency-domain features in multiple directions and at multiple scales; Dynamically enhance the frequency-domain components corresponding to crack features through a learnable frequency band selection matrix.
7. The fault diagnosis method for the gearbox of an offshore wind turbine based on image recognition according to claim 6, wherein, The spatial attention branch of the spatio-temporal feature extraction network includes: An attention weight generation module based on the prior of gear stress distribution, which dynamically adjusts the weights of the feature map according to the position of the gear meshing point; Cascaded Transformer encoders for capturing the long-range spatial dependencies of crack features; In the attention weight generation module, the stress distribution prior obtains the maximum principal stress distribution map during the gear meshing process through finite element simulation and encodes it as a spatial weight template; The module performs a Hadamard product operation on the template and the real-time image feature map to generate attention weights.
8. The method for fault diagnosis of an offshore wind turbine gearbox based on image recognition according to claim 7, wherein In the frequency-domain analysis branch of step S4, the frequency-domain coefficients after Fourier transform are weighted by a learnable frequency band selection matrix to dynamically highlight the frequency-domain components corresponding to cracks, including: Perform a two-dimensional discrete Fourier transform on the preprocessed artifact-free image X to obtain frequency-domain coefficients: Where X(p,q) represents the pixel value at spatial coordinates (p,q), H and W represent the height and width of the image respectively, u = 0,…,H-1, v = 0,…,W-1 represent frequency-domain coordinates, j represents the imaginary unit, and F(u,v) represents the frequency-domain coefficients; Divide the frequency domain plane into N b non-overlapping frequency band sets and calculate the average amplitude of each frequency band: Among them, B i represents the coordinate set of the i-th frequency band, |B i | represents the set size, |F(u, v)| represents the amplitude of the frequency domain coefficient, and f i represents the characteristic quantity of the i-th frequency band; Using the band feature vector as the input, through the learnable matrix W m and the bias b m generate the band weight vector: w = softmax(W m f + b m ), Among them, represents a learnable band selection matrix, represents a learnable bias vector, N b represents the preset number of band partitions, and softmax(·) represents the component-wise normalization function, represents the weights of each band; Matrix W m The elements of which are represented as: where, w ij represents the mapping coefficient of the j-th band feature to the i-th output weight; For each frequency band, the weight w i is applied to the corresponding frequency domain coefficient to enhance the crack signal: Among them, F en (u, v) represents the weighted frequency-domain coefficient, and w i represents the weighting factor of the i-th frequency band; Restore the weighted frequency-domain coefficients to a spatial-domain image for subsequent processing by the spatial domain branch: Among them, X en (p, q) represents the enhanced spatial domain pixel value.
9. The fault diagnosis method for the gearbox of an offshore wind turbine based on image recognition according to claim 1, wherein, The classification model fuses image features and vibration time series features, and the specific implementation method is: Construct a dual-channel input structure, where the first channel receives spatio-temporal image features and the second channel receives the spectral features of the vibration signals collected synchronously; Implement cross-modal feature interaction through a cross-attention mechanism to generate a joint representation vector.
10. A fault diagnosis method for an offshore wind turbine gearbox based on image recognition according to claim 1, wherein, The generation of the maintenance decision instruction includes: Match the preset maintenance priority rules according to the fault type; Combine the offshore meteorological data and the operating status of the unit to dynamically adjust the maintenance time window; Output a three-dimensional visualization report containing the fault location mark.
Citation Information
Patent Citations
Gear box portable vibration monitoring device and monitoring method thereof
CN111982268A
Intelligent crack detection method based on wind turbine surface blurred image
CN112465776A
Hydraulic plunger pump intelligent fault diagnosis method based on pressure signals
CN114359663A
Small sample fault classification method and system for gearbox
CN116578907A
Gearbox fault diagnosis method based on time-frequency domain image and convolutional neural network
CN116796224A