Wind power blade damage intelligent identification method and system fused with multi-scale attention

By collecting wind turbine blade images using drones and performing multi-dimensional data enhancement, combined with the lightweight YOLO11 and SNMSDA-YOLO11 algorithm models with multi-scale attention mechanisms, the problems of inaccurate wind turbine blade damage identification and slow calculation speed are solved, and efficient online detection and identification are achieved.

CN120673279APending Publication Date: 2025-09-19BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510580671.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing intelligent identification methods for wind turbine blade damage have optimization problems such as inaccurate identification, slow calculation speed, and the need for a large number of samples.

Method used

An intelligent wind turbine blade damage recognition method integrating multi-scale attention is adopted. UAV images are collected and multi-dimensional data enhancement is performed. The SNMSDA-YOLO11 algorithm model is constructed, and the lightweight YOLO11 and multi-scale attention mechanism are combined for online detection and recognition.

Benefits of technology

The accuracy and calculation speed of damage detection are improved, making it suitable for large-scale real-time monitoring of wind turbine blades and overcoming the problem of limited sample size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673279A_ABST
    Figure CN120673279A_ABST
Patent Text Reader

Abstract

The invention discloses a wind turbine blade damage intelligent identification method and system fused with multi-scale attention, and the method comprises the steps: collecting an image of a wind turbine blade through an unmanned plane during inspection; performing multi-dimensional data enhancement on the acquired image; a fusion multi-scale attention mechanism is fused into the lightweight YOLO11, and an algorithm model of SNMSDA-YOLO11 is constructed; and the trained SNMSDA-YOLO11 algorithm model is utilized to carry out online detection and identification on the damage of the wind power blade. The calculation speed and efficiency are remarkably improved while the damage is accurately recognized, and the method has huge potential in application of large-scale real-time monitoring of wind turbine blade damage detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of damage identification technology, and specifically to a method and system for intelligently identifying wind turbine blade damage by integrating multi-scale attention. Background Art

[0002] Currently, many wind turbines have exceeded their designed service life and frequently experience failures, resulting in reduced power generation efficiency. This not only impacts the economic benefits of wind farms but also poses challenges to power dispatch. The high production costs and ongoing maintenance expenses of wind turbines are also significant. Blades, as core components of wind turbines, carry the highest cost and a high failure rate, making them a key factor affecting safe wind turbine operation. Monitoring blade operation and diagnosing faults present significant challenges.

[0003] Damage to wind turbine blades is mainly caused by two reasons. On the one hand, defects in the manufacturing process will lead to the failure to fully inspect the blades before production and delivery. During use, the damaged parts may further expand under high load and external environment, eventually leading to structural damage; on the other hand, climatic factors such as wind, sand, ice and snow cause serious damage to the surface of the blades, especially in offshore wind farms. Severe convective weather such as thunderstorms and typhoons and salt spray corrosion have aggravated the wear and damage of the blades. Surface defect detection of wind turbine blades is crucial in the operation and maintenance of wind turbines. How to achieve timely and accurate blade damage detection and identification is still a technical problem in the current field of wind power operation and maintenance:

[0004] (1) Although the wind farm has collected a large amount of wind turbine blade image data, most of them are normal image sample data, and there are very few fault samples. In addition, the types of fault samples are unbalanced, that is, some blades have more fault image samples, while some blades have fewer. The scarce blade fault image data brings challenges to model training.

[0005] (2) The scales of damage such as cracks, skin peeling, and glass fiber breakage on wind turbine blades vary, and there are damage forms of different scales. It is challenging to accurately and effectively extract multi-scale blade damage information.

[0006] (3) For wind turbines in actual wind farms, the existence of a large amount of wind turbine inspection image data taken by drones requires an algorithm model that can detect and identify wind turbine blade damage online and in real time. However, image-based damage detection and identification algorithm models usually have deep model network layers, complex parameters, and slow calculation speed. It is challenging to design an algorithm model that maintains a lightweight design while still having good damage detection and identification performance. Summary of the Invention

[0007] In view of the above-mentioned problems, the present invention is proposed.

[0008] Therefore, the technical problem solved by the present invention is that the existing intelligent identification method for wind turbine blade damage has optimization problems such as inaccurate identification, slow calculation speed, and the need for a large number of samples.

[0009] To solve the above technical problems, the present invention provides the following technical solution: a wind turbine blade damage intelligent identification method integrating multi-scale attention, comprising:

[0010] Use drones to collect images of wind turbine blades during inspections;

[0011] Perform multi-dimensional data enhancement on the collected images;

[0012] Integrate the multi-scale attention mechanism into the lightweight YOLO11 to build the SNMSDA-YOLO11 algorithm model;

[0013] The trained SNMSDA-YOLO11 algorithm model is used to detect and identify wind turbine blade damage online.

[0014] As a preferred solution of the wind turbine blade damage intelligent identification method integrating multi-scale attention described in the present invention, the multi-dimensional data enhancement is enhanced by adding random noise, sharpening the image and adjusting the image saturation.

[0015] As a preferred solution of the wind turbine blade damage intelligent identification method integrating multi-scale attention described in the present invention, wherein: the random noise enhancement includes simulating shooting distortion by adding Gaussian noise and salt and pepper noise;

[0016] I noisy1 (x,y)=I(x,y)+N(x,y)

[0017]

[0018] Where x represents the horizontal coordinate of the pixel in the image; y represents the vertical coordinate of the pixel in the image; I(x,y) represents the original pixel value at the coordinate (x,y); N(x,y) represents the Gaussian noise at the coordinate (x,y), N(x,y)~N(μ,σ 2 ) means that N(x,y) has a mean of μ and a variance of σ 2 Gaussian distribution with variance σ 2 Randomly select in the range [10,50]; I noisy1 (x,y) represents the pixel value at the coordinate (x,y) after Gaussian noise processing; I noisy2 (x,y) represents the pixel value at coordinate (x,y) after salt and pepper noise processing; P salt +P pepper =P amount , P amount∈[0.01,0.1] represents the noise ratio; P salt represents the probability of salt noise appearing in salt and pepper noise; P pepper Represents the probability of pepper noise appearing in salt and pepper noise.

[0019] As a preferred solution of the wind turbine blade damage intelligent identification method integrating multi-scale attention of the present invention, the image sharpening includes highlighting the image edges and details using the Laplacian operator;

[0020]

[0021] Among them, K is the convolution kernel, defined as:

[0022]

[0023] Among them, I sharpened (x,y) represents the pixel value at the coordinate (x,y) after sharpening; K(i,j) represents the value of the coordinate position (i,j) in the convolution kernel matrix K; j represents the column index in the convolution kernel matrix, and i represents the row index in the convolution kernel matrix.

[0024] As a preferred solution of the wind turbine blade damage intelligent identification method integrating multi-scale attention according to the present invention, wherein: the saturation adjustment includes adjusting the saturation in the HSV color space;

[0025] S'(x,y)=clip(S(x,y)·α,0,255)

[0026] Where S(x,y) represents the original saturation value; S'(x,y) represents the adjusted saturation value, α∈[0.8,1.2] represents the randomly generated scaling factor; clip represents the constraint function to ensure that the result is within the valid range of [0,255].

[0027] After completing the adjustments, convert the adjusted image back to the BGR color space.

[0028] As a preferred solution of the wind turbine blade damage intelligent identification method integrating multi-scale attention described in the present invention, the lightweight YOLO11 includes integrating the SlimNeck structure into YOLO11, using the SlimNeck structure in the neck of YOLO11, replacing the C3K2 module in the neck of YOLO11 with the existing VoVGSCSP module, and replacing the conv module in the neck of YOLO11 with the existing GSConv module;

[0029] The SlimNeck structure is mainly composed of a GSConv module and a VoVGSCSP module.

[0030] As a preferred solution of the wind turbine blade damage intelligent identification method integrating multi-scale attention described in the present invention, the multi-scale attention mechanism adopts a multi-head design to divide the channels of the feature map into n different heads, and different heads use different void rates to perform the sliding window void attention operation;

[0031] Different heads use different dilation rates to perform sliding window void attention, which sparsely selects keys and values ​​within a sliding window around the query block and performs self-attention operations;

[0032] For a given position (i, j), the hole rate r is defined to control sparsity; the query in the input feature map is sparsely selected in a sliding window of size w×w, and the self-attention operation is performed; the output feature map X of the sliding window hole attention operation corresponds to the component x ij Defined as:

[0033]

[0034] 1≤i≤W

[0035] 1≤j≤H

[0036] Among them, x ij is the component at coordinate (i, j) in the output feature map X of the sliding window hole attention operation; H and W are the height and width of the feature map respectively, and K r and V r Represents the key and value selected from K and V according to the void rate r, d k is the dimension of K;

[0037] Given a query at (i, j), (i', j') represents the coordinates corresponding to the key and value selected according to the void rate r within the sliding window centered at the query location (i, j). The specific coordinates are determined by the formula:

[0038] {(i',j')|i'=i+p×r,j'=j+q×r}

[0039]

[0040] The key and value at position (i',j') will be used to match the query Q ij Perform self-attention calculation to obtain the component x of the output feature map at the (i, j) position ij ;

[0041] For the feature map X, the query, key, and value corresponding to the feature map X are obtained through linear projection; the channel of the feature map is divided into n different heads, and a multi-scale sliding window hole attention operation is performed in each head with different hole rates, which can be expressed as:

[0042] h k =SWDA(Q k ,K k ,V k ,r k )

[0043] 1≤k≤n

[0044] X=Lincar(Concat[h1,…,h n ])

[0045] Among them, h i represents the features output by the k-th head after the sliding window hole attention operation; r k is the hole rate of the kth head, Q k , K k and V k represents the feature map slice input to the kth head; n represents the number of heads; Concat is the concatenation operation, which concatenates the outputs h1,...,h n Spliced ​​together along the channel dimension; Lincar is a linear transformation operation implemented by a fully connected layer;

[0046] The fusion multi-scale attention module adopts a multi-head design, which divides the channels of the feature map into n different heads. Each head uses a different hole rate to perform a sliding window hole attention operation, thereby achieving multi-scale feature extraction; the output are connected together and fed into the linear layer for feature aggregation;

[0047] X represents the final feature map after the multi-head attention mechanism is processed. It is the result obtained by concatenating and linearly transforming the outputs of each head, integrating the multi-scale feature information learned by different heads;

[0048] The SNMSDA-YOLO11 algorithm model includes, on the basis of the SN-YOLO11 model, embedding a multi-head attention mechanism when performing preliminary processing on the input features after the C2PSA module receives the input features.

[0049] An intelligent wind turbine blade damage identification system integrating multi-scale attention using the method of the present invention is characterized by:

[0050] The acquisition unit collects images of wind turbine blades through drones during inspections;

[0051] A processing unit performs multi-dimensional data enhancement on the collected images;

[0052] The recognition unit integrates the multi-scale attention mechanism into the lightweight YOLO11 to build the SNMSDA-YOLO11 algorithm model; the trained SNMSDA-YOLO11 algorithm model is used to detect and identify wind turbine blade damage online.

[0053] A computer device comprises: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, the steps of any one of the methods of the present invention are implemented.

[0054] A computer-readable storage medium stores a computer program, wherein: when the computer program is executed by a processor, the steps of any one of the methods of the present invention are implemented.

[0055] Beneficial effects of the present invention: The intelligent wind turbine blade damage identification method provided by the present invention, which integrates multi-scale attention, has many groundbreaking advantages. First, by enhancing the diversified data of limited samples, the problem of the limited number of wind turbine damage samples is overcome. Second, the innovative introduction of a multi-scale attention mechanism improves detection performance. Finally, a thin-neck structure is used to optimize the model, which significantly improves the computing speed and efficiency while ensuring accurate damage identification. It has great potential in the application of large-scale real-time monitoring of wind turbine blade damage detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0057] Figure 1 This is an overall flow chart of the wind turbine blade damage intelligent identification method integrating multi-scale attention provided by the first embodiment of the present invention;

[0058] Figure 2 The first embodiment of the present invention provides a method for intelligently identifying wind turbine blade damage by integrating multi-scale attention, and integrates a lightweight YOLO11 architecture with SlimNeck. DETAILED DESCRIPTION

[0059] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0060] Example 1, with reference to Figure 1 、 Figure 2 , which is an embodiment of the present invention, provides a wind turbine blade damage intelligent identification method integrating multi-scale attention, including:

[0061] S1: Collect images of wind turbine blades during inspection using drones.

[0062] S2: Perform multi-dimensional data enhancement on the collected images.

[0063] Furthermore, wind turbine blades operate in complex environments, including strong sunlight, shadows, dust, and other atmospheric factors, which can degrade the quality of images captured by unmanned aerial vehicles (UAVs). These challenges hinder accurate blade damage detection. To simulate real-world conditions and improve the robustness of the detection model, three image enhancement methods were used: random noise enhancement, image sharpening enhancement, and saturation adjustment enhancement.

[0064] (1) Random noise enhancement simulates the distortion introduced by environmental interference (such as particles in the air, vibration) or compression artifacts during the image capture process of the drone. The specific random noise is as follows:

[0065] Filming distortion is simulated by sequentially adding Gaussian noise and then salt and pepper noise.

[0066] I noisy1 (x,y)=I(x,y)+N(x,y)

[0067]

[0068] Among them, I noisy1 (x,y) represents the pixel value at the coordinate (x,y) after Gaussian noise processing; x represents the horizontal coordinate of the pixel in the image, and y represents the vertical coordinate of the pixel in the image; I(x,y) represents the original pixel value at the coordinate (x,y); N(x,y) represents the Gaussian noise at the coordinate (x,y), and N(x,y) to N(μ,σ 2 ) means that N(x,y) has a mean of μ and a variance of σ 2 Gaussian distribution; μ represents the mean, and variance σ 2 Randomly select in the range [10,50]. I noisy2(x,y) represents the pixel value at coordinate (x,y) after salt and pepper noise processing; P salt +P pepper =P amount , P amount ∈[0.01,0.1] represents the noise ratio, P salt Represents the probability of "salt" noise occurring, P pepper Indicates the probability of "pepper" noise occurring.

[0069] (2) During drone-based inspections, motion blur or atmospheric interference may obscure key features of wind turbine blade damage, such as cracks, delamination, or scratches. To highlight these details, image sharpening is applied as follows:

[0070]

[0071] Among them, K is the convolution kernel, defined as:

[0072]

[0073] Among them, I sharpened (x,y) represents the pixel value at the coordinate (x,y) after sharpening; K(i,j) represents the value of the coordinate position (i,j) in the convolution kernel matrix K; j represents the column index in the convolution kernel matrix, and i represents the row index in the convolution kernel matrix.

[0074] (3) Under different weather and lighting conditions, the color characteristics of the leaf damage area may change significantly in the drone image. In order to enhance the robustness of the model in capturing damage characteristics under different lighting conditions, saturation adjustment is used. This operation is performed in the HSV color space and is defined as follows:

[0075] S'(x,y)=clip)S(x,y)·α,0,255)

[0076] Where S(x,y) represents the original saturation value; S'(x,y) represents the adjusted saturation value, α∈[0.8,1.2] represents a randomly generated scaling factor; clip represents a constraint function that ensures that the result is within the valid range of [0,255]; after the adjustment is completed, the adjusted image is converted back to the BGR color space.

[0077] S3: Integrate the multi-scale attention mechanism into the lightweight YOLO11 to build the SNMSDA-YOLO11 algorithm model. This lightweight YOLO model improves the algorithm's computational capabilities. The multi-scale attention mechanism is introduced into the lightweight YOLO model to enable it to adaptively extract damage feature information at different scales for wind turbine blades. The details are as follows:

[0078] (1) Lightweight YOLO model

[0079] Figure 2 This is a lightweight YOLO11 architecture integrated with SlimNeck. The lightweight SN-YOLO11 model uses the YOLO11 framework as its foundation and incorporates the existing SlimNeck architecture. The SlimNeck architecture primarily consists of the GSConv module and the VoVGSCSP module. These modules work together to significantly improve inference speed and computational efficiency while maintaining object detection accuracy. The SlimNeck architecture is used in the YOLO11 neck, replacing the C3K2 module in the YOLO11 neck with the existing VoVGSCSP module, and the conv module in the YOLO11 neck with the existing GSConv module.

[0080] GSConv is the core module of the SlimNeck architecture. It combines standard convolution (SC) and depthwise separable convolution (DSC) and uses the shuffle operation to fuse features, maximizing the preservation of effective connections between channels. This approach reduces computational cost while avoiding the information loss common in traditional sparse convolution. Specifically, GSConv first downsamples the input using standard convolution, then extracts features using depthwise separable convolution, and finally mixes the feature map information through the shuffle operation, thereby improving inference speed and feature representation capabilities.

[0081] The GSbottleneck module further expands upon GSConv, introducing a cross-level local network structure to optimize computational efficiency. By combining standard convolution with depthwise convolution, GSbottleneck effectively reduces the number of parameters and computational complexity. Building on this foundation, the VoV-GSCSP module employs a single aggregation technique.

[0082] The YOLO11-SlimNeck architecture combines the advantages of the SlimNeck structure, enabling the model to not only achieve high accuracy in object detection tasks but also significantly improve inference speed and computational efficiency. By integrating the SlimNeck structure into YOLO11 and combining it with the GSConv, GSBottleneck, and VoV-GSCSP modules, the network's ability to handle high-dimensional data is further enhanced.

[0083] (2) Introducing a multi-scale attention mechanism into the lightweight YOLO model

[0084] A multi-scale attention mechanism is used to extract leaf damage feature information at different scales. Specifically, it uses sliding window void attention (SWDA), in which keys and values ​​are sparsely selected in a sliding window centered on the query block. First, the input features are mapped into a three-dimensional matrix consisting of query (Q), keys (K), and values ​​(V). SWDA sparsely selects keys and values ​​within a sliding window around the query block, and then performs a self-attention operation on these. This method not only expands the receptive field of window attention, but also increases the computational cost.

[0085] MSDA uses a multi-head design, splitting the channels of the feature map into n different heads. Different heads use different dilation rates to perform SWDA. Given a feature map X, the corresponding query, key, and value of the feature map X are obtained through linear projection. The channels of the feature map are then split into n different heads, and multi-scale SWDA is performed in each head using a different dilation rate. Specifically:

[0086] For a given position (i, j), the hole rate r is defined to control sparsity; the query in the input feature map is sparsely selected in a sliding window of size w×w, and the self-attention operation is performed; the output feature map X of the sliding window hole attention operation corresponds to the component x ij Defined as:

[0087]

[0088] 1≤i≤W

[0089] 1≤j≤H

[0090] Among them, x ij is the component at coordinate (i, j) in the output feature map X of the sliding window hole attention operation; H and W are the height and width of the feature map respectively, and K r and V r Represents the key and value selected from K and V according to the void rate r, d k is the dimension of K; the dilation rate r is a parameter used to control the sparsity when selecting keys and values ​​within the sliding window. In traditional convolution or attention mechanisms, adjacent elements are usually selected continuously. In dilation attention, by setting the dilation rate, keys and values ​​can be selected at intervals of r within the sliding window. For example, when r = 1, it is a normal continuous selection, similar to traditional convolution; when r = 2, one element is skipped for selection, which can expand the receptive field and allow the model to capture a wider range of contextual information. In the MSDA module, different heads can use different dilation rates to achieve multi-scale feature extraction.

[0091] Given a query at (i, j), (i', j') represents the coordinates corresponding to the key and value selected according to the void rate r within the sliding window centered at the query location (i, j). The specific coordinates are determined by the formula:

[0092] {(i',j')|i'=i+p×r,j'=j+q×r}

[0093]

[0094] The key and value at position (i',j') will be used to match the query Q ij Perform self-attention calculation to obtain the component x of the output feature map at the (i, j) position ij .

[0095] For the feature map X, the query, key, and value corresponding to the feature map X are obtained through linear projection; the channel of the feature map is divided into n different heads, and a multi-scale sliding window hole attention operation is performed in each head with different hole rates, which can be expressed as:

[0096] h k =SWDA(Q k ,K k ,V k ,r k )

[0097] 1≤k≤n

[0098] X=Lincar(Concat[h1,…,h n ])

[0099] Among them, h i represents the features output by the k-th head after the sliding window hole attention operation; r k is the hole rate of the kth head, Q k , K k and V k represents the feature map slice input to the kth head; n represents the number of heads; Concat is the concatenation operation, which concatenates the outputs h1,...,h n Concatenate together along the channel dimension; Lincar is a linear transformation operation implemented by a fully connected layer.

[0100] The fusion multi-scale attention module adopts a multi-head design, which divides the channels of the feature map into n different heads. Each head uses a different hole rate to perform a sliding window hole attention operation, thereby achieving multi-scale feature extraction; the output are connected together and fed into the linear layer for feature aggregation.

[0101] X represents the final feature map after processing by the multi-head attention mechanism. It is the result obtained by concatenating and linearly transforming the outputs of each head, integrating the multi-scale feature information learned by different heads.

[0102] The SNMSDA-YOLO11 algorithm model includes, on the basis of the SN-YOLO11 model, embedding a multi-head attention mechanism when performing preliminary processing on the input features after the C2PSA module receives the input features.

[0103] S4: Use the trained SNMSDA-YOLO11 algorithm model to detect and identify wind turbine blade damage online.

[0104] For the lightweight YOLO algorithm model that integrates the multi-scale attention mechanism, the model is trained using wind turbine blade damage data enhanced by multi-dimensional data, and the model parameters are reversely optimized and adjusted based on evaluation indicators such as accuracy, ultimately constructing a high-precision online wind turbine blade damage detection and identification method.

[0105] On the other hand, this embodiment also provides a wind turbine blade damage intelligent identification system integrating multi-scale attention, which includes:

[0106] The acquisition unit collects images of wind turbine blades through drones during inspections.

[0107] The processing unit performs multi-dimensional data enhancement on the collected images.

[0108] The recognition unit integrates the multi-scale attention mechanism into the lightweight YOLO11 to build the SNMSDA-YOLO11 algorithm model; the trained SNMSDA-YOLO11 algorithm model is used to detect and identify wind turbine blade damage online.

[0109] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0110] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0111] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0112] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0113] Example 2 is an embodiment of the present invention, which provides an intelligent identification method for wind turbine blade damage integrating multi-scale attention. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0114] Using images of blade damage in an actual wind turbine operating environment as data, the method proposed in this invention is used to identify and detect wind turbine damage categories, specifically including the following steps:

[0115] The first step is data increment:

[0116] We collected 1,534 images of wind turbine blades captured by drones in actual wind turbine operating environments. Among these images, 85 included images of damaged blades, primarily including skin loss, fiberglass breakage, cracks, and pockmarks. To expand the data size and improve its diversity, we first performed three rotation operations at different angles on the original images, preprocessing the dataset to 510 images.

[0117] Perform multi-dimensional data enhancement on the collected images:

[0118] Random noise enhancement: For each image, Gaussian noise is added with a probability of 0.5, with a mean of 0 and a variance randomly selected between 0.001 and 0.005. Salt and pepper noise is also added with a probability of 0.3, with a noise density randomly selected between 0.01 and 0.05. For example, adding Gaussian noise with a mean of 0 and a variance of 0.003 to an image of a normal leaf will blur the image and cause fine noise artifacts, simulating the distortion introduced by environmental interference or compression artifacts when shooting with a drone.

[0119] (2) Image sharpening: Use the Laplace operator to sharpen the image. By calculating the Laplace transform of the image, the edges and details in the image are highlighted. For an image of a blade with a fine crack, the crack edge becomes clearer after sharpening, and the originally blurred crack features become more distinct, which helps the subsequent model extract damage features.

[0120] (3) Saturation adjustment enhancement: In the HSV color space, the image's saturation is adjusted with a probability of 0.6. The saturation adjustment factor is randomly selected between 0.8 and 1.2. For example, in an image of leaf damage taken on a cloudy day, by adjusting the saturation to 1.1, the color of the damaged area becomes more vivid, and the damage characteristics can be better captured by the model under different lighting conditions.

[0121] The data was augmented using the three data augmentation methods above, expanding to 4440 images. Here are the results of the three types of data augmentation, using fiberglass breakage as an example.

[0122] The second step is to build the model:

[0123] Lightweight YOLO model construction:

[0124] First, the GSConv module is constructed. When processing the input image, it first uses a 3×3 standard convolution for downsampling, reducing the input image size by half while increasing the number of channels. Depthwise separable convolution is then used to extract features from the downsampled feature map. Depthwise separable convolution is divided into depthwise convolution and pointwise convolution. Depthwise convolution performs convolution operations on each channel independently, while pointwise convolution is used to fuse channel information. Finally, a shuffle operation is used to shuffle the feature map, recombining features from different channels to improve inference speed and feature representation capabilities.

[0125] The GSBottleneck module is built on GSConv, introducing a cross-level local network structure. For example, in a specific network layer, the GSBottleneck module combines 1×1 standard convolution with 3×3 depthwise convolution to process the feature map output by the previous layer, effectively reducing the number of parameters and computational complexity while maintaining good feature extraction capabilities.

[0126] Construct the VoV-GSCSP module and use single aggregation technology to aggregate the outputs of multiple GSBottleneck modules to further optimize the network structure.

[0127] The SlimNeck structure is integrated into YOLO11, and combined with the GSConv, GSBottleneck and VoV-GSCSP modules to build a lightweight YOLO11 model, which has high accuracy in target detection tasks while improving inference speed and computational efficiency.

[0128] Introduction of multi-scale attention mechanism:

[0129] Using sliding window void attention (SWDA), the input features are mapped into a three-dimensional matrix to obtain the query (Q), key (K), and value (V). For a leaf image feature map, when performing SWDA, the sliding window size is set, and the keys and values ​​are sparsely selected within the sliding window surrounding the query block. The self-attention operation is then performed on these keys and values, expanding the receptive field of the window attention.

[0130] Using a multi-headed MSDA design, the channels of the feature map are divided into four different heads. Different heads perform SWDA using different dilation rates, for example, the first head has a dilation rate of 1, the second head has a dilation rate of 2, the third head has a dilation rate of 3, and the fourth head has a dilation rate of 4. Given a feature map X, the corresponding query, key, and value of the feature map X are obtained through linear projection. The channels of the feature map are then split into four different heads, and multi-scale SWDA is performed in each head using a different dilation rate, thus adaptively extracting damage feature information at different scales of wind turbine blades.

[0131] The third step is to calculate the results based on the lightweight fusion multi-scale attention mechanism:

[0132] The image data after multi-dimensional data augmentation was divided into training and test sets in an 8:2 ratio. The training set was used to train the constructed lightweight YOLO algorithm model that integrates the multi-scale attention mechanism. Stochastic gradient descent (SGD) was used as the optimizer during training, with a learning rate of 0.001, momentum of 0.9, and weight decay of 0.0005. The number of training epochs was set to 500.

[0133] During training, we reverse-optimize and adjust model parameters based on evaluation metrics such as accuracy, recall, and mean average precision (mAP). After 500 rounds of training, the trained model was tested on the test set. The final model, SNMSDA-YOLO11, achieved a mAP of 94.9% on the test set, as shown in Table 1.

[0134] Table 1 Data comparison table

[0135]

[0136]

[0137] The results of this experiment show that although the mAP@0.5 of SNMSDA-YOLO11 is 92.6%, which is slightly lower than that of MSDA-YOLO11 (95.5%), its inference speed is faster, reaching 92.7 frames per second (FPS), which is 4FPS faster than MSDA-YOLO11. SNMSDA-YOLO11 has fewer parameters and lower computational complexity, with 2.48 million parameters and 5.8 gigaflops (GFLOPs), showing better computational efficiency and inference speed. In comparison, the inference speeds of SNMSDA-YOLO11 and MSDA-YOLO11 are 78.1FPS and 88.7FPS, respectively. For different types of wind turbine blade damage, four types of damage, namely cracks, glass fiber breakage, epidermal loss and pitting, can be accurately identified and detected, achieving high-precision online wind turbine blade damage detection and identification.

[0138] Overall, SNMSDA-YOLO11 is more suitable for real-time detection while maintaining high accuracy, especially in resource-constrained environments.

[0139] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An intelligent wind turbine blade damage identification method integrating multi-scale attention is characterized by: include: Use drones to collect images of wind turbine blades during inspections; Perform multi-dimensional data enhancement on the collected images; Integrate the multi-scale attention mechanism into the lightweight YOLO11 to build the SNMSDA-YOLO11 algorithm model; The trained SNMSDA-YOLO11 algorithm model is used to detect and identify wind turbine blade damage online.

2. The wind turbine blade damage intelligent identification method integrating multi-scale attention as claimed in claim 1 is characterized by: The multi-dimensional data enhancement is performed by adding random noise, sharpening the image, and adjusting the image saturation.

3. The wind turbine blade damage intelligent identification method integrating multi-scale attention as claimed in claim 2 is characterized by: The random noise enhancement includes simulating shooting distortion by adding Gaussian noise and salt and pepper noise; I noisy1 (x,y)=I(x,y)+N(x,y) Where x represents the horizontal coordinate of the pixel in the image; y represents the vertical coordinate of the pixel in the image; I(x,y) represents the original pixel value at the coordinate (x,y); N(x,y) represents the Gaussian noise at the coordinate (x,y), N(x,y)~N(μ,σ 2 ) means that N(x,y) has a mean of μ and a variance of σ 2 Gaussian distribution with variance σ 2 Randomly select in the range [10,50]; I noisy1 (x,y) represents the pixel value at the coordinate (x,y) after Gaussian noise processing; I noisy2 (x,y) represents the pixel value at coordinate (x,y) after salt and pepper noise processing; P salt +P pepper =P amount , P amount ∈[0.01,0.1] represents the noise ratio; P salt represents the probability of salt noise appearing in salt and pepper noise; P pepper Represents the probability of pepper noise appearing in salt and pepper noise.

4. The wind turbine blade damage intelligent identification method integrating multi-scale attention as claimed in claim 3 is characterized by: The image sharpening includes highlighting the image edges and details using the Laplacian operator; Among them, K is the convolution kernel, defined as: Among them, I sharpened (x,y) represents the pixel value at the coordinate (x,y) after sharpening; K(i,j) represents the value of the coordinate position (i,j) in the convolution kernel matrix K; i represents the row index in the convolution kernel matrix, and j represents the column index in the convolution kernel matrix.

5. The wind turbine blade damage intelligent identification method integrating multi-scale attention as claimed in claim 4 is characterized by: The saturation adjustment includes adjusting the saturation in the HSV color space; S'(x,y)=clip(S(x,y)·α,0,255) Where S(x,y) represents the original saturation value; S'(x,y) represents the adjusted saturation value, α∈[0.8,1.2] represents the randomly generated scaling factor; clip represents the constraint function to ensure that the result is within the valid range of [0,255]. After completing the adjustments, convert the adjusted image back to the BGR color space.

6. The intelligent wind turbine blade damage identification method integrating multi-scale attention as claimed in claim 5 is characterized by: The lightweight YOLO11 includes integrating the SlimNeck structure into YOLO11, using the SlimNeck structure in the neck of YOLO11, replacing the C3K2 module in the neck of YOLO11 with the existing VoVGSCSP module, and replacing the conv module in the neck of YOLO11 with the existing GSConv module; The SlimNeck structure is mainly composed of a GSConv module and a VoVGSCSP module.

7. The wind turbine blade damage intelligent identification method integrating multi-scale attention as claimed in claim 6 is characterized by: The multi-scale attention mechanism adopts a multi-head design to divide the channels of the feature map into n different heads. Different heads use different void rates to perform sliding window void attention operations. Different heads use different dilation rates to perform sliding window void attention, which sparsely selects keys and values ​​within a sliding window around the query block and performs self-attention operations; For a given position (i, j), the hole rate r is defined to control sparsity; the query in the input feature map is sparsely selected in a sliding window of size w×w, and the self-attention operation is performed; the output feature map X of the sliding window hole attention operation corresponds to the component x ij Defined as: 1≤i≤W 1≤j≤H Among them, x ij is the component at coordinate (i, j) in the output feature map X of the sliding window hole attention operation; H and W are the height and width of the feature map respectively, and K r and V r Represents the key and value selected from K and V according to the void rate r, d k is the dimension of K; Given a query at (i, j), (i', j') represents the coordinates corresponding to the key and value selected according to the void rate r within the sliding window centered at the query location (i, j). The specific coordinates are determined by the formula: {(i',j')|i'=i+p×r,j'=j+q×r} The key and value at position (i',j') will be used to match the query Q ij Perform self-attention calculation to obtain the component x of the output feature map at the (i, j) position ij ; For the feature map X, the query, key, and value corresponding to the feature map X are obtained through linear projection; the channel of the feature map is divided into n different heads, and a multi-scale sliding window hole attention operation is performed in each head with different hole rates, which can be expressed as: h k =SWDA(Q k ,K k ,V k ,r k ) 1≤k≤n X=Lincar(Concat[h1,…,h n )] Among them, h k represents the features output by the k-th head after the sliding window hole attention operation; r k is the hole rate of the kth head, Q k , K k and V k represents the feature map slice input to the kth head; n represents the number of heads; Concat is the concatenation operation, which concatenates the outputs h1,...,h n Spliced ​​together along the channel dimension; Lincar is a linear transformation operation implemented by a fully connected layer; The fusion multi-scale attention module adopts a multi-head design, which divides the channels of the feature map into n different heads. Each head uses a different hole rate to perform a sliding window hole attention operation, thereby achieving multi-scale feature extraction; the output are connected together and fed into the linear layer for feature aggregation; X represents the final feature map after the multi-head attention mechanism is processed. It is the result obtained by concatenating and linearly transforming the outputs of each head, integrating the multi-scale feature information learned by different heads; The SNMSDA-YOLO11 algorithm model includes, on the basis of the SN-YOLO11 model, embedding a multi-head attention mechanism when performing preliminary processing on the input features after the C2PSA module receives the input features.

8. An intelligent wind turbine blade damage identification system integrating multi-scale attention using the method according to any one of claims 1 to 7, characterized in that: The acquisition unit collects images of wind turbine blades through drones during inspections; A processing unit performs multi-dimensional data enhancement on the collected images; The recognition unit integrates the multi-scale attention mechanism into the lightweight YOLO11 to build the SNMSDA-YOLO11 algorithm model; the trained SNMSDA-YOLO11 algorithm model is used to detect and identify wind turbine blade damage online.

9. A computer device comprising: A memory and a processor; the memory stores a computer program, wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.