Vision-based wind blade damage defect detection method and system

By constructing a damage association knowledge graph and an improved YOLOv11 detection network, combined with an adaptive input unit, the problems of long detection cycles and low accuracy of wind turbine blades are solved, realizing automated, rapid location and efficient detection of wind turbine blade damage.

CN121686077BActive Publication Date: 2026-05-12VOCATIONAL & TECH COLLEGE OF INNER MONGOLIA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VOCATIONAL & TECH COLLEGE OF INNER MONGOLIA AGRI UNIV
Filing Date
2025-12-12
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In the current technology, the inspection of wind turbine blades mainly relies on manual visual inspection, which has the problems of long inspection cycle, difficulty in quantitative evaluation, and insufficient inspection of high-risk areas, making it impossible to detect damage in time and determine maintenance priorities.

Method used

A vision-based wind turbine blade damage and defect detection method is adopted. By acquiring the working environment features and damage defects of multiple sample blades, a damage association knowledge graph is constructed. An improved YOLOv11 detection network is used, combined with an adaptive input unit to dynamically adjust the image resolution, so as to achieve automated detection.

Benefits of technology

It enables rapid location and efficient detection of wind turbine blade damage, reduces computational load, and improves detection accuracy and efficiency. It can quickly locate high-risk components based on real-time environmental information, achieving a balance between accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121686077B_ABST
    Figure CN121686077B_ABST
Patent Text Reader

Abstract

The application provides a visual-based wind blade damage defect detection method and system, and relates to the field of image recognition.The method comprises the following steps: constructing a damage correlation knowledge graph based on the working environment characteristics and damage defects of a plurality of sample wind blades; constructing and training an improved YOLOv11 detection network, wherein the improved YOLOv11 detection network comprises an adaptive input unit, and the adaptive input unit is used for adjusting the resolution of an image based on the image features of the image and the damage correlation knowledge graph; constructing damage detection paths corresponding to various working environment types based on the damage correlation knowledge graph and the improved YOLOv11 detection network; determining a damage detection path based on the working environment characteristics of a wind blade to be detected and the damage detection paths corresponding to various working environment types; and performing damage defect detection on the wind blade to be detected based on the damage detection path and the improved YOLOv11 detection network, thereby achieving the advantage of automatic detection of wind blade damage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition, and in particular to a vision-based method and system for detecting damage and defects in wind turbine blades. Background Technology

[0002] As the core functional component of wind turbine generators for capturing wind energy, the structural reliability of wind turbine blades directly determines the generator's power generation efficiency and service life. Ideally, the blades can efficiently convert wind energy into mechanical energy, thereby driving the generator to produce electricity. However, in actual operating environments, wind turbine blades face numerous complex and harsh conditions. In extreme environments like sandstorms, a large number of high-speed dust particles continuously scour and impact the blade surface, causing gradual wear and tear, resulting in scratches, pits, and other damage. Salt spray corrosion is also a common corrosion phenomenon, especially in coastal areas or offshore wind farms. Salt in the air forms salt spray deposits on the blade surface, triggering electrochemical corrosion, damaging the protective coating, and causing the blade material to gradually lose its original performance. Furthermore, the constant fluctuations in temperature and humidity also adversely affect the blades. Differences in the coefficients of thermal expansion between different materials can cause internal stress within the blades, potentially leading to the formation and propagation of microcracks over time. Under the combined effect of these complex environmental stresses, the blade surface is highly susceptible to progressive damage such as delamination and cracking. These damages may not be obvious at first, but they will gradually worsen over time. Once such defects occur, the aerodynamic performance of the blades will decrease significantly, leading to a reduction in aerodynamic efficiency and power generation. More seriously, these damages may also trigger catastrophic structural failures, such as blade breakage, which will not only cause huge economic losses, but also pose a serious threat to the surrounding environment and human safety.

[0003] Currently, the maintenance of wind turbine blades primarily relies on traditional methods, mainly manual visual inspection. While this method can detect obvious damage to the blade surface to some extent, it has several insurmountable drawbacks. Firstly, the inspection cycle for manual visual inspection is too long. Because wind farms are typically widely distributed, with numerous blades installed at high altitudes, manual inspection requires significant time and manpower. Furthermore, to ensure comprehensiveness, all blades often need to be inspected periodically, further extending the inspection cycle and making it difficult to detect damage that occurs between inspections. Secondly, manual visual inspection lacks quantitative evaluation. Inspectors rely mainly on experience to judge the extent of blade damage, lacking objective and accurate quantitative indicators. Different inspectors may have different evaluations of the same damage, complicating subsequent maintenance decisions and hindering the precise determination of maintenance priorities and plans. In addition, manual visual inspection does not adequately cover high-risk areas. Some blade parts, such as blade tips and roots, are difficult to observe closely due to their unique locations, creating blind spots. Damage to these high-risk areas may be more severe, but because they cannot be detected in time, they are prone to causing more serious consequences.

[0004] Therefore, there is a need to provide a vision-based method and system for detecting wind turbine blade damage and defects, so as to achieve automated detection of wind turbine blade damage. Summary of the Invention

[0005] This invention provides a vision-based method for detecting damage and defects in wind turbine blades, comprising: acquiring the working environment characteristics and damage defects of multiple sample wind turbine blades, wherein the working environments of the multiple sample wind turbine blades differ; constructing a damage association knowledge graph based on the working environment characteristics and damage defects of the multiple sample wind turbine blades, wherein the damage association knowledge graph is used to record the damage association relationships of multiple components of the wind turbine blades corresponding to various working environment types; and constructing and training an improved YOLOv11 detection network, wherein the improved YOLOv11 detection network includes an adaptive input unit, the adaptive input unit being used for detection based on the graph. The image resolution is adjusted based on image features and a damage association knowledge graph. Damage detection paths corresponding to various working environment types are constructed based on the damage association knowledge graph and an improved YOLOv11 detection network. These damage detection paths determine the detection order of multiple components of the wind turbine blade. Images and working environment information of the wind turbine blade to be detected are acquired. Damage detection paths are determined based on the working environment features of the wind turbine blade and the damage detection paths corresponding to various working environment types. Damage defects of the wind turbine blade are detected based on the damage detection paths and the improved YOLOv11 detection network.

[0006] Furthermore, based on the working environment characteristics and damage defects of multiple sample wind turbine blades, a damage association knowledge graph is constructed, including: clustering multiple sample wind turbine blades based on their working environment characteristics to determine multiple blade classes, where each blade class corresponds to a working environment type; for each blade class, determining the damage association coefficient between any two components of the wind turbine blade based on the damage defects of the multiple sample wind turbine blades included in the blade class; determining the damage-associated components of each component corresponding to the blade class based on the damage association coefficient between any two components of the wind turbine blade corresponding to each blade class, and constructing a subgraph corresponding to the blade class; and constructing a damage association knowledge graph based on the subgraph corresponding to each blade class.

[0007] Furthermore, the adaptive input unit adjusts the image resolution based on the image features and damage association knowledge graph of the image, including: locating the target in the image and extracting the image features, wherein the image features include at least background complexity, target density, and average target size; determining the scene complexity of the image based on the image features; determining the component corresponding to the image; obtaining the damage and defect detection results of the damage association components of the component corresponding to the image; and adjusting the image resolution based on the scene complexity of the image and the damage and defect detection results of the damage association components of the component corresponding to the image.

[0008] Furthermore, based on the damage association knowledge graph and the improved YOLOv11 detection network, damage detection paths corresponding to various working environment types are constructed, including: for each working environment type, determining the sampling probability of each component based on the sub-graph corresponding to the blade class and the damage defects of multiple sample wind turbine blades included in the blade class; initializing the population based on the sampling probability of each component and the damage association knowledge graph, wherein the population includes multiple individuals, each individual corresponding to a damage detection path; constructing a fitness function, wherein the fitness function is related to the damage detection efficiency and the total number of pixels corresponding to the individual; sampling from multiple sample wind turbine blades included in the blade class to obtain multiple sample wind turbine blades; for each individual, generating the damage detection result of each sample wind turbine blade based on the damage detection path corresponding to the individual through the improved YOLOv11 detection network; calculating the fitness value of each individual based on the damage detection result of each sample wind turbine blade and the fitness function; and performing selection, crossover, and mutation operations based on the fitness value of each individual to update the population and iteratively optimize until the termination condition is met.

[0009] Furthermore, based on the sub-map corresponding to the blade class and the damage and defects of multiple sample wind turbine blades included in the blade class, the sampling probability of each component is determined, including: determining the damage and defect probability of each component based on the damage and defects of multiple sample wind turbine blades included in the blade class; and determining the sampling probability of each component based on the sub-map corresponding to the blade class, the damage and defect probability of the component, and the damage and defect probability of the component's associated components.

[0010] Furthermore, based on the sampling probability of each component and the damage association knowledge graph, population initialization is performed, including: constructing path constraints, wherein the path constraints include at least two adjacent components in the damage detection path being damage-associated components, component repetition constraints, and component coverage constraints; for each individual, generating the current probability of each component, and generating the individual based on the current probability and sampling probability of each component and the path constraints.

[0011] Furthermore, the intermediate layer of the improved YOLOv11 detection network includes an initial feature fusion unit, an iterative feature fusion unit, a multi-scale interaction unit, and a feature refinement and output unit. The initial feature fusion unit is used to scale-expand the multi-scale feature map output by the backbone of the improved YOLOv11 detection network to generate an expanded feature map. The iterative feature fusion unit is used to generate three sets of bidirectionally fused feature maps based on the expanded feature map. The multi-scale interaction unit is used to perform cross-scale feature interaction and symmetric scale transformation on the three sets of bidirectionally fused feature maps to generate a processed multi-scale feature map. The feature refinement and output unit is used to perform multi-level refinement and scale expansion on the processed multi-scale feature map to generate five sets of feature maps.

[0012] Furthermore, the detection head of the improved YOLOv11 detection network includes a feature preprocessing unit, a spatial adaptive upsampling unit, a dual-path feature fusion unit, a four-level progressive feature fusion network, and a three-level optimized feature map output layer. Specifically, the feature preprocessing unit dynamically adjusts the receptive field and enhances the features of the five sets of feature maps output from the intermediate layer of the improved YOLOv11 detection network to generate a semantically unified basic feature map. The spatial adaptive upsampling unit upsamples the basic feature map to generate an upsampled feature map. The dual-path feature fusion unit fuses the basic feature map and the upsampled feature map to generate a semantically focused feature map. The four-level progressive feature fusion network generates multi-level enhanced feature maps based on the semantically focused feature map. The three-level optimized feature map output layer generates a fused feature map based on the multi-level enhanced feature maps.

[0013] Furthermore, the loss function used to train the improved YOLOv11 detection network includes a damage / defect classification loss and a bounding box loss, wherein the damage / defect classification loss is related to the sampling probability of each part corresponding to each working environment type.

[0014] This invention provides a vision-based wind turbine blade damage and defect detection system, comprising: a data acquisition module for acquiring the working environment characteristics and damage defects of multiple sample wind turbine blades, wherein the working environments of the multiple sample wind turbine blades differ; a graph construction module for constructing a damage association knowledge graph based on the working environment characteristics and damage defects of the multiple sample wind turbine blades, wherein the damage association knowledge graph records the damage association relationships of multiple components of the wind turbine blades corresponding to various working environment types; and a model building module for constructing and training an improved YOLOv11 detection network, wherein the improved YOLOv11 detection network includes an adaptive input unit, which is used to detect damage based on image features and... The system includes a damage association knowledge graph for adjusting image resolution; a path optimization module for constructing damage detection paths corresponding to various operating environment types based on the damage association knowledge graph and an improved YOLOv11 detection network, wherein the damage detection paths are used to determine the detection order of multiple components of the wind turbine blade; a data acquisition module for acquiring images and operating environment information of the wind turbine blade to be inspected; a defect detection module for determining damage detection paths based on the operating environment characteristics of the wind turbine blade to be inspected and the damage detection paths corresponding to various operating environment types; and a defect detection module for detecting damage and defects of the wind turbine blade to be inspected based on the damage detection paths and the improved YOLOv11 detection network.

[0015] Compared with existing technologies, the vision-based wind turbine blade damage and defect detection method and system provided by this invention have at least the following advantages:

[0016] 1. By integrating sample data from multiple working conditions, a knowledge graph was constructed that records the correlation between different working environment types and blade component damage, providing a dynamic basis for detection path planning and enabling the model to quickly locate high-risk components based on real-time environmental information;

[0017] 2. By embedding an adaptive input unit in YOLOv11, this method achieves dynamic adjustment of image resolution: the network can automatically select the optimal resolution for preprocessing based on the clarity, noise level, and damage association knowledge graph of the input image. On the one hand, by accurately matching the resolution with the scene requirements, it avoids the overcomputation of traditional fixed-resolution models for simple scenes (such as allocating high-resolution resources to defect-free background areas) and the loss of accuracy for complex scenes (such as missing small defects due to insufficient resolution). While maintaining detection accuracy, it significantly reduces the amount of computation and improves the inference speed. On the other hand, the introduction of damage association knowledge makes resolution adjustment more task-oriented, achieving a balance between accuracy and efficiency. Attached Figure Description

[0018] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0019] Figure 1 This is a flowchart illustrating a vision-based wind turbine blade damage and defect detection method according to some embodiments of this specification.

[0020] Figure 2 This is a schematic flowchart illustrating the process of adjusting image resolution according to some embodiments of this specification;

[0021] Figure 3 This is a flowchart illustrating the construction of a damage detection path corresponding to a working environment type, based on some embodiments of this specification.

[0022] Figure 4 This is a schematic diagram of experimental results for an improved YOLOv11 detection network according to some embodiments shown in this specification;

[0023] Figure 5 This is a schematic diagram of experimental results for an existing YOLOv11 detection network, as shown in some embodiments of this specification.

[0024] Figure 6 This is a schematic diagram of a vision-based wind turbine blade damage and defect detection system according to some embodiments of this specification. Detailed Implementation

[0025] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0026] Figure 1 This is a flowchart illustrating a vision-based wind turbine blade damage and defect detection method according to some embodiments of this specification, such as... Figure 1 As shown, a vision-based method for detecting damage and defects in wind turbine blades may include the following steps.

[0027] Step 110: Obtain the working environment characteristics and damage defects of multiple sample wind turbine blades.

[0028] Among them, the working environments of the wind turbine blades in multiple samples differed.

[0029] Specifically, the characteristics of the working environment include multiple aspects, such as wind speed, wind direction, temperature, humidity, and the amount of dust in the air. Wind speed and direction directly affect the stress on the wind turbine blades. Different combinations of wind speed and direction may lead to different stress distributions in different parts of the blades, thus affecting the occurrence and development of damage. Temperature changes cause thermal expansion and contraction of the blade materials, which may lead to material fatigue and structural damage under long-term effects. Excessive humidity may cause the blade surface materials to become damp, reducing their performance and increasing the risk of corrosion and damage. When the dust content in the air is high, the erosion and wear of the blade surface by dust will be intensified, affecting the service life of the blades. These environmental parameters can be monitored and recorded in real time by installing various sensors near the wind turbine generator. For example, anemometers and wind direction meters are used to measure wind speed and direction, temperature sensors and humidity sensors measure temperature and humidity respectively, and dust detectors are used to detect the amount of dust in the air. These sensors transmit the collected data to a data storage system for subsequent analysis and processing.

[0030] For each sample wind turbine blade, detailed information such as the location, type, and extent of any detected damage or defects needs to be recorded. The location of the damage can be described using specific coordinates or relative positions on the blade; the types of damage include cracks, wear, corrosion, and deformation; and the extent of damage can be measured using quantitative indicators such as crack length and wear depth.

[0031] Step 120: Based on the working environment characteristics and damage defects of multiple sample wind turbine blades, construct a damage association knowledge graph.

[0032] Among them, the damage association knowledge graph is used to record the damage association relationships of multiple components of wind turbine blades corresponding to various working environment types.

[0033] Specifically, it includes:

[0034] Based on the working environment characteristics of multiple sample wind turbine blades, multiple sample wind turbine blades are clustered to determine multiple blade classes, where each blade class corresponds to a working environment type.

[0035] For each blade class, based on the damage defects of multiple sample wind turbine blades included in the blade class, the damage correlation coefficient between any two components of the wind turbine blade is determined. Based on the damage correlation coefficient between any two components of the wind turbine blade corresponding to each blade class, the damage-related components of each component corresponding to the blade class are determined, and a sub-map corresponding to the blade class is constructed.

[0036] A damage association knowledge graph is constructed based on the subgraph corresponding to each blade class.

[0037] Specifically, since different working environments have significantly different effects on blade damage, grouping sample wind turbine blades with similar working environments into one category allows for a more accurate analysis of the commonalities and characteristics of blade damage under that environment. For example, sample blades operating in high wind speed and high humidity environments can be grouped into one category, while sample blades operating in low wind speed and dry environments can be grouped into another.

[0038] For each sample wind turbine blade, an environmental feature vector can be constructed. The cosine similarity between the environmental feature vectors of any two sample wind turbine blades is calculated. Then, a clustering algorithm (e.g., K-means clustering) is used to cluster multiple sample wind turbine blades based on the cosine similarity between their environmental feature vectors, thus determining multiple blade classes.

[0039] Wind turbine blades can be divided into multiple components based on their structure and function; for example, the root, middle, and tip of the blade can be considered different components. For multiple sample wind turbine blades within each blade type, the damage and defect detection results are analyzed, and the damage correlation coefficient between any two components is calculated. This coefficient reflects the degree of correlation between the occurrence of damage in the two components.

[0040] For each sample wind turbine blade, the damage and defects of each component can be numerically coded. For example, each damage and defect can be coded as a number to identify the degree of damage to the component. For instance, a graded coding system (e.g., levels 1-5) can be used. The damage and defect codes for each sample wind turbine blade corresponding to the two components are then substituted into the correlation coefficient calculation formula (e.g., Pearson correlation coefficient, Spearman rank correlation coefficient, etc.) to calculate the damage correlation coefficient between the two components.

[0041] For any two components, if the damage correlation coefficient between the two components is greater than a damage correlation coefficient threshold (e.g., 0.5), then the two components are considered to be damage-related components. The subgraph corresponding to the blade class can include nodes representing the components. The nodes corresponding to the two components that are damage-related components are connected by edges, and the weight of the edges represents the damage correlation coefficient between the two components.

[0042] The damage association knowledge graph can include a subgraph corresponding to each blade class.

[0043] Step 130: Build and train the improved YOLOv11 detection network.

[0044] The improved YOLOv11 detection network includes an adaptive input unit, which adjusts the image resolution based on image features and a damage association knowledge graph.

[0045] Figure 2 This is a schematic flowchart illustrating the process of adjusting image resolution according to some embodiments of this specification, such as... Figure 2 As shown, it specifically includes:

[0046] The image is used to locate the target and extract the image features, wherein the image features include at least background complexity, target density and average target size;

[0047] Determine the scene complexity of an image based on its image features;

[0048] Identify the component corresponding to the image;

[0049] Obtain the damage and defect detection results of the component corresponding to the image;

[0050] Based on the scene complexity of the image and the damage and defect detection results of the corresponding component, the image resolution is adjusted.

[0051] Specifically, the adaptive input unit locates targets in the image using object detection algorithms (such as YOLO and Faster R-CNN). Background complexity is quantified by calculating the texture entropy or edge density of non-target regions (such as using the LBP operator to statistically analyze local texture changes). The more complex the background (such as the presence of interfering objects, noise, or blurred areas), the higher the value. Target density is measured by the number of targets per unit area or the reciprocal of the average distance between targets. High-density scenes (such as dense cracks or crowds) require higher resolution to support detection. The average target size is reflected by statistically analyzing the mean area or median aspect ratio of all target bounding boxes.

[0052] The background complexity, target density, and average target size can be normalized. Based on the normalized background complexity, target density, and average target size, the scene complexity of the image can be calculated. The higher the normalized background complexity, the higher the target density, and the smaller the average target size, the higher the scene complexity.

[0053] Identify the corresponding parts of an image (such as the root, middle, and tip of a leaf) through semantic segmentation (such as U-Net) or instance segmentation (such as Mask R-CNN).

[0054] The image resolution can be adjusted based on the scene complexity of the image and the damage and defect detection results of the damage-related components of the corresponding component. For example, the higher the scene complexity of the image and the higher the proportion of damage and defects detected in the damage-related components of the corresponding component, the higher the image resolution.

[0055] As an example only, the image resolution can be adjusted based on the following formula:

[0056] ,

[0057] in, To adjust the resolution of the image, For the scene complexity of the image, Given the preset scene complexity, Let be the damage correlation coefficient between the component corresponding to the image and the nth damage-related component. This is the damage defect identifier for the nth damage-associated component. The value is 1 when a damage defect is detected in the nth damage-associated component, and 0 otherwise. These are preset parameters. Greater than 1, The total number of damaged and associated components. The initial image resolution, To Values.

[0058] This formula achieves a balance between accuracy and efficiency by dynamically adjusting the image resolution through quantifying scene complexity and damage correlation strength. In the formula, This represents the ratio of the current scene complexity to a preset baseline, reflecting the resolution requirements of features such as background interference and target density. The damage correlation coefficients (weights) and defect identifiers (0 or 1) of the target component and each associated component are summed to measure the risk of damage propagation. The sum of these two factors, divided by a parameter (which controls the adjustment range), is used as a resolution scaling factor, which is then multiplied by the initial resolution to obtain the target resolution. Its function is as follows: when the scene complexity is high (e.g., densely packed and small targets) or the risk of damage to associated components is high (e.g., multiple components detect defects), the scaling factor increases, and the resolution automatically increases (e.g., from 640×640 to 960×960) to capture detailed features; conversely, the resolution decreases (e.g., reduced to 320×320) to reduce redundant calculations.

[0059] Understandably, the process begins by locating the target in the image and extracting key features such as background complexity, target density, and average target size to quantitatively assess the scene's complexity level (e.g., simple, medium, complex). Simultaneously, a damage association knowledge graph is used to identify damage detection results of the target component and its associated components, analyzing the risk of damage propagation. Finally, the input resolution is dynamically switched based on a comprehensive assessment of scene complexity and associated damage intensity. For images with high complexity (e.g., strong background interference, dense and small targets) or high associated damage risk (e.g., many associated components showing damage defects), the resolution is automatically increased to a higher resolution (e.g., from 640×640 to 960×960) to capture detailed features; for images with low complexity and low risk, the resolution is reduced (e.g., scaled to 320×320) to minimize redundant computation. The mechanism has significant benefits: on the one hand, by accurately matching the resolution with the scene requirements, it avoids the overcomputation of simple scenes (such as allocating high-resolution resources to defect-free background areas) and the loss of accuracy in complex scenes (such as missing small defects due to insufficient resolution) of traditional fixed-resolution models. While maintaining detection accuracy, it significantly reduces the amount of computation and improves the inference speed. On the other hand, the introduction of damage correlation knowledge makes resolution adjustment more task-oriented, achieving a balance between accuracy and efficiency.

[0060] The improved YOLOv11 detection network also includes a backbone network, intermediate layers, and a detection head.

[0061] The backbone network employs a dual strategy of dynamic channel pruning and cross-layer feature reuse.

[0062] The core component of the backbone network, the novel CSPDarknetV2 network, employs a dual strategy of dynamic channel pruning and cross-layer feature reuse, significantly reducing model parameters and computational resource consumption while ensuring detection performance. The input RGB image, after preprocessing, enters the feature encoding stage. The initial stage uses a two-level feature extraction unit based on improved Ghost convolutions. This unit integrates adaptive normalization layers and nonlinear enhanced activation functions, reducing computational cost while improving local feature discrimination capabilities. Subsequently, the feature maps are fed into the upgraded C3-X module, which innovatively introduces a gated attention mechanism. This module dynamically assigns feature aggregation weights based on feature importance evaluation, achieving approximately 15% improvement in feature fusion efficiency compared to the traditional C3 module. The four-stage cascaded C3-X module group constructs a multi-scale feature pyramid with spatial awareness through depthwise separable convolutions and cross-layer feature interaction. The SPPF++ multi-scale feature fusion module at the backbone end achieves an architectural breakthrough, incorporating a Transformer self-attention mechanism into the classic spatial pyramid pooling framework. By constructing a collaborative network that associates local feature responses with global semantics, this module effectively enhances the robustness of target representation in complex scenarios. The feature reduction process employs a progressive step-size optimization strategy, dynamically adjusting the convolution kernel parameters during the five-stage convolution operation to gradually compress the input resolution to 1 / 32 of the original image, ensuring efficient semantic abstraction while minimizing spatial information loss. Experimental results show that this backbone network achieves a better balance between model complexity and feature representation capability through innovative module combinations and optimization strategies. Its output multi-level feature maps possess rich spatial details and semantic information, laying the foundation for high-precision multi-scale target localization in the detection head network.

[0063] The intermediate layer completes information transmission and optimization through a multi-stage feature integration mechanism. In some embodiments, the intermediate layer of the improved YOLOv11 detection network includes an initial feature fusion unit, an iterative feature fusion unit, a multi-scale interaction unit, and a feature refinement and output unit. The initial feature fusion unit is used to scale-expand the multi-scale feature map output by the backbone network of the improved YOLOv11 detection network to generate an expanded feature map. The iterative feature fusion unit is used to generate three sets of bidirectionally fused feature maps based on the expanded feature map. The multi-scale interaction unit is used to perform cross-scale feature interaction and symmetric scale transformation on the three sets of bidirectionally fused feature maps to generate a processed multi-scale feature map. The feature refinement and output unit is used to perform multi-level refinement and scale expansion on the processed multi-scale feature map to generate five sets of feature maps.

[0064] Specifically, the initial feature fusion unit takes three sets of multi-scale feature maps (1 / 8, 1 / 16, and 1 / 32 resolution) output by the backbone network as input. For the high-level features (1 / 32), deformable convolutional kernels are used. By dynamically adjusting the sampling position of the convolutional kernels (e.g., introducing an offset learning branch), the model's adaptability to target geometric deformations (such as rotation and scaling) is improved. Subsequently, bilinear interpolation is used to upsample the high-level features by a factor of 2, expanding their scale to match that of the mid-level features (1 / 16), providing a spatial alignment basis for subsequent cross-level fusion. This process mitigates feature distortion caused by resolution differences by preserving high-level semantic information while introducing shallow spatial details.

[0065] The iterative feature fusion unit constructs a bidirectional feature propagation path, both top-down and bottom-up:

[0066] Top-down fusion: The expanded high-level features (1 / 16) and mid-level features are aligned in the Feature AlignmentModule (FAM). The FAM module employs a dynamic feature selection strategy based on attention weights. It generates a feature importance weight map through a lightweight attention sub-network (such as a sequential combination of channel attention and spatial attention), dynamically weighting the input features to effectively suppress background noise (such as redundant textures in complex backgrounds) and highlight target-related features. The fused mid-level features continue to be iteratively fused with the shallow features (1 / 8) to form a complete top-down path.

[0067] Bottom-up aggregation: To enhance the transmission of shallow spatial details, the network injects spatial details from shallow features (1 / 8) into deep semantic features (1 / 16, 1 / 32) through downsampling (e.g., convolution with a stride of 2) and feature concatenation operations, constructing a multi-directional information interaction network. Each fusion node is equipped with a channel pruning unit, employing a learnable channel importance scoring mechanism (e.g., channel contribution evaluation based on gradient or statistical information) to dynamically remove low-contribution channels, achieving dimensionality compression while preserving key information and reducing redundant computation.

[0068] Multi-scale interaction units achieve deep inter-level collaboration through cross-scale feature interaction mechanisms, comprising two key stages:

[0069] Multi-granularity context extraction: Spatial pyramid pooling (SPP) is used to perform multi-scale pooling operations (such as average pooling with different kernel sizes) on the three sets of input feature maps to extract local and global context features, thereby enhancing the model's ability to perceive the surrounding environment of the target.

[0070] Cross-level association modeling: A multi-head cross-attention mechanism is introduced, using high-level features (1 / 32) as the query and mid-level (1 / 16) and shallow-level (1 / 8) features as the key and value, respectively, to construct a cross-level association matrix. Through adaptive feature weighting (such as feature aggregation based on attention weights), dynamic fusion of features at different scales is achieved, improving the model's ability to detect multi-scale targets.

[0071] The entire processing flow includes four symmetrical scaling operations (such as two upsampling and two downsampling). By optimizing the convolution kernel parameters (such as kernel size and dilation rate) and stride configuration (such as convolution with stride of 1 or 2), the integrity of information during the feature map size transformation process is ensured, and feature loss due to scale changes is avoided.

[0072] The feature refinement and output unit performs multi-layer refinement on the processed multi-scale feature maps: first, high-level semantic features are further extracted through convolutional operations; then, a progressive scale expansion strategy (such as stepwise upsampling) is used to generate five sets of feature maps (1 / 8, 1 / 16, 1 / 32, 1 / 64, and 1 / 128 resolutions). The 1 / 8 and 1 / 16 resolution feature maps retain rich spatial details (such as edge information of small targets), while the 1 / 32, 1 / 64, and 1 / 128 resolution feature maps encode high-order semantic information (such as target category and contextual relationships). This multi-scale output design enables the detection head network to process targets of different sizes simultaneously, improving the detection accuracy for both small targets (such as fine cracks) and large targets (such as large fissures). Experiments show that this intermediate layer design improves target localization accuracy by approximately 2.3 AP (average accuracy) while maintaining low computational complexity (FLOPs reduced by approximately 18%), especially in complex scenes (such as cluttered backgrounds and dense targets), significantly outperforming the traditional FPN architecture.

[0073] In some embodiments, the detection head of the improved YOLOv11 detection network includes a feature preprocessing unit, a spatial adaptive upsampling unit, a dual-path feature fusion unit, a four-level progressive feature fusion network, a three-level optimized feature map output layer, and a prediction layer. The feature preprocessing unit dynamically adjusts the receptive field and enhances the features of the five sets of feature maps output from the intermediate layers of the improved YOLOv11 detection network to generate a semantically unified basic feature map. The spatial adaptive upsampling unit upsamples the basic feature map to generate an upsampled feature map. The dual-path feature fusion unit fuses the basic feature map and the upsampled feature map to generate a semantically focused feature map. The four-level progressive feature fusion network generates multi-level enhanced feature maps based on the semantically focused feature map. The three-level optimized feature map output layer generates a fused feature map based on the multi-level enhanced feature maps.

[0074] Specifically, the feature preprocessing unit takes five sets of multi-scale feature maps (1 / 8, 1 / 16, 1 / 32, 1 / 64, and 1 / 128 resolution) output from the intermediate layer as input and uses an improved SPPF module (a fast version of spatial pyramid pooling) for initial feature enhancement:

[0075] Kernel parameter optimization convolutional layers: By dynamically adjusting the convolutional kernel size (e.g., 3×3, 5×5, 7×7) and dilation rate, the receptive field is reconstructed, enabling the model to adaptively capture local details (e.g., edge textures of small objects) and global context (e.g., scene associations of large objects) at different scales. Experiments show that this design improves fine-grained feature extraction efficiency by 17.3%.

[0076] Semantic unification processing: A lightweight feature alignment module is introduced, which uses deformable convolution to spatially align the five sets of feature maps, eliminates semantic shifts caused by scale differences, and generates semantically consistent basic feature maps, providing a unified benchmark for subsequent fusion.

[0077] To address the drawback of traditional interpolation methods (such as bilinear interpolation) in easily losing high-frequency information, the spatial adaptive upsampling unit proposes a differentiated upsampling strategy based on semantic distribution:

[0078] Semantic density awareness: By calculating the channel attention weights of the feature map (e.g., using the SE module to generate channel importance scores), high semantic density regions (e.g., target subjects) and low semantic density regions (e.g., background) are identified.

[0079] Dynamic upsampling kernel selection: Small stride (e.g., stride=1) is used for upsampling in high semantic regions to preserve details, while large stride (e.g., stride=2) is used for upsampling in low semantic regions to reduce redundant computation. For example, when upsampling from 1 / 32 to 1 / 16 of the feature map, a 4×4 convolution kernel is used in the target edge region and a 2×2 convolution kernel is used in the background region, reducing the loss of detail information by 31.2%.

[0080] The dual-path feature fusion unit employs a dual-path architecture to achieve deep fusion of basic features and upsampled features:

[0081] Path 1: Cross-scale stitching preserves original information: directly stitching the base feature map and the upsampled feature map to preserve the original distribution of multi-scale features (such as shallow spatial details and deep semantic information).

[0082] Path Two: Channel Attention-Guided Weighted Fusion: This approach introduces a learnable feature weight allocation matrix (generated via 1×1 convolution) and dynamically adjusts the fusion ratio based on the importance score of each channel dimension. For example, when fusing 1 / 16 base features with upsampled features, higher weights are assigned to channels relevant to the target category (e.g., weight = 0.7), while background noise channels are suppressed (e.g., weight = 0.3), achieving intelligent feature selection. Experiments show that this hybrid fusion strategy improves feature representation capability by 23.6%.

[0083] The four-level progressive feature fusion network optimizes feature representation step by step through four levels of progressive fusion components (levels 12 / 16 / 19 / 22):

[0084] Level 12: Residual Enhancement Mechanism: Introducing a cross-layer skip connection, shallow features (such as 1 / 8) are added to the current level features to strengthen feature correlation and alleviate the gradient vanishing problem. For example, in Level 12 fusion, the 1 / 8 feature is adjusted for channel number through a 1×1 convolution and then added to the current feature, improving the small object detection recall rate by 9.5%.

[0085] Level 16: Dynamic weight allocation algorithm: Automatically adjusts the fusion ratio based on feature saliency (e.g., calculated by gradient magnitude). Higher weights (e.g., weight = 0.8) are assigned to regions with high saliency (e.g., the target center), while weights are reduced (e.g., weight = 0.5) for edge regions, achieving adaptive feature aggregation.

[0086] Level 19: Multi-scale convolutional kernel group: Parallel 3×3, 5×5, and 7×7 convolutional kernels are used to extract spatial-semantic joint features at different scales, and a multi-scale representation system is constructed by concatenating channels. For example, when detecting large targets, the 7×7 convolutional kernel captures the global context, while the 3×3 convolutional kernel focuses on local details, improving the accuracy of large target recognition by 12.1%.

[0087] Level 22: Feature Enhancement Operator: Deploys a channel attention-guided feature reorganization module, which generates channel descriptors through global average pooling, and then generates channel weights through a fully connected layer to perform weighted reorganization of the feature map, thereby improving information density. For example, after level 22 fusion, the inter-channel correlation of the feature map is reduced by 28.7%, and the concentration of key information is increased by 34.2%.

[0088] The third-level optimized feature map output layer is based on multi-level enhanced feature maps generated by a four-level fusion network. The output layer is specifically optimized for different scale detection requirements.

[0089] Low-level features (17 layers): A small target sensitivity enhancement algorithm is introduced. The recall rate of small targets (such as targets with an area <32×32 in the COCO dataset) is improved by using a high-frequency information compensation mechanism (such as Laplace pyramid filtering). Experiments show that the recall rate is improved by 14.3%.

[0090] Mid-layer features (20 layers): A spatial alignment matching strategy is employed to construct an error correction model for target location regression. Deformable ROI pooling is used to align the target region, reducing localization bias and improving bounding box localization accuracy by 2.8%.

[0091] High-level features (23 layers): A semantic enhancement module was developed to improve the recognition accuracy of large targets (such as targets with an area > 96×96 in the COCO dataset) through context-aware mechanisms (such as non-local neural networks). Experiments show that the accuracy is improved by 11.7%.

[0092] The prediction layer identifies damage and defects in wind turbine blade components based on the fused feature map and outputs the identification results.

[0093] The improved YOLOv11 detection network also includes an attention mechanism, which can be CBAM, ASPP, SKAttention (SKNet), ShuffleAttention, or UFOAttention (UFO-ViT).

[0094] Step 140: Based on the damage association knowledge graph and the improved YOLOv11 detection network, construct damage detection paths corresponding to various working environment types.

[0095] The damage detection path is used to determine the detection sequence of multiple components of the wind turbine blade.

[0096] Figure 3 This is a flowchart illustrating the construction of damage detection paths corresponding to different work environment types, based on some embodiments of this specification. In some embodiments, damage detection paths corresponding to various work environment types are constructed based on damage association knowledge graphs and an improved YOLOv11 detection network, including:

[0097] For each type of working environment, the sampling probability of each component is determined based on the sub-map corresponding to the blade class and the damage defects of multiple sample wind turbine blades included in the blade class.

[0098] Based on the sampling probability of each component and the damage association knowledge graph, the population is initialized. The population includes multiple individuals, and each individual corresponds to a damage detection path.

[0099] Construct a fitness function, where the fitness function is related to the damage detection efficiency and the total number of pixels for each individual.

[0100] Multiple sample wind turbine blades are obtained by sampling from multiple sample wind turbine blades included in the blade class.

[0101] For each individual, the improved YOLOv11 detection network generates damage detection results for each sampled wind turbine blade based on the individual's corresponding damage detection path.

[0102] Based on the damage detection results and fitness function of each sampled wind turbine blade, the fitness value of each individual is calculated.

[0103] Based on the fitness value of each individual, selection, crossover, and mutation operations are performed to update the population and iteratively optimize it until the termination condition is met.

[0104] In some embodiments, the sampling probability of each component is determined based on the sub-map corresponding to the blade class and the damage defects of multiple sample wind turbine blades included in the blade class, including:

[0105] Based on the damage defects of multiple sample wind turbine blades included in the blade class, the probability of damage defects for each component is determined.

[0106] Based on the sub-map corresponding to the blade class, and based on the damage and defect probability of the component and the damage and defect probability of the component associated with the component, the sampling probability of each component is determined.

[0107] Specifically, the frequency of damage (such as cracks, corrosion, wear, etc.) to each component (e.g., blade edge, blade root, blade tip, etc.) in multiple sample wind turbine blades included in the blade category is statistically analyzed. For example, if the blade root shows 50 damage defects in 100 sample wind turbine blades included in the blade category, then its damage defect probability is 50%.

[0108] For each component, the damage correlation coefficient between the component and the damaged associated components is used as a weight, and the damage defect probabilities of the component and the damaged associated components are weighted and summed to obtain the sampling probability of the component.

[0109] In some embodiments, population initialization is performed based on the sampling probability of each component and the damage association knowledge graph, including:

[0110] Construct path constraints, which include at least the following: adjacent components in the damage detection path are damage-related components, component repetition constraints, and component coverage constraints.

[0111] For each individual, generate the current probability of each component, and generate the individual based on the current probability and sampling probability of each component and the path constraints.

[0112] Specifically, the component repetition constraint can be that the number of times the same component appears in the path is less than a threshold (e.g., 3 times).

[0113] The coverage constraint can be that the damage detection path must cover at least 80% of the components.

[0114] Specifically, when generating an individual, a starting component is randomly selected from the components according to sampling probability, and the path list and component counter (recording the number of times each component appears in the path) are initialized. Then, a recursive expansion phase begins: at each step, based on the component at the end of the current path, all candidate components with direct damage associations are selected from the damage association knowledge graph; for each candidate component, its current probability is randomly generated; if the current probability is lower than a preset threshold, the component is considered a candidate component; further constraints are applied: it is checked whether the candidate component satisfies the component repetition constraint (counter value less than the threshold, e.g., 3 times) and coverage constraint (if the current path coverage is less than 80%, uncovered key components are prioritized); if a candidate component passes all constraints, it is added to the path and the counter is updated; if all candidate components fail to meet the constraints, the process backtracks to the previous node and reselects. This expansion process is repeated until the path length reaches a preset upper limit or cannot be expanded further. The finally generated individual must pass a coverage check: if the proportion of covered key components is less than 80%, nodes are added from the uncovered key components according to sampling probability, ensuring that the uncovered key components corresponding to the added nodes are damage-associated components of the component at the end of the path.

[0115] Specifically, for each sampled wind turbine blade, the improved YOLOv11 detection network can be used to complete the damage and defect detection of the sampled wind turbine blade according to the individual damage detection path. The longer the time taken, the lower the damage detection efficiency of the sampled wind turbine blade. For each component of the sampled wind turbine blade, it can be processed according to... Figure 2 The algorithm shown calculates the resolution of the adjusted image for each component, sums the resolutions of the adjusted images for each component, and obtains the result. The larger the sum of the pixels corresponding to the sampled wind turbine blades, the better. It can be understood that the images of each component of the same sample wind turbine blade can be images acquired separately.

[0116] For each individual, calculate the mean of the damage detection efficiency corresponding to each sampled wind turbine blade, and use it as the damage detection efficiency for that individual. Also calculate the mean of the sum of the number of pixels corresponding to each sampled wind turbine blade, and use it as the sum of the number of pixels for that individual.

[0117] Higher damage detection efficiency results in fewer pixels and a higher fitness value. The fitness function is also related to the accuracy of damage and defect detection; higher accuracy leads to a larger fitness value.

[0118] Based on the fitness value of each individual, selection, crossover, and mutation operations are performed to update the population and iteratively optimize until the termination condition is met. This is the existing technology and will not be elaborated here.

[0119] Understandably, dynamically calculating sampling probabilities based on sample damage frequency and component correlation allows the detection path to focus on high-risk components and their associated areas, improving the targeting of defect detection. Secondly, path constraints ensure the physical rationality and comprehensiveness of the detection path, avoiding invalid traversal or local repetitive detection. Furthermore, the fitness function comprehensively considers detection efficiency, total pixel count, and detection accuracy, guiding the population to evolve towards high efficiency, low computational cost, and high accuracy, balancing detection speed and quality. In addition, recursive expansion and backtracking mechanisms combined with coverage verification dynamically adjust the path structure, prioritizing coverage of key components even with limited path length, enhancing the algorithm's robustness. Finally, iterative optimization of the population through genetic operations such as selection, crossover, and mutation enables automatic generation and continuous improvement of detection paths. This method not only overcomes the limitations of traditional detection paths relying on manual design but also adapts to differences in leaf damage patterns under different working environments, significantly improving detection efficiency and defect identification accuracy, while reducing computational resource consumption by optimizing image resolution.

[0120] In some embodiments, the loss function used to train the improved YOLOv11 detection network includes a damage / defect classification loss and a bounding box loss, wherein the damage / defect classification loss is related to the sampling probability of each part corresponding to each working environment type.

[0121] Specifically, the damage / defect classification loss can be:

[0122] ,

[0123] in, Classify losses for damage and defects. The total number of samples, The total number of components. The total number of damage / defect categories. Let be the sampling probability of the j-th component in the i-th sample. The label is a binary label. If the actual damage / defect category of the j-th component in the i-th sample belongs to category c, the label is 1; otherwise, it is 0. Let c be the probability that the damage / defect category of the j-th component in the i-th sample belongs to category c.

[0124] The above formula employs a three-layer nested summation structure: the outer layer iterates through samples, the middle layer iterates through components, and the inner layer iterates through damage categories, ensuring that each category prediction for each component in each sample is evaluated independently. Sampling probability is used as a weight to dynamically adjust the contribution of different components to the total loss—if a component is assigned a higher sampling probability due to its working environment, its classification error will dominate the gradient update direction, guiding the model to prioritize optimizing the identification ability of high-risk components.

[0125] The bounding box loss can be a CIoU (Complete Intersection over Union) loss function.

[0126] The total loss is a weighted sum of the damage / defect classification loss and the bounding box loss.

[0127] Step 150: Obtain images of the wind turbine blades to be inspected and information about the working environment.

[0128] Step 160: Determine the damage detection path based on the working environment characteristics of the wind turbine blade to be detected and the damage detection paths corresponding to various working environment types.

[0129] Specifically, the similarity between the operating environment characteristics of the wind turbine blade to be detected and the corresponding operating environment characteristics of various operating environment types is calculated. Cluster analysis or a rule engine is then used to match the closest operating environment type. The damage detection path corresponding to the closest operating environment type is then used as the damage detection path for the wind turbine blade to be detected.

[0130] Step 170: Based on the damage detection path and the improved YOLOv11 detection network, perform damage and defect detection on the wind turbine blade to be detected.

[0131] The following section, based on experiments, explains the beneficial effects of the improved YOLOv11 detection network.

[0132] The existing YOLOv11 detection network was used as a comparison model, trained for 200 epochs to achieve optimal accuracy. The initial learning rate was set to 0.001, the training image size was set to 640×640, the batch size was set to 64, and the number of worker threads was 12. The SGD optimizer was used for training, with automatic mixed-precision training enabled, and adjustments were made continuously based on the training results. A rich library of tools was installed using Anaconda to facilitate algorithm development and model training.

[0133] Experimental results are as follows Figure 4 , Figure 5As shown, in the early stages of training (the first few dozen epochs), the accuracy and recall rates increase significantly faster than the comparison model. This indicates that the improved YOLOv11 detection network can learn effective features from the data more quickly, reducing the training time required to achieve higher performance. During training, the accuracy and recall rates of the improved YOLOv11 detection network fluctuate less, especially in the later stages of training, where the curves are smoother. This suggests that the improved YOLOv11 detection network is more stable during training and less affected by data noise or hyperparameter adjustments.

[0134] Figure 5 This is a schematic diagram of a vision-based wind turbine blade damage and defect detection system according to some embodiments of this specification, such as... Figure 5 As shown, a vision-based wind turbine blade damage and defect detection system may include a data acquisition module, a map construction module, a model building module, a path optimization module, and a defect detection module.

[0135] The data acquisition module is used to acquire the working environment characteristics and damage defects of multiple sample wind turbine blades, where the working environments of the multiple sample wind turbine blades are different;

[0136] The graph construction module is used to construct a damage association knowledge graph based on the working environment characteristics and damage defects of multiple sample wind turbine blades. The damage association knowledge graph is used to record the damage association relationships of multiple components of wind turbine blades corresponding to various working environment types.

[0137] The model building module is used to construct and train an improved YOLOv11 detection network, wherein the improved YOLOv11 detection network includes an adaptive input unit, which is used to adjust the resolution of the image based on the image features and the damage association knowledge graph.

[0138] The path optimization module is used to construct damage detection paths corresponding to various working environment types based on the damage association knowledge graph and the improved YOLOv11 detection network. The damage detection path is used to determine the detection order of multiple components of the wind turbine blade.

[0139] The data acquisition module is also used to acquire images of the wind turbine blades to be detected and information about their working environment;

[0140] The defect detection module is used to determine the damage detection path based on the working environment characteristics of the wind turbine blade to be detected and the damage detection paths corresponding to various working environment types.

[0141] The defect detection module is also used to detect damage and defects in the wind turbine blades to be tested based on the damage detection path and the improved YOLOv11 detection network.

[0142] For a more detailed description of the vision-based wind turbine blade damage and defect detection system, please refer to the relevant description of the vision-based wind turbine blade damage and defect detection method, which will not be repeated here.

[0143] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A vision-based method for detecting damage and defects in wind turbine blades, characterized in that, include: The working environment characteristics and damage defects of multiple sample wind turbine blades were obtained, and the working environments of the multiple sample wind turbine blades were different; Based on the working environment characteristics and damage defects of multiple sample wind turbine blades, a damage association knowledge graph is constructed. The damage association knowledge graph is used to record the damage association relationships of multiple components of wind turbine blades corresponding to various working environment types. An improved YOLOv11 detection network is constructed and trained, wherein the improved YOLOv11 detection network includes an adaptive input unit, which is used to adjust the resolution of the image based on the image features and the damage association knowledge graph of the image; Based on the damage association knowledge graph and the improved YOLOv11 detection network, damage detection paths corresponding to various working environment types are constructed. The damage detection paths are used to determine the detection order of multiple components of the wind turbine blade. Acquire images and working environment information of the wind turbine blades to be inspected; Based on the working environment characteristics of the wind turbine blades to be tested and the damage detection paths corresponding to various working environment types, the damage detection paths are determined. Damage and defect detection of wind turbine blades is performed based on the damage detection path and the improved YOLOv11 detection network.

2. The vision-based wind turbine blade damage defect detection method according to claim 1, characterized in that, Based on the working environment characteristics and damage defects of multiple sample wind turbine blades, a damage association knowledge graph is constructed, including: Based on the working environment characteristics of multiple sample wind turbine blades, multiple sample wind turbine blades are clustered to determine multiple blade classes, where each blade class corresponds to a working environment type. For each blade class, based on the damage defects of multiple sample wind turbine blades included in the blade class, the damage correlation coefficient between any two components of the wind turbine blade is determined. Based on the damage correlation coefficient between any two components of the wind turbine blade corresponding to each blade class, the damage-related components of each component corresponding to the blade class are determined, and a sub-map corresponding to the blade class is constructed. A damage association knowledge graph is constructed based on the subgraph corresponding to each blade class.

3. The vision-based wind turbine blade damage defect detection method according to claim 2, characterized in that, The adaptive input unit adjusts the image resolution based on image features and a damage association knowledge graph, including: The image is used to locate the target and extract the image features, wherein the image features include at least background complexity, target density and average target size; Determine the scene complexity of an image based on its image features; Identify the component corresponding to the image; Obtain the damage and defect detection results of the component corresponding to the image; Based on the scene complexity of the image and the damage and defect detection results of the corresponding component, the image resolution is adjusted.

4. The vision-based wind turbine blade damage defect detection method according to claim 3, characterized in that, Based on a damage association knowledge graph and an improved YOLOv11 detection network, damage detection paths are constructed for various working environment types, including: For each type of working environment, the sampling probability of each component is determined based on the sub-map corresponding to the blade class and the damage defects of multiple sample wind turbine blades included in the blade class. Based on the sampling probability of each component and the damage association knowledge graph, the population is initialized. The population includes multiple individuals, and each individual corresponds to a damage detection path. Construct a fitness function, where the fitness function is related to the damage detection efficiency and the total number of pixels for each individual. Multiple sample wind turbine blades are obtained by sampling from multiple sample wind turbine blades included in the blade class. For each individual, the improved YOLOv11 detection network generates damage detection results for each sampled wind turbine blade based on the individual's corresponding damage detection path. Based on the damage detection results and fitness function of each sampled wind turbine blade, the fitness value of each individual is calculated. Based on the fitness value of each individual, selection, crossover, and mutation operations are performed to update the population and iteratively optimize it until the termination condition is met.

5. The vision-based wind turbine blade damage defect detection method according to claim 4, characterized in that, Based on the sub-maps corresponding to the blade class and the damage defects of multiple sample wind turbine blades included in the blade class, the sampling probability of each component is determined, including: Based on the damage defects of multiple sample wind turbine blades included in the blade class, the probability of damage defects for each component is determined. Based on the sub-map corresponding to the blade class, and based on the damage and defect probability of the component and the damage and defect probability of the component associated with the component, the sampling probability of each component is determined.

6. The vision-based wind turbine blade damage defect detection method according to claim 5, characterized in that, Population initialization is performed based on the sampling probability of each component and the damage association knowledge graph, including: Construct path constraints, which include at least the following: adjacent components in the damage detection path are damage-related components, component repetition constraints, and component coverage constraints. For each individual, generate the current probability of each component, and generate the individual based on the current probability and sampling probability of each component and the path constraints.

7. The vision-based wind turbine blade damage defect detection method according to any one of claims 1-6, characterized in that, The intermediate layer of the improved YOLOv11 detection network includes an initial feature fusion unit, an iterative feature fusion unit, a multi-scale interaction unit, and a feature refinement and output unit. The initial feature fusion unit scales the multi-scale feature maps output by the backbone of the improved YOLOv11 detection network to generate expanded feature maps. The iterative feature fusion unit generates three sets of bidirectionally fused feature maps based on the expanded feature maps. The multi-scale interaction unit performs cross-scale feature interaction and symmetric scale transformation on the three sets of bidirectionally fused feature maps to generate processed multi-scale feature maps. The feature refinement and output unit performs multi-level refinement and scale expansion on the processed multi-scale feature maps to generate five sets of feature maps.

8. The vision-based wind turbine blade damage defect detection method according to any one of claims 1-6, characterized in that, The improved YOLOv11 detection network's detection head includes a feature preprocessing unit, a spatial adaptive upsampling unit, a dual-path feature fusion unit, a four-level progressive feature fusion network, and a three-level optimized feature map output layer. The feature preprocessing unit dynamically adjusts the receptive field and enhances the features of the five sets of feature maps output from the intermediate layers of the improved YOLOv11 detection network, generating a semantically unified basic feature map. The spatial adaptive upsampling unit upsamples the basic feature map to generate an upsampled feature map. The dual-path feature fusion unit fuses the basic feature map and the upsampled feature map to generate a semantically focused feature map. The four-level progressive feature fusion network generates multi-level enhanced feature maps based on the semantically focused feature map. The three-level optimized feature map output layer generates a fused feature map based on the multi-level enhanced feature maps.

9. The vision-based wind turbine blade damage defect detection method according to claim 5, characterized in that, The loss function used to train the improved YOLOv11 detection network includes a damage / defect classification loss and a bounding box loss, wherein the damage / defect classification loss is related to the sampling probability of each part corresponding to each working environment type.

10. A vision-based wind turbine blade damage and defect detection system, characterized in that, The vision-based wind turbine blade damage defect detection method according to claim 1 includes: The data acquisition module is used to acquire the working environment characteristics and damage defects of multiple sample wind turbine blades, where the working environments of the multiple sample wind turbine blades are different; The graph construction module is used to construct a damage association knowledge graph based on the working environment characteristics and damage defects of multiple sample wind turbine blades. The damage association knowledge graph is used to record the damage association relationships of multiple components of wind turbine blades corresponding to various working environment types. The model building module is used to construct and train an improved YOLOv11 detection network, wherein the improved YOLOv11 detection network includes an adaptive input unit, which is used to adjust the resolution of the image based on the image features and the damage association knowledge graph. The path optimization module is used to construct damage detection paths corresponding to various working environment types based on the damage association knowledge graph and the improved YOLOv11 detection network. The damage detection path is used to determine the detection order of multiple components of the wind turbine blade. The data acquisition module is also used to acquire images of the wind turbine blades to be detected and information about their working environment; The defect detection module is used to determine the damage detection path based on the working environment characteristics of the wind turbine blade to be detected and the damage detection paths corresponding to various working environment types. The defect detection module is also used to detect damage and defects in the wind turbine blades to be tested based on the damage detection path and the improved YOLOv11 detection network.