Wind turbine surface damage detection method and system based on segmentation prior

CN122821134APending Publication Date: 2026-09-25SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611051389.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

但该方法存在局限性:需要分别训练和推理两个网络,导致流程繁琐,且分割结果与检测过程缺乏联合优化

Benefits of technology

通过跨任务交互实现风力发电机的高效定位与精准损伤识别。该方法借鉴人类视觉注意力机制——先锁定复杂场景中的目标区域再进行细节搜索——设计了并行分割分支以精准提取风力发电机轮廓信息。与传统将分割和检测视为独立任务的方法不同,本申请利用像素级分割质量信息作为精确先验知识,显式引导检测任务,从而实现两者的深度协同与信息增益。通过这种显式背景抑制策略,网络能有效过滤无人机影像中常见的无关背景噪声,迫使检测头将计算资源集中于风力发电机叶片区域,显著提升复杂背景中微小损伤的特征提取能力与检测精度。本申请提出了基于语言大模型的新框架,用于部分信息下的旋转机械可靠性评估。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821134A_ABST
    Figure CN122821134A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of wind generator surface damage detection, in particular to a wind generator surface damage detection method and system based on segmentation prior, which comprises the following steps: inputting a to-be-detected image into a trained multi-task multi-learning network model to obtain a wind generator surface damage detection result; performing multi-scale feature fusion on shallow layer features to obtain shallow layer fusion features; performing multi-scale feature fusion on global features to obtain global fusion features; performing multi-scale feature fusion on deep layer features to obtain deep layer fusion features; performing interactive feature fusion on the global fusion features to obtain interactive fusion features; performing selective global self-adaptive injection on the shallow layer fusion features, the deep layer fusion features and the interactive fusion features to obtain shallow layer injection features and deep layer injection features; and obtaining a final detection result based on the shallow layer injection features, the deep layer injection features and a segmentation mask. The application can effectively suppress background interference, thereby ensuring accurate surface damage detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of surface damage detection technology for wind turbine generators, and in particular to a method and system for surface damage detection of wind turbine generators based on segmentation prior. Background Technology

[0002] Wind turbines are expensive to build and maintain, and their operational status directly affects the efficiency and economic benefits of the power generation system. However, long-term exposure to harsh environments such as strong winds and salt spray makes wind turbines highly susceptible to various types of damage. If not repaired in time, initial defects may worsen under stress concentration, jeopardizing the structural stability of the equipment. Minor defects can reduce power generation, and in severe cases, may even lead to catastrophic structural failure. Therefore, developing efficient and reliable detection technologies to accurately identify surface damage to wind turbines is of great engineering value and practical significance for ensuring the safe and stable operation of the equipment, reducing the total life-cycle maintenance cost, and extending its service life.

[0003] Large wind turbines are typically deployed in remote areas with complex terrain and poor transportation, their blades rotating at high speeds and at high altitudes. Traditional manual inspections are not only costly and time-consuming, but also pose safety hazards. Due to over-reliance on operator judgment, these methods are prone to fatigue-related misjudgments and weather-related interference leading to missed or false alarms, making their efficiency and coverage insufficient for the needs of large wind farms. In the early stages before the widespread application of drone technology, non-destructive testing mainly relied on various signal acquisition and analysis methods such as vibration analysis, acoustic emission monitoring, ultrasonic testing, and infrared thermal imaging. While these methods can reflect structural conditions, they are heavily reliant on specialized equipment and struggle to detect subtle surface defects such as coating damage, hindering comprehensive assessments. In recent years, breakthroughs in computer vision technology have provided new avenues for these applications. Among them, object detection technology, a core research direction in this field, has consistently held a crucial position. This technology automatically identifies specific targets in images and uses bounding boxes to accurately label their spatial locations, providing a reliable basis for subsequent analysis. With the increasing maturity of drone technology equipped with high-definition imaging devices and the rapid development of deep learning in the field of object detection, wind turbine defect identification schemes based on deep learning algorithms have gradually become a research hotspot, demonstrating broad application prospects in the wind power operation and maintenance field. Most existing single-task models directly search for damage in the entire image. However, drone-captured turbine images have cluttered and wide-ranging backgrounds, with the turbine occupying only a small portion of the image. Without auxiliary localization functions, the global search is severely affected by background interference, leading to inaccurate localization, false alarms due to background interference, and insufficient sensitivity to minor damage. Rizvi et al. adopted a two-stage strategy of "segmentation and detection separation": first, pixel-level U-Net is used to perform binary segmentation of the blade region, and then the masked region is input into YOLO for defect detection to reduce background interference. However, this method has limitations: it requires separate training and inference of two networks, resulting in a cumbersome process, and the segmentation results and detection process lack joint optimization. Recently, to address the background interference problem, Feng et al. proposed an "edge cropping" strategy. They used the traditional Canny operator and Hough transform to extract blade edges and physically crop the background. However, this strategy is essentially a fragmented process of "preprocessing + detection" rather than true multi-task collaborative learning. Because its foreground extraction process relies entirely on traditional image processing algorithms, it cannot interact with or jointly optimize features with subsequent detection networks, resulting in the entire system lacking end-to-end adaptive capabilities. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this application provides a method and system for detecting surface damage of wind turbines based on segmentation priors. By utilizing the segmentation results of wind turbine components, the detection attention is guided to focus on the target area. This network can effectively suppress background interference, thereby ensuring accurate detection of surface damage.

[0005] On the one hand, a method for detecting surface damage of wind turbines based on segmentation priors is provided, including: Acquire the image of the wind turbine to be inspected; The image to be detected is input into a trained multi-task multi-learning network model to obtain the surface damage detection result of the wind turbine. The trained multi-task multi-learning network model is used to extract multi-scale features from the image to be detected, obtaining shallow features, global features, and deep features. Multi-scale feature fusion is performed on the shallow features to obtain shallow fused features; multi-scale feature fusion is performed on the global features to obtain global fused features; multi-scale feature fusion is performed on the deep features to obtain deep fused features; interactive feature fusion is performed on the global fused features to obtain interactive fused features; selective global adaptive injection is performed on the shallow fused features and interactive fused features to obtain shallow injected features; selective global adaptive injection is performed on the deep fused features and interactive fused features to obtain deep injected features; the shallow injected features and the upsampled deep injected features are concatenated along the channel dimension, and segmentation is performed on the concatenation result to obtain a segmentation mask; based on the shallow injected features, deep injected features, and segmentation mask, the final detection result is obtained.

[0006] On the other hand, a wind turbine surface damage detection system based on segmentation prior is provided, including: The acquisition module is configured to acquire an image of the wind turbine to be detected. The detection module is configured to: input the image to be detected into a trained multi-task multi-learning network model to obtain the surface damage detection result of the wind turbine; wherein, the trained multi-task multi-learning network model is used to extract multi-scale features from the image to be detected to obtain shallow features, global features, and deep features; perform multi-scale feature fusion on the shallow features to obtain shallow fused features; perform multi-scale feature fusion on the global features to obtain global fused features; perform multi-scale feature fusion on the deep features to obtain deep fused features; perform interactive feature fusion on the global fused features to obtain interactive fused features; perform selective global adaptive injection on the shallow fused features and interactive fused features to obtain shallow injected features; perform selective global adaptive injection on the deep fused features and interactive fused features to obtain deep injected features; concatenate the shallow injected features and the upsampled deep injected features along the channel dimension, perform segmentation on the concatenation result to obtain a segmentation mask; and obtain the final detection result based on the shallow injected features, deep injected features, and segmentation mask.

[0007] The above technical solution has the following advantages or beneficial effects: This method achieves efficient localization and accurate damage identification of wind turbines through cross-task interaction. Borrowing from human visual attention mechanisms—first locking onto the target region in a complex scene and then searching for details—it designs a parallel segmentation branch to accurately extract the wind turbine's contour information. Unlike traditional methods that treat segmentation and detection as independent tasks, this application utilizes pixel-level segmentation quality information as precise prior knowledge to explicitly guide the detection task, thereby achieving deep synergy and information gain between the two. Through this explicit background suppression strategy, the network can effectively filter out common irrelevant background noise in UAV imagery, forcing the detection head to concentrate computational resources on the wind turbine blade area, significantly improving the feature extraction capability and detection accuracy of minor damage in complex backgrounds. This application also proposes a novel framework based on a large language model for the reliability assessment of rotating machinery with partial information. Attached Figure Description

[0008] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application.

[0009] Figure 1 This is a flowchart of the method in Example 1; Figure 2 The neck architecture of the YOLO series typically uses the traditional FPN structure. Detailed Implementation

[0010] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0011] Wind turbines are exposed to harsh environments for extended periods, making them highly susceptible to surface damage. Accurate identification of these defects is crucial for ensuring power generation efficiency and extending equipment lifespan. Existing methods typically employ a decoupled two-stage "segmentation-detection" strategy to address complex background interference in detection images, but this approach fails to achieve interaction between features. To overcome this limitation, this application proposes an end-to-end multi-task learning network based on segmentation priors. This network synchronously completes detection and segmentation tasks within a unified framework, achieving bidirectional reinforcement between tasks by utilizing segmentation features to guide the detection process. Specifically, the architecture integrates a dedicated semantic segmentation branch, which accurately segments the wind turbine region while retaining the main target detection branch for damage identification. A three-stage progressive fusion (TSPF) network replaces the traditional feature pyramid network (FPN) structure, effectively mitigating information attenuation during cross-scale feature propagation. Furthermore, a novel loss function is innovatively introduced to dynamically balance the multi-task loss weights and resolve gradient conflict issues during training. Finally, an adaptive guidance mechanism guided by semantic quality entropy is integrated to further optimize the semantic guidance effect of the detection task. Experimental results show that this proposed method significantly outperforms the baseline YOLOv8 model in complex contexts. In the detection task, mAP@50, mAP@75, and mAP@50:95 improvements are achieved by 2.2%, 4.2%, and 3.5%, respectively; in the semantic segmentation task, mIoU and Dice coefficients reach 97.09% and 97.88%, respectively. This framework validates the effectiveness of multi-task collaboration and provides a robust solution for the intelligent maintenance of wind energy infrastructure.

[0012] Example 1 This embodiment provides a method for detecting surface damage of wind turbines based on segmentation priors; like Figure 1 As shown, the wind turbine surface damage detection method based on segmentation prior includes: S101: Acquire the image of the wind turbine to be inspected; S102: Input the image to be detected into the trained multi-task multi-learning network model to obtain the surface damage detection result of the wind turbine. The trained multi-task multi-learning network model is used to extract multi-scale features from the image to be detected, obtaining shallow features, global features, and deep features. Multi-scale feature fusion is performed on shallow features to obtain shallow fused features; multi-scale feature fusion is performed on global features to obtain global fused features; multi-scale feature fusion is performed on deep features to obtain deep fused features; and interactive feature fusion is performed on global fused features to obtain interactive fused features. Selective global adaptive injection is performed on shallow fusion features and interactive fusion features to obtain shallow injected features; selective global adaptive injection is performed on deep fusion features and interactive fusion features to obtain deep injected features. The shallow injection features and the upsampled deep injection features are concatenated along the channel dimension. The concatenation result is then segmented to obtain a segmentation mask. The final detection result is obtained based on shallow injection features, deep injection features, and segmentation mask.

[0013] Furthermore, the multi-scale feature extraction of the image to be detected, yielding shallow features, global features, and deep features, specifically includes: The CSPDark network is used to extract multi-scale features from the image to be detected. The CSPDark network includes: a first convolutional layer, a second convolutional layer, a first C2f layer, a third convolutional layer, a second C2f layer, a fourth convolutional layer, a third C2f layer, a fifth convolutional layer, a fourth C2f layer, and an SPPF layer connected in sequence. The first C2f layer output feature map The second C2f layer outputs a feature map. The third C2f layer outputs a feature map. SPPF layer output feature map ; feature map Feature map and feature map The feature set formed by these features is considered as shallow features; feature map Feature map Feature map and feature map The feature set formed by these features is considered as global features; feature map Feature map and feature map The resulting feature set is considered as deep features.

[0014] Furthermore, the internal structures of the first C2f layer, the second C2f layer, the third C2f layer, and the fourth C2f layer are consistent. The first C2f layer includes: a sixth convolutional layer, a Split layer, n Bottleneck layers, a first concatenation unit, and a seventh convolutional layer connected in sequence. Each Bottleneck layer includes: an eighth convolutional layer, a ninth convolutional layer, and a first adder connected in sequence, wherein the input of the eighth convolutional layer is also connected to the input of the first adder; The output of the Split layer and the outputs of the n Bottleneck layers are all connected to the input of the first splicing unit concat.

[0015] Furthermore, the SPPF layer includes: a tenth convolutional layer, a first max pooling layer, a second max pooling layer, a third max pooling layer, a second concatenation unit concat, and an eleventh convolutional layer connected in sequence. The outputs of the tenth convolutional layer, the first max pooling layer, and the second max pooling layer are all connected to the input of the second concatenation unit concat.

[0016] To facilitate efficient cross-scale feature interaction in the intermediate layers of the network, this application designs a shallow multi-scale feature fusion module. The core concept of the shallow multi-scale feature fusion module is to select the feature map at the intermediate scale as the spatial alignment anchor point. By unifying the feature information of the deep (low resolution) and shallow (high resolution) layers to the intermediate scale, the shallow multi-scale feature fusion module achieves balanced information fusion, enhancing the ability of the shallow detection branch to perceive small targets and local structural information.

[0017] Furthermore, the multi-scale feature fusion of shallow features to obtain shallow fused features specifically includes: Let the three input features of the shallow multi-scale feature fusion module be as follows: 、 、 The dimensions are represented as follows: , , ;in, Indicates batch size, and These represent the height and width of the image, respectively. 、 、 These represent the number of channels for the three input features; With intermediate scale features Using spatial resolution as a unified alignment benchmark, let... Then there is For resolutions higher than shallow scale features Adaptive average pooling is used to downsample it to For resolutions lower than Deep-scale features Bilinear interpolation is used to upsample it to ; The spatial dimensions remain unchanged; the spatial scale alignment process is represented as: ; ; ; in, This indicates an adaptive average pooling operation. This indicates a bilinear interpolation upsampling operation; This represents the shallow scale features after spatial scale alignment in the shallow multi-scale feature fusion module. This represents the intermediate scale features after spatial scale alignment in the shallow multi-scale feature fusion module. This represents the deep-scale features after spatial scale alignment in the shallow multi-scale feature fusion module. After completing spatial scale alignment, use Convolution maps the number of channels of the three features to a unified value. Let the number of output channels after unification be . Then we have:

[0018]

[0019]

[0020] in, , and These represent the channel mapping operations on the corresponding branches. After mapping, the three features have the same number of channels and spatial size. ; This represents the shallow-scale features after mapping. This represents the intermediate scale features after mapping. This represents the deep-scale features after mapping; Subsequently, the three aligned features are concatenated along the channel dimension to obtain the concatenated features. : ; in, .

[0021] Finally, a convolutional fusion operation is used to re-encode the concatenated features to obtain shallow fused features. ;in, This indicates a convolution fusion operation; therefore, the final output is... The convolution fusion operation involves first performing a convolution operation through a convolutional layer (Conv), and then activating the convolution result using the ReLU activation function.

[0022] To effectively integrate feature information at different scales, this application designs a global feature fusion module. The global feature fusion module aims to address the spatial inconsistencies of multi-scale feature maps, thereby achieving effective feature aggregation.

[0023] Furthermore, the multi-scale feature fusion of the global features to obtain the global fused features specifically includes: Let the four input features of the global feature fusion module be as follows: 、 、 and For input size The dimensions of the four input features of the image are represented as follows: , , , ;in, Indicates batch size, 、 、 、 These represent the number of channels for features at different levels; First, spatial scale alignment is performed on the four input features to achieve the desired spatial scale. Using spatial resolution as a unified benchmark, its spatial size is defined as... For resolutions higher than shallow scale features and mid-scale features Adaptive average pooling is used for downsampling; for resolutions lower than [specific resolution], [the following is a separate, unrelated sentence:] Deep-scale features Upsampling is performed using bilinear interpolation; The process remains unchanged; the specific process is as follows: ; ; ; ; in, This indicates the Adaptive Average Pooling operation. This indicates a bilinear interpolation upsampling operation; This represents the shallow scale features after spatial scale alignment in the global feature fusion module. This represents the mid-scale features after spatial scale alignment in the global feature fusion module. Represents the first spatially scale-aligned feature in the global feature fusion module. Scale characteristics This represents the deep scale features after spatial scale alignment in the global feature fusion module; After scale alignment, all four feature maps have consistent spatial dimensions. Its dimensions can be uniformly represented as Subsequently, the features aligned to the four scales are concatenated along the channel dimension to obtain the globally fused features. , Therefore, the dimension of the global fusion feature is .

[0024] The core principle of multi-scale feature fusion for global features lies in constructing a unified feature representation with higher information density. In the fused feature map, the feature vector corresponding to any spatial location can simultaneously integrate contextual information across multiple receptive fields—including both local high-frequency details extracted from small-scale feature maps and global semantic information captured by large-scale feature maps. Through the synergistic effect of this complementary feature, subsequent network layers can make more comprehensive and reliable judgments on object recognition and localization.

[0025] In the feature extraction process of deep neural networks, although deep feature maps have rich semantic information, their low spatial resolution may lead to the loss of fine-grained details. In order to effectively integrate features from different levels at a deeper level of the network, this application designs a deep multi-scale feature fusion module. Unlike the shallow multi-scale feature fusion module, the deep multi-scale feature fusion module uses the deep features with the lowest resolution as the scale alignment benchmark, unifying the mid-level and high-level features to the deep semantic scale, thereby forming a fused representation oriented towards large-scale targets and global context modeling.

[0026] Furthermore, the multi-scale feature fusion of deep features to obtain deep fused features specifically includes: Let the three input features of the deep multi-scale feature fusion module be as follows: The dimensions are represented as follows: , , ;in, 、 and These represent the number of channels for the three input features; With deep features Using spatial resolution as a unified alignment benchmark, let... Then there is ;because and The spatial resolution is higher than Adaptive average pooling is used to downsample them to... ,and The entity itself remains unchanged, represented as: ; ; ; in, This represents the scale-aligned first feature in the deep multi-scale feature fusion module. One characteristic, This represents the scale-aligned first feature in the deep multi-scale feature fusion module. One characteristic, This represents the scale-aligned first feature in the deep multi-scale feature fusion module. One feature; After scale alignment, the three features have the same spatial resolution. , , .

[0027] Subsequently, convolution operations were used to unify the three features to the same number of channels. : ; ; ; in, 、 and This indicates the channel mapping operation on the corresponding branch; This represents the first channel after channel mapping in the deep multi-scale feature fusion module. One characteristic, This represents the first channel after channel mapping in the deep multi-scale feature fusion module. One characteristic, This represents the first channel after channel mapping in the deep multi-scale feature fusion module. There are 1 feature; after mapping, there are 1... ; Then, the three features are concatenated along the channel dimension. ,in, The concatenated features are re-encoded using a fusion convolution operation to obtain deep fusion features. Therefore, the final output is The convolution fusion operation involves first performing a convolution operation through a convolutional layer (Conv), and then activating the convolution result using the ReLU activation function.

[0028] The beneficial effects of the above technical solution are as follows: After completing the global feature fusion, this application further designs an interactive feature fusion module to re-encode the globally fused features and generate the global guiding features required for subsequent cross-scale semantic injection. Unlike the global feature fusion module, which only performs scale alignment and channel concatenation, the interactive feature fusion module further fuses and models the concatenated high-dimensional multi-scale features through convolutional mapping and multi-layer feature interaction units, thereby enhancing the information interaction capability between features at different levels.

[0029] Furthermore, the interactive feature fusion of the global fusion features to obtain interactive fusion features specifically includes: The interactive feature fusion module first uses convolutional mapping to globally fuse features. Projecting onto a unified embedding space yields initial embedding features. , ;in, This represents a convolution operation, initially embedding features. The dimension is , Indicates the number of embedded channels; Subsequently, multiple consecutive feature interaction units, RepVGGBlock, were used to... Deep recoding is performed. The feature interaction unit RepVGGBlock enhances local context modeling and inter-channel information interaction while maintaining the same spatial size; Let the number of feature interaction units RepVGGBlock be . Then, multiple consecutive feature interaction units RepVGGBlock are used to... Deep recoding can be recursively represented as ;in, Indicates the first Each feature interaction unit; after After the interaction modeling, the recoded fusion features are obtained. ; Finally, through convolution... Mapping to interactive fusion features ;in, This represents the convolution operation. This is an interactive fusion feature.

[0030] The output channel of the interactive feature fusion module is and It consists of two parts: ; The output channel of the interactive feature fusion module is set to ; Therefore, the final output dimension of the interactive feature fusion module is .

[0031] Combination , The final output dimension is .

[0032] Furthermore, the feature interaction unit RepVGGBlock includes: The twelfth convolutional layer, the second adder, the thirteenth convolutional layer, the third adder, the fourteenth convolutional layer, the fourth adder, the fifteenth convolutional layer, the fifth adder, the sixteenth convolutional layer, and the sixth adder are connected in sequence; the input of the twelfth convolutional layer is the input of the feature interaction unit RepVGGBlock, and the output of the sixth adder is the output of the feature interaction unit RepVGGBlock; The input of the second adder is also connected to the output of the seventeenth convolutional layer, and the input of the seventeenth convolutional layer is connected to the input of the thirteenth convolutional layer. The input of the third adder is also connected to the output of the eighteenth convolutional layer, and the input of the eighteenth convolutional layer is connected to the input of the fourteenth convolutional layer; the input of the third adder is also connected to the input of the fourteenth convolutional layer. The input of the fourth adder is also connected to the output of the nineteenth convolutional layer, and the input of the nineteenth convolutional layer is connected to the input of the fifteenth convolutional layer; the input of the fourth adder is also connected to the input of the fifteenth convolutional layer. The input of the fifth adder is also connected to the output of the twentieth convolutional layer, and the input of the twentieth convolutional layer is connected to the input of the sixteenth convolutional layer; the input of the fifth adder is also connected to the input of the sixteenth convolutional layer.

[0033] Functionally, the interactive feature fusion module follows the global feature fusion module, and its input is the globally fused features that have already undergone scale alignment and channel concatenation. Therefore, this module not only performs channel compression but also interactively recodes features from different levels, enabling the full fusion of shallow spatial details, mid-level structural information, and deep semantic information in a unified feature space. Unlike directly using concatenated features for subsequent detection, the interactive feature fusion module reduces the computational burden of redundant channels and generates a more compact and discriminative global guided representation.

[0034] In the overall network, the output of the interactive feature fusion module is... This will serve as the global guiding feature for subsequent semantic injection modules, and will be fused with shallow multi-scale features. and deep multi-scale fusion features To proceed with further interaction. Specifically, The focus is on preserving fine-grained spatial information at higher resolutions. It focuses on aggregating high-level semantic information at low resolution, while This is obtained by global multi-scale feature recoding, which can provide unified contextual guidance for different detection branches. Thus, the interactive feature fusion module establishes an effective information transmission bridge between the globally fused features and the multi-scale detection branches, providing a foundation for subsequent cross-scale semantic enhancement.

[0035] After obtaining shallow multi-scale fusion features, deep multi-scale fusion features, and interactive global guidance features, a selective global adaptive injection module was further designed. Its core function is to select the corresponding semantic subspace from the global guidance features output by the interactive feature fusion module according to the needs of different detection branches, and adaptively inject it into the local multi-scale fusion features, thereby realizing the dynamic enhancement of local branches by the global context.

[0036] Furthermore, the selective global adaptive injection of shallow fusion features and interactive fusion features to obtain shallow injected features, and the selective global adaptive injection of deep fusion features and interactive fusion features to obtain deep injected features, specifically includes: The interactive feature fusion module outputs interactive fused features. ,in, Interactive fusion features The channel is divided into two semantic subspaces. Among them, the first global guiding feature Second global guiding feature ; and They are used for global semantic injection at different scales; For deep fusion features, select As a global guiding feature; For shallow fusion features, select As a global guiding feature; The selection process is represented as follows: ;in, This indicates the selection index corresponding to the current branch.

[0037] ; ; in, This indicates that the current branch is for deep fusion features; where, This indicates that the current branch is for shallow fusion features; therefore, the selective global adaptive injection module has a "selective" feature, that is, different detection branches do not directly share the complete global guiding features, but instead select the corresponding global semantic subspace for injection according to the branch requirements.

[0038] For any local branch, let its local features be... The selected global guidance feature is Local features include: shallow fusion features or deep fusion features; First, we analyze the local features respectively. and global guidance features Perform channel mapping; Local features go through Convolutional mapping is ;in, This represents a channel mapping operation for a local branch. For local embedding features; Global guidance features Each through two independent Convolutional branches generate global modulation features. and global compensation features ;in, This represents the global activation branch, used to generate modulation information for local features; This represents the global embedding branch, used to generate directly injected global semantic features; and All passed This is achieved through convolutional layers; Due to local features With global guidance features The two have different spatial resolutions, and the scaling method is adaptively selected based on their spatial dimensions.

[0039] Let the spatial dimensions of the local feature be... The spatial size of the global guiding feature is ; When the spatial resolution of local features is lower than that of global guiding features Adaptive average pooling is used to extract global modulation features. and global compensation features Downsampling to local branch scale:

[0040]

[0041] When the spatial resolution of local features is higher than or equal to that of global guiding features, bilinear interpolation is used to upsample global information to the local branch scale, and modulation weights are generated through an activation function.

[0042]

[0043] in, Indicates adaptive average pooling. This indicates bilinear interpolation upsampling. This represents the gated activation function. The gated activation function is an approximation of the sigmoid function. Activation function used to generate adaptive modulation weights for local features; ; After scale alignment is completed, local embedding features are obtained. Global modulation features after scale alignment Global compensation features after scale alignment Having the same spatial dimensions: ;in, This represents the number of output channels for the current branch. Subsequently, global semantic injection is achieved through multiplicative modulation and additive compensation to obtain injected features. : ;in, This indicates element-wise multiplication.

[0044] If the interactive feature fusion module takes shallow fused features as input, the corresponding output will be shallow injected features. ; If the interactive feature fusion module takes deep fusion features as input, the corresponding output will be deep injection features. .

[0045] formula It includes two complementary processes: on the one hand, As a global modulation weight, it is used for local features On the one hand, selective enhancement or inhibition; on the other hand... As a global compensation feature, contextual semantic information is directly injected into local branches. Therefore, the selective global adaptive injection module can not only preserve the original spatial structure information of local branches, but also dynamically adjust the local response according to the global context.

[0046] The selective global adaptive injection module is designed to inject globally guided features into local features at different scales through branch selection. Compared to simple feature concatenation or layer-by-layer addition, the selective global adaptive injection module can adaptively adjust the injection method of global information according to the spatial scale and semantic requirements of different detection branches. As a result, the network gains stronger small target detail perception in shallow branches and more comprehensive global semantic modeling ability in deep branches, thus forming a more coordinated multi-scale detection representation.

[0047] Furthermore, the shallow injection features and the upsampled deep injection features are concatenated along the channel dimension, and the concatenation result is segmented to obtain a segmentation mask, specifically including: The result of the splicing operation is used as the input fusion feature of the segmentation branch, denoted as: ,in, Indicates the batch size of the input images. Indicates the height of the input image. Indicates the width of the input image. This represents the number of channels in the input fusion feature; For the input fusion features Perform the first convolution mapping to obtain the first segmentation feature. ;in, This represents the first convolutional mapping module; the first convolutional mapping module includes a two-dimensional convolutional layer, a batch normalization layer, and an activation function layer connected in sequence, used to extract local spatial features from the input fused features and adjust the number of channels; the first segmentation feature satisfies ; For the first segmentation feature Perform double nearest neighbor upsampling to obtain the second segmentation feature. ;in, This indicates a 2x nearest neighbor upsampling operation, where the second segmentation feature satisfies... ; For the second segmentation feature Perform the first feature recoding to obtain the third segmentation feature. ;in, This represents the first feature recoding operation module; the first feature recoding operation module is implemented based on the C2f layer and is used to perform channel transformation, branch feature extraction, feature concatenation, and fusion reconstruction on the input features to enhance the expressive power of the segmentation features. The third segmentation feature satisfies: ; For the third segmentation feature Perform a second convolution mapping to obtain the fourth segmentation feature: ;in, This represents the second convolutional mapping module, which includes a two-dimensional convolutional layer, a batch normalization layer, and an activation function layer connected in sequence, and a fourth segmentation feature. satisfy ; For the fourth segmentation feature Perform double nearest neighbor upsampling to obtain the fifth segmentation feature. : ;in: ; For the fifth segmentation feature Perform the third convolution mapping to obtain the sixth segmentation feature. : ;in, This refers to the third convolutional mapping module, which includes a two-dimensional convolutional layer, a batch normalization layer, and an activation function layer connected in sequence, and a sixth segmentation feature. satisfy ; For the sixth segmentation feature Perform second feature recoding to obtain the seventh segmentation feature: ;in, This represents the second feature recoding operation, which is implemented based on the C2f layer. The seventh segmentation feature satisfies... ; For the seventh segmentation feature Performing double nearest neighbor upsampling yields the eighth segmentation feature: ;in, ; For the eighth segmentation feature Perform segmentation prediction convolution to obtain the segmentation mask. ;in, This represents the segmentation prediction convolutional module, which includes convolutional layers used to map the eighth segmentation feature to segmentation category channels; This represents the output segmentation mask. ;in, This indicates the number of segmentation categories. In this embodiment, .

[0048] The above segmentation process can be represented as a whole as follows:

[0049] in, 、 and These represent the first convolution mapping operation, the second convolution mapping operation, and the third convolution mapping operation, respectively. and These represent the first feature recoding operation and the second feature recoding operation, respectively. This indicates a 2x nearest neighbor upsampling operation; This represents the segmentation prediction convolution operation. Through the above processing, the segmentation branch restores the stitching operation results to the scale of the input image step by step, and outputs a segmentation mask with the same spatial size as the input image.

[0050] Furthermore, the final detection result obtained based on shallow injection features, deep injection features, and segmentation mask specifically includes: The shallow injection features and deep injection features are fused to obtain P3 scale fused detection features, P4 scale fused detection features and P5 scale fused detection features, respectively. The detection features and segmentation mask fused at the P3 scale are input into the first segmentation quality entropy-guided adaptive guidance mechanism module, and the detection results at the P3 scale are output. The detection features and segmentation mask fused at the P4 scale are input into the second segmentation quality entropy-guided adaptive guidance mechanism module, and the detection results at the P4 scale are output. The detection features and segmentation mask fused at the P5 scale are input into the third segmentation quality entropy-guided adaptive guidance mechanism module, and the detection results at the P5 scale are output. The detection results at the P3, P4, and P5 scales are fused to obtain the final detection result.

[0051] After obtaining shallow and deep injection features, fused detection features at three scales (P3, P4, and P5) are constructed based on upsampling, feature concatenation, C2f feature recoding, and downsampling operations. First, deep semantic information is transferred to a higher resolution scale via a top-down upsampling path, and then shallow detail information is transferred back to the deep scale via a bottom-up downsampling path, thus achieving bidirectional fusion between features at different scales.

[0052] Furthermore, the shallow injection features and deep injection features are fused to obtain P3-scale fused detection features, P4-scale fused detection features, and P5-scale fused detection features, specifically including: Let the shallow injection characteristics be: Deep injection features are ,in, Indicates shallow injection characteristics. Indicates deep injection characteristics, Indicates the batch size of the input images. and These represent the height and width of the input image, respectively. and These represent the number of channels for shallow injection features and deep injection features, respectively; First, deep injection features The feature is upsampled by 2x, adjusted from the P5 scale to the P4 scale, and then concatenated with the lateral features at the P4 scale. Afterwards, it is re-encoded through a C2f layer to obtain the P4 intermediate fused feature. ; in, Representation of feature map , This indicates a 2x upsampling operation. This indicates a channel-level concatenation operation. This represents the P4-scale C2f feature recoding operation in the top-down path. This process can be understood as upsampling deep semantic information to the P4 scale and then merging it laterally with the P4-scale features.

[0053] Subsequently, the intermediate features of P4 are further fused along the top-down path. Upsampling to the P3 scale, while simultaneously injecting deep features After two upsampling adjustments to the P3 scale, and combined with shallow injection features... Channel stitching is performed, followed by obtaining P3-scale fused detection features through the C2f layer. :

[0054] in, This represents the P3-scale C2f feature recoding operation. This step enables the P3-scale features to simultaneously fuse detailed information from shallow-injected features, mid-level structural information from P4-fused features, and high-level semantic information from deep-injected features, thereby enhancing the detection representation at high resolution scales.

[0055] After obtaining the P3-scale fused detection features, a bottom-up downsampling path is adopted. First, the P3-scale fused detection features... Downsampling is performed, the feature is adjusted to the P4 scale, and then fused with the P4 intermediate features obtained from the top-down path. Channel stitching is performed, and then the final P4-scale fused detection features are obtained through the C2f module. :

[0056] in, This indicates a 2x downsampling operation. This represents the P4-scale C2f feature recoding operation in the bottom-up path. This step backpropagates the fine-grained spatial information from the P3 scale to the P4 scale, so that the P4 features contain both spatial details and mid-to-high-level semantic information; Detection features fused at the P4 scale Downsampling was performed, the data was adjusted to the P5 scale, and then combined with deep injection features. Channel stitching is performed, and then P5-scale fusion detection features are obtained through the C2f module. :

[0057] in, This represents the P5-scale C2f feature recoding operation. This step further passes the information fused from the shallow and mid-level layers to the deep scale and fuses it with the deep-injected features, thereby enhancing the P5-scale's ability to express global semantics and large-scale targets.

[0058] To enhance the detection branch's ability to perceive target structures and boundary regions, a segmentation quality entropy-guided adaptive guidance mechanism was designed. Its core idea is to utilize the uncertainty information output by the segmentation branch to evaluate the segmentation quality and adaptively modulate the detection features based on this quality information. Unlike directly concatenating segmentation features to the detection branch, this mechanism does not unconditionally introduce segmentation information; instead, it dynamically controls the guidance strength of the segmentation prediction on the detection features based on the reliability of the segmentation prediction.

[0059] Furthermore, the working processes of the first segmentation mass entropy-guided adaptive guidance mechanism module, the second segmentation mass entropy-guided adaptive guidance mechanism module, and the third segmentation mass entropy-guided adaptive guidance mechanism module are consistent. The first segmentation mass entropy-guided adaptive guidance mechanism module specifically includes: Let the current scale fusion detection features be The segmentation mask is ; By using the Sigmoid function After processing, the segmentation probability is obtained: ; To measure the reliability of segmentation predictions, a segmentation quality map is constructed based on information entropy. For segmentation probabilities... Its pixel-wise entropy Defined as: ; when When the value is close to 0 or 1, the prediction is more certain and the entropy is lower; when... When the value approaches 0.5, the prediction uncertainty is the highest, and the entropy value is relatively high.

[0060] Therefore, the confidence level of the segmentation quality will be... Defined as: ; Achieving an overall segmentation quality map by averaging the quality confidence maps of all segmentation categories along the channel dimension:

[0061] in, Indicates the first The quality confidence map is corresponding to each segmentation category. The overall segmentation quality map is obtained by averaging the quality confidence scores of all category channels. The overall segmentation quality map is used to measure the overall predictive reliability of segmentation branches at various spatial locations.

[0062] Subsequently, the overall segmentation quality map will be generated. Adaptive adjustment to the spatial resolution of the detected features:

[0063] in, This indicates an adaptive average pooling operation.

[0064] Simultaneously, a spatial-channel attention map is generated based on the detection features at the current scale. ;in, This represents the attention generation function composed of convolutional layers, and its output is... ;

[0065] in, 、 and These represent the convolution mapping operation. The attention function uses two layers... Convolution extracts local contextual information of the detected features, and then... Convolution restores the channel dimension, and the Sigmoid function is used to generate normalized attention weights; Finally, the segmentation quality map With attention map detection Together they form an adaptive modulation factor, which enhances the detection features. ;in, This represents element-wise multiplication. These are learnable guidance strength parameters.

[0066] During the training process, Defined as a learnable scalar parameter, it is automatically updated as the network optimizes. This application sets different initial guidance strengths at different detection scales, with the P3 branch initialized to 0.8 and the P4 and P5 branches initialized to 0.5. A larger initial P3 value enhances the utilization of segmentation boundary information by high-resolution small object detection features, while smaller initial P4 / P5 values ​​prevent excessive segmentation guidance interference on deep semantic features.

[0067] This formula shows that the enhancement of detection features is constrained by both the attention of the detection branch itself and the confidence of the segmentation quality: when the segmentation prediction quality is high, The larger the value, the more the mechanism enhances the guiding role of segmentation information on detection features; when segmentation prediction is uncertain, The smaller the value, the more the guidance intensity is automatically suppressed, thereby reducing the interference of low-quality segmentation results on the detection branch.

[0068] This mechanism enables the detection branch to selectively utilize the structural information provided by the segmentation branch through "entropy quality assessment - detection attention generation - quality constraint modulation", thereby improving the discriminative ability of detection features in scenarios with blurred target boundaries, strong background interference, or dense small targets.

[0069] The adaptive segmented guidance mechanism based on quality entropy transforms the traditional static feature fusion into a dynamic intelligent guidance process based on quality assessment, significantly improving the performance and robustness of the multi-task learning framework.

[0070] Furthermore, the fusion of the first, second, and third detection results to obtain the final detection result specifically includes: The enhancement detection features at the three scales are as follows:

[0071]

[0072]

[0073] in,

[0074]

[0075]

[0076] The enhanced detection features at the three scales are then fed into the detection head for target prediction:

[0077] The detection head is the original YOLOv8 detection head, mainly composed of a bounding box regression branch and a class prediction branch. Multi-scale feature maps are input into the detection head, and the model predicts the target location and class information at different scales. Higher-resolution feature maps are primarily used for small target detection, while lower-resolution feature maps are more suitable for medium and large target detection.

[0078] Furthermore, the loss function used during training: To achieve adaptive weight balancing between detection and segmentation tasks, learnable parameters are introduced. and Construct the Segmentation Prior Guided Detection Loss (SPGD) function:

[0079] This represents the total loss function during network training, which is the joint optimization objective of the detection task loss and the segmentation task loss after adaptive weighting. This represents the loss of the detection task, used to measure the difference between the prediction results of the object detection branch and the true annotation. This represents the segmentation task loss, used to measure the difference between the pixel-level prediction results of the segmentation branch output and the true segmentation annotations. and These represent the uncertainty parameters for the detection and segmentation tasks, respectively, and are used to adaptively adjust the weights of the corresponding loss terms according to the learning difficulty of different tasks. and This is a regularization constraint term used to prevent the uncertainty parameter from increasing indefinitely and to ensure the stability of the multi-task joint optimization process.

[0080] Detection loss It is a weighted combination of classification loss and bounding box regression loss:

[0081]

[0082] Classification loss The binary cross-entropy loss (BCE) is used to constrain the consistency between the predicted class probabilities and the true class labels, and its form is as follows:

[0083] in, Indicates the number of training samples. Indicates the number of categories. and They represent the first The sample at the th The true label and predicted probability on the class.

[0084] The bounding box regression part is composed of and Together they constitute, among which CIoU loss is used to simultaneously constrain the overlap area, center point distance, and aspect ratio difference between the predicted and ground truth bounding boxes. Its calculation form is as follows:

[0085]

[0086]

[0087] in, and These represent the predicted bounding box and the ground truth bounding box, respectively. This represents the Euclidean distance between the centers of the two points. This represents the diagonal length of the smallest bounding box. 、 as well as 、 These represent the width and height of the predicted bounding box and the ground truth bounding box, respectively.

[0088] Distribution focus loss Discrete probability distribution used to optimize boundary coordinates, if the true continuous position Located at adjacent discrete positions and Between, it is represented as:

[0089] in, and These represent the target position predicted by the model falling within adjacent discrete positions. and The probability of it.

[0090] For the segmentation branch, a pixel-level cross-entropy loss function is used. To perform supervision, its expression is:

[0091] in, and These represent the height and width of the segmented output feature map, respectively. and Representing pixels The true label and its probability of being predicted as the foreground.

[0092] The core challenge of multi-task learning lies in effectively coordinating the interactions between tasks to achieve performance improvements beyond independent single-task training, i.e., creating a synergistic effect of "1+1>2". To address this challenge, an optimization strategy is proposed that integrates a loss balancing method based on weight uncertainty and CAGrad, aiming to alleviate gradient conflicts. The former dynamically adjusts the loss weights for each task, while the latter optimizes the direction of gradient updates.

[0093] To address the long-standing challenge of balancing loss terms in multi-task learning, this application employs a loss fusion scheme based on homoscedasticity uncertainty, combining the losses of detection (regression) and segmentation (classification) heads. The resulting multi-task objective is:

[0094] in, Indicates the regression loss of the detection, and The classification loss (cross entropy) represents the segmentation. Uncertainty parameter. and With network parameters Joint optimization enables adaptive task balancing during training.

[0095] To mitigate gradient conflicts between object detection and semantic segmentation during joint training, this application employs conflict-avoiding gradient descent CAGrad to encourage collaborative optimization. Instead of predicting and modifying gradients for each task individually (as in PCGrad), CAGrad searches for a single update direction within a trust region around the average gradient, thereby reducing the risk of compromising any task.

[0096] CAGrad and weight uncertainty optimize multi-task learning from different perspectives by considering gradient direction and loss function weights, respectively. Therefore, this application combines them to improve the overall performance of the network.

[0097] Since the network contains both a detection head and a segmentation head, the multi-task loss consists of two parts.

[0098] In terms of architecture, the entire network consists of four core components: a shared feature extraction backbone network, a feature fusion neck, two task-specific branches (for object detection and semantic segmentation, respectively), and an object detection module guided by segmentation information. Specifically, the network is built upon the CSPDarknet shared feature extraction backbone network. This component extracts multi-scale feature maps hierarchically, capturing both fine spatial details and deep semantic information from the input image. These features are then input into the feature fusion neck. This component first introduces three global-to-local multi-scale feature fusion modules for information integration, followed by deep feature optimization through an interactive feature fusion module. Finally, a selective global adaptive injection module is employed. The neck layer injects "global" contextual information into "local" features. Finally, by strengthening the bidirectional feature pyramid path from top to bottom and bottom to top, the neck layer achieves efficient complementarity and fusion of deep and shallow features, constructing a robust feature representation suitable for subsequent multi-task processing. For the detection head, the classification task uses binary cross-entropy (BCE) loss for supervision, while the regression task combines CIOU and DFL losses. The segmentation head is explicitly connected to the high-resolution branch of the feature fusion network. Unlike the processing method in deep neural networks where spatial information is greatly compressed and downsampled, this layer can fully preserve rich geometric details and boundary features. Starting the segmentation task at this resolution can minimize the loss of fine spatial structure—which is crucial for generating accurate pixel-level masks and avoiding edge blurring in the final output. To reconstruct spatial resolution while maintaining semantic richness, this application adopts a lateral connection strategy. Specifically, upsampled high-level features are concatenated with intermediate features of the backbone network. This operation can effectively compensate for the spatial information loss inherent in the downsampling-upsampling process, enabling the decoder to recover the fine geometric details that were originally lost in the deep network from the backbone network.

[0099] This application proposes an end-to-end multi-task learning network, called the Segmentation Prior Guided Multi-Task Learning (SPGML) framework, which can simultaneously achieve segmentation of wind turbine regions and surface damage detection.

[0100] To mitigate the inherent information loss during information transmission in traditional FPN structures, this application proposes and constructs a three-stage progressive fusion (TSPF) network. This framework achieves efficient processing of cross-scale information interaction through a continuous process of panoramic feature aggregation, cross-layer directional injection, and bidirectional feature pyramid.

[0101] To address the task conflict and imbalance issues in multi-task learning, this application proposes a segmentation prior detection loss (SPGD). This method integrates task uncertainty weights with the CAGrad strategy to balance loss weights and alleviate multi-head gradient conflicts during training.

[0102] The output of the segmentation task is used as high-quality prior knowledge to explicitly guide the detection branch. This mechanism helps the model focus on the target region more effectively, thereby suppressing background interference.

[0103] like Figure 2 As shown, the neck architecture of the YOLO series typically adopts a traditional FPN structure, which contains multiple branches for multi-scale feature fusion. Figure 2The information fusion mechanism of the standard FPN is demonstrated, in which the existing layers (layer 1, layer 2, and layer 3) are arranged from top to bottom. FPN facilitates information fusion between these different layers. When layer 1 attempts to obtain information from the other two layers, two different scenarios occur: (1) Direct access: If layer 1 needs to utilize information from layer 2, it can directly access and fuse the data. (2) Indirect recursive access: If layer 1 needs information from layer 3, it must recursively call the fusion module of the adjacent layer. Specifically, the information from layer 2 and layer 3 must be fused first, and then layer 1 can indirectly obtain the information from layer 3 by integrating the combined output of layer 2.

[0104] This transmission mode can lead to severe information loss during computation. Since inter-layer interaction is limited to exchanging selected information from intermediate layers, unselected information is discarded during transmission. To alleviate the inherent information loss problem in the recursive transmission of traditional FPN structures, this application proposes a three-stage progressive fusion framework. This framework achieves efficient processing of cross-scale information interaction by employing a sequential pipeline that includes panoramic feature aggregation, cross-layer directional injection, and a bidirectional feature pyramid.

[0105] To validate the proposed model, this application utilizes two widely adopted and complementary public datasets: DTU and Blade30. The DTU dataset, published by the Technical University of Denmark, includes high-resolution close-range UAV inspection images, providing the model with rich features of blade appearance. Conversely, the Blade30 dataset contains high-resolution images collected from 30 turbines in diverse global environments, significantly enhancing the model's adaptability to complex background scenarios. To leverage the synergistic advantages of these multi-source datasets, this application implements a joint training scheme. By fusing datasets with different damage types, imaging conditions, and environments, this strategy simulates real-world variability and provides comprehensive samples for feature learning. Therefore, this joint approach significantly improves the model's generalization ability and robustness, enabling it to effectively address various challenges in complex industrial applications.

[0106] To ensure accurate feature learning and provide a reliable data foundation, this application uses LabelMe for comprehensive damage annotation and turbine segmentation. Addressing the limitations of existing datasets in capturing the full spectrum of real-world industrial damage, this application improves and expands the labels, constructing a systematic classification framework. Based on inducing factors, visual appearance, and dynamic evolution, this framework establishes six distinct categories: Lightning strike: This is caused by extreme weather such as thunderstorms. It manifests as large black ablation areas on the blade surface, exhibiting distinct characteristics different from other damage types. Damage: This is mainly caused by external influences such as sandstorms and hail. The damage presents an irregular shape, is primarily yellowish-black, and usually only affects the surface protective layer. Rust and cracks: Rust damage appears as pale yellow spots in the initial stage of gel coating or paint peeling; crack damage is black linear cracks caused by material aging. Peeling and corrosion: Peeling damage manifests as the peeling of paint or gel coating, exposing a silvery-white metal surface, but without rust; corrosion occurs when significant rust appears in the separated area, presenting as brownish-red rust spots.

[0107] To mitigate the risk of overfitting due to insufficient dataset size and enhance the model's adaptability to complex scenes, this application introduces data augmentation techniques during the data preprocessing stage. This method not only significantly expands the size of the training samples but also simulates different shooting conditions and environmental backgrounds. Therefore, it improves the model's generalization ability, enabling it to better adapt to the variable and complex environments encountered in real-world applications. The four specific data augmentation methods employed are noise injection, brightness adjustment, geometric transformation, and mirror flipping.

[0108] For each original sample, this application applied four random data augmentations, ensuring that at least one augmentation effect was included in each iteration. This process produced a different set of augmented samples. By utilizing these augmentation techniques, the dataset was effectively expanded in terms of scale and diversity, thereby more accurately simulating the complex and varied environmental conditions encountered in real-world applications and enhancing the model's adaptability to various scenarios. Table 1 shows detailed statistics of the dataset after data augmentation.

[0109] Table 1: Distribution of Damage Target Categories in the Dataset

[0110] The experiments were conducted in a high-performance computing environment. The hardware configuration included one Intel(R) Xeon(R) Platinum 8473C CPU and four Nvidia RTX A6000 GPUs. The operating system used was Ubuntu 20.04.6LTS. The deep learning framework used was PyTorch 2.3.1, GPU acceleration was provided by CUDA 11.8, and computational performance was optimized using cuDNN 8.9.7. The input image resolution was uniformly set to 640×640 pixels; this resolution ensured a balance between model performance and computational efficiency. Detailed training parameters of the network are shown in Table 2.

[0111] Table 2: Detailed training parameters of the network

[0112] To comprehensively evaluate the damage detection performance of the model, this application selected several mainstream evaluation metrics for analysis. For the damage detection task, mean accuracy (mAP) was used as the primary metric. This application reports mAP@50 (IoU threshold 0.5), mAP@75 (IoU threshold 0.75), and mAP@50:95 (averaged over IoU thresholds from 0.5 to 0.95). mAP was calculated by plotting the mean accuracy (AP) across all N categories:

[0113] AP for each category comes from accuracy ( ) and recall rate ( ), which is defined as: , ,

[0114] in , and These represent the number of true positives, false positives, and false negatives, respectively.

[0115] For the wind turbine segmentation task, the joint mean intersection (mIoU) and Dice similarity coefficient were used. Given the binary nature of the task, these metrics provide a robust assessment of segmentation quality and boundary consistency. mIoU measures the predicted segmentation quality in class C (C=2). ) and ground reality ( Average overlap between regions:

[0116] The Dice coefficient is used to further evaluate overlap accuracy, especially for slender structures such as turbine blades:

[0117] Furthermore, this experiment uses a confusion matrix to analyze the model's predictions. Unlike image classification, this object detection matrix takes into account missed detections that are misclassified as background. This intuitively evaluates the model's performance across different damage types and locations.

[0118] Table 3 presents the model's accuracy results on damage detection and assessment metrics. Notably, the model achieved an excellent accuracy of over 96% on the important assessment metric map@50.

[0119] Table 3: Model Damage Detection and Evaluation Results

[0120] The model proposed in this application demonstrates high detection rates for fractures and cracks, and also exhibits superior performance in detecting other damage categories, maintaining high overall recognition accuracy. These results indicate that the proposed model has significant advantages in overall performance. It not only effectively identifies challenging damage types but also enhances the detection capabilities of other categories, thereby achieving more balanced and comprehensive detection results.

[0121] Through systematic analysis of the target bounding boxes output by the model, the proposed model demonstrates significant advantages in handling difficult targets. This is evident in the key categories of fractures and cracks, characterized by irregular shapes and large dimensional variations. Comprehensive analysis combining confusion matrices and detection visualization fully validates the efficiency and reliability of the proposed model in solving the problem of surface damage detection on wind turbines in real-world scenarios, showcasing its enormous potential for practical applications.

[0122] The proposed model is highly sensitive to the target detection region, but pays significantly less attention to the background region. This indicates that the proposed model can more effectively focus on important feature regions related to the damage location and successfully suppress background interference.

[0123] Table 4: Evaluation Results of Model Semantic Segmentation

[0124] As shown in Table 4, the proposed method achieves excellent results on two key metrics. Specifically, the proposed method obtains an mIoU of 0.9709 and a Dice coefficient of 0.9788. This indicates that the improved model has a stronger ability to extract semantic features and restore edge details, thereby generating a more accurate segmentation mask.

[0125] Comparative analysis of the segmentation results shows that for tasks where the target shape is regular and occupies a large proportion of the image, the model can usually achieve satisfactory segmentation results. However, when the target region is small or exhibits a slender and irregular shape, the model proposed in this application maintains stable and superior segmentation performance even in these complex scenes, demonstrating stronger robustness and significant performance advantages. These results strongly demonstrate that multi-task learning significantly improves the model's ability to discriminate targets with complex structures by effectively utilizing complementary information between detection and segmentation tasks.

[0126] To systematically evaluate the contribution of each proposed innovation, this application conducted a series of ablation experiments, progressively analyzing how individual components affect damage detection performance. Specifically, four comparative models were designed: Model A introduces a basic multi-task branch learning framework using a traditional feature pyramid network; Model B introduces a multi-task learning optimization strategy based on Model A; Model C introduces a three-stage progressive fusion network structure; and Model D, as the complete control model, combines an adaptive guidance mechanism guided by segmentation quality entropy with all the above innovations.

[0127] Table 5: Results of target detection and semantic segmentation evaluation metrics in ablation experiments.

[0128]

[0129] As shown in Table 5, although baseline model A achieved simultaneous output for multiple tasks, there was still room for improvement in various metrics. After introducing weight uncertainty modeling and the CAGrad strategy to construct model B, all evaluation metrics showed significant leaps. Specifically, the improvement in segmentation performance was particularly outstanding: mIoU increased from 0.9316 to 0.9685, and the Dice coefficient increased by approximately 1.9%. Meanwhile, the detection metric mAP@50 steadily increased from 0.939 to 0.958. This dramatic change in the data is strong evidence of severe gradient conflicts and task-dominated imbalance in the original baseline model. The addition of the optimization strategy effectively dynamically balanced the gradient magnitudes of the detection and segmentation tasks, enabling the model to avoid local optima, significantly improving convergence quality, and particularly addressing the difficulty of refining the segmentation task.

[0130] Building upon the optimization problem, Model C replaces the traditional FPN with a three-stage progressive fusion network. Experimental results demonstrate that this improvement further unlocks the model's performance potential. Model C achieves optimal performance across all experiments on the segmentation task, with mIoU and Dice increasing to 0.9756 and 0.9822, respectively. Simultaneously, mAP@75 significantly increases from 0.765 to 0.786 (an increase of 2.1%). The simultaneous increase and decrease of the mAP@75 segmentation metric indicates that the progressive fusion structure effectively preserves deep semantic information and shallow detail texture through multi-scale feature interaction. This enhanced feature representation capability enables the model to handle pixels with damaged edges more accurately, resulting in breakthroughs in both high-threshold detection metrics and pixel-level segmentation metrics.

[0131] Model D, as a complete model integrating all innovations, aims to utilize segmentation information to assist in the detection task. Experimental data show that Model D achieves peak performance on the detection task, with mAP@50, mAP@75, and mAP@50:95 reaching maximum values ​​of 0.963, 0.797, and 0.722, respectively. Notably, compared to Model C, Model D's mIoU is slightly lower ("0.9756→0.9709", a decrease of only about 0.4%). This phenomenon of "improved detection with slight segmentation regression" aligns with the expectations of multi-task collaborative learning. The main motivation for introducing the segmentation quality entropy guidance module is to feed back segmentation branch information to the detection branch to correct localization bias. Data shows that the model successfully reallocated computational resources originally used for perfect pixel segmentation to accurate regression of the target bounding box. Although the segmentation metric fluctuates slightly, it remains at an exceptionally high level. This trade-off significantly improves the accuracy of the core detection task, validating the effectiveness of the guidance module in facilitating cross-task information flow.

[0132] This application demonstrates a broad application prospect for solving key challenges such as wind turbine safety monitoring, maintenance decision-making, and operational efficiency improvement in complex environments based on an end-to-end multi-task learning network for segmentation. By establishing a collaborative mechanism between object detection and semantic segmentation, the network achieves complementary advantages at the feature level. Test results on a self-built system wind turbine damage dataset show that the proposed method performs excellently. Compared with the baseline YOLOv8 model, all metrics are significantly improved: mAP@50 increases by 2.2%, mAP@75 increases by 4.2%, and mAP@50:95 increases by 3.5%. These results fully validate the accuracy and reliability of the proposed model in the task of wind turbine surface damage detection.

[0133] Example 2 This embodiment provides a wind turbine surface damage detection system based on segmentation prior, including: The acquisition module is configured to acquire an image of the wind turbine to be detected. The detection module is configured to: input the image to be detected into a trained multi-task multi-learning network model to obtain the surface damage detection result of the wind turbine; wherein, the trained multi-task multi-learning network model is used to extract multi-scale features from the image to be detected to obtain shallow features, global features, and deep features; perform multi-scale feature fusion on the shallow features to obtain shallow fused features; perform multi-scale feature fusion on the global features to obtain global fused features; perform multi-scale feature fusion on the deep features to obtain deep fused features; perform interactive feature fusion on the global fused features to obtain interactive fused features; perform selective global adaptive injection on the shallow fused features and interactive fused features to obtain shallow injected features; perform selective global adaptive injection on the deep fused features and interactive fused features to obtain deep injected features; concatenate the shallow injected features and the upsampled deep injected features along the channel dimension, perform segmentation on the concatenation result to obtain a segmentation mask; and obtain the final detection result based on the shallow injected features, deep injected features, and segmentation mask.

[0134] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for detecting surface damage of wind turbine generators based on segmentation priors, characterized in that, include: Acquire the image of the wind turbine to be inspected; The image to be detected is input into a trained multi-task multi-learning network model to obtain the surface damage detection result of the wind turbine. The trained multi-task multi-learning network model is used to extract multi-scale features from the image to be detected, obtaining shallow features, global features, and deep features. Multi-scale feature fusion is performed on the shallow features to obtain shallow fused features; multi-scale feature fusion is performed on the global features to obtain global fused features; multi-scale feature fusion is performed on the deep features to obtain deep fused features; interactive feature fusion is performed on the global fused features to obtain interactive fused features; selective global adaptive injection is performed on the shallow fused features and interactive fused features to obtain shallow injected features; selective global adaptive injection is performed on the deep fused features and interactive fused features to obtain deep injected features; the shallow injected features and the upsampled deep injected features are concatenated along the channel dimension, and segmentation is performed on the concatenation result to obtain a segmentation mask; based on the shallow injected features, deep injected features, and segmentation mask, the final detection result is obtained.

2. The wind turbine surface damage detection method based on segmentation prior as described in claim 1, characterized in that, The process of extracting multi-scale features from the image to be detected, obtaining shallow features, global features, and deep features, specifically includes: The CSPDark network is used to extract multi-scale features from the image to be detected. The CSPDark network includes: a first convolutional layer, a second convolutional layer, a first C2f layer, a third convolutional layer, a second C2f layer, a fourth convolutional layer, a third C2f layer, a fifth convolutional layer, a fourth C2f layer, and an SPPF layer connected in sequence. The first C2f layer output feature map The second C2f layer outputs a feature map. The third C2f layer outputs a feature map. SPPF layer output feature map ; feature map Feature map and feature map The resulting feature set is considered a shallow feature set; the feature map is then... Feature map Feature map and feature map The feature set formed is considered as global features; the feature map is... Feature map and feature map The resulting feature set is considered as deep features.

3. The wind turbine surface damage detection method based on segmentation prior as described in claim 1, characterized in that, The process of fusing shallow features at multiple scales to obtain shallow fused features specifically includes: Let the three input features of the shallow multi-scale feature fusion module be as follows: 、 、 The dimensions are represented as follows: , , ;in, Indicates batch size, and These represent the height and width of the image, respectively. 、 、 These represent the number of channels for the three input features; using intermediate scale features... Using spatial resolution as a unified alignment benchmark, let... Then there is For resolutions higher than shallow scale features Adaptive average pooling is used to downsample it to For resolutions lower than Deep-scale features Bilinear interpolation is used to upsample it to ; The spatial dimensions remain unchanged; the spatial scale alignment process is represented as: ; ; ;in, This indicates an adaptive average pooling operation. This indicates a bilinear interpolation upsampling operation; This represents the shallow scale features after spatial scale alignment in the shallow multi-scale feature fusion module. This represents the intermediate scale features after spatial scale alignment in the shallow multi-scale feature fusion module. This represents the deep-scale features after spatial scale alignment in the shallow multi-scale feature fusion module. After completing spatial scale alignment, use Convolution maps the number of channels in the three feature paths to a unified number. Let the number of output channels after the mapping be... Then we have: ; ; in, , and These represent the channel mapping operations on the corresponding branches. After mapping, the three features have the same number of channels and spatial size. ; This represents the shallow-scale features after mapping. This represents the intermediate scale features after mapping. This represents the deep-scale features after mapping; subsequently, the three-way aligned features are concatenated along the channel dimension to obtain the concatenated features. : ;in, ; Finally, a convolutional fusion operation is used to re-encode the concatenated features to obtain shallow fused features. ;in, This represents the convolution fusion operation; therefore, the final output is... The convolution fusion operation involves first performing a convolution operation through a convolutional layer (Conv), and then activating the convolution result using the ReLU activation function.

4. The wind turbine surface damage detection method based on segmentation prior as described in claim 3, characterized in that, The process of fusing global features at multiple scales to obtain global fused features specifically includes: Let the four input features of the global feature fusion module be as follows: 、 、 and For input size The dimensions of the four input features of the image are represented as follows: , , , ;in, Indicates batch size, 、 、 、 These represent the number of channels for features at different levels; First, spatial scale alignment is performed on the four input features to achieve the desired spatial scale. Using spatial resolution as a unified benchmark, its spatial size is defined as... For resolutions higher than shallow scale features and mid-scale features Adaptive average pooling is used for downsampling; for resolutions lower than [specific resolution], [the following is a separate, unrelated sentence:] Deep-scale features Upsampling is performed using bilinear interpolation; The process remains unchanged; the specific process is as follows: ; ; ; ; in, This indicates an adaptive average pooling operation. This indicates a bilinear interpolation upsampling operation; This represents the shallow scale features after spatial scale alignment in the global feature fusion module. This represents the mid-scale features after spatial scale alignment in the global feature fusion module. Represents the first spatially scale-aligned feature in the global feature fusion module. Scale characteristics This represents the deep scale features after spatial scale alignment in the global feature fusion module; After scale alignment, all four feature maps have consistent spatial dimensions. Its dimensions can be uniformly represented as Subsequently, the features aligned to the four scales are concatenated along the channel dimension to obtain the globally fused features. , Therefore, the dimension of the global fusion feature is .

5. The wind turbine surface damage detection method based on segmentation prior as described in claim 4, characterized in that, The process of multi-scale feature fusion of deep features to obtain deep fused features specifically includes: Let the three input features of the deep multi-scale feature fusion module be as follows: The dimensions are represented as follows: , , ;in, 、 and These represent the number of channels for the three input features; With deep features Using spatial resolution as a unified alignment benchmark, let... Then there is ;because and The spatial resolution is higher than Adaptive average pooling is used to downsample them to... ,and The entity itself remains unchanged, represented as: ; ; ; in, This represents the scale-aligned first feature in the deep multi-scale feature fusion module. One characteristic, This represents the scale-aligned first feature in the deep multi-scale feature fusion module. One characteristic, This represents the scale-aligned first feature in the deep multi-scale feature fusion module. One feature; After scale alignment, the three features have the same spatial resolution. , , ; Subsequently, convolution operations were used to unify the three features to the same number of channels. : ; ; ; in, 、 and This indicates the channel mapping operation on the corresponding branch; This represents the first channel after channel mapping in the deep multi-scale feature fusion module. One characteristic, This represents the first channel after channel mapping in the deep multi-scale feature fusion module. One characteristic, This represents the first channel after channel mapping in the deep multi-scale feature fusion module. There are 1 feature; after mapping, there are 1... ; Then, the three features are concatenated along the channel dimension. ,in, The concatenated features are re-encoded using a fusion convolution operation to obtain deep fusion features. Therefore, the final output is The convolution fusion operation involves first performing a convolution operation through a convolutional layer (Conv), and then activating the convolution result using the ReLU activation function.

6. The wind turbine surface damage detection method based on segmentation prior as described in claim 5, characterized in that, The interactive feature fusion of the global fused features to obtain interactive fused features specifically includes: the interactive feature fusion module first uses convolutional mapping to fuse the global fused features... Projecting onto a unified embedding space yields initial embedding features. , ;in, This represents a convolution operation, initially embedding features. The dimension is , Indicates the number of embedded channels; Subsequently, multiple consecutive feature interaction units, RepVGGBlock, were used to... Perform deep recoding; Let the number of feature interaction units RepVGGBlock be . Then, multiple consecutive feature interaction units RepVGGBlock are used to... Deep recoding can be recursively represented as ;in, Indicates the first Each feature interaction unit; after After the interaction modeling, the recoded fusion features are obtained. ; Finally, through convolution... Mapping to interactive fusion features ;in, This represents the convolution operation. For interactive fusion features; the output channel of the interactive feature fusion module is composed of and It consists of two parts: The output channel of the interactive feature fusion module is set to... Therefore, the final output dimension of the interactive feature fusion module is: ; combination , The final output dimension is .

7. The wind turbine surface damage detection method based on segmentation prior as described in claim 6, characterized in that, The process of obtaining the final detection result based on shallow injection features, deep injection features, and a segmentation mask specifically includes: fusing the shallow injection features and deep injection features to obtain P3-scale fused detection features, P4-scale fused detection features, and P5-scale fused detection features, respectively; inputting the P3-scale fused detection features and the segmentation mask into a first segmentation quality entropy-guided adaptive guidance mechanism module to output the detection result at the P3 scale; inputting the P4-scale fused detection features and the segmentation mask into a second segmentation quality entropy-guided adaptive guidance mechanism module to output the detection result at the P4 scale; inputting the P5-scale fused detection features and the segmentation mask into a third segmentation quality entropy-guided adaptive guidance mechanism module to output the detection result at the P5 scale; and fusing the detection results at the P3, P4, and P5 scales to obtain the final detection result.

8. The wind turbine surface damage detection method based on segmentation prior as described in claim 7, characterized in that, The shallow injection features and deep injection features are fused to obtain P3-scale fused detection features, P4-scale fused detection features, and P5-scale fused detection features, specifically including: Let the shallow injection characteristics be: Deep injection features are ,in, Indicates shallow injection characteristics. Indicates deep injection characteristics, Indicates the batch size of the input images. and These represent the height and width of the input image, respectively. and These represent the number of channels for shallow injection features and deep injection features, respectively; First, deep injection features The feature is upsampled by 2x, adjusted from the P5 scale to the P4 scale, and then concatenated with the lateral features at the P4 scale. Afterwards, it is re-encoded through a C2f layer to obtain the P4 intermediate fused feature. ; in, Representation of feature map , This indicates a 2x upsampling operation. This indicates a channel-level concatenation operation. This represents the P4-scale C2f feature recoding operation in the top-down path; Subsequently, the intermediate features of P4 are further fused along the top-down path. Upsampling to the P3 scale, while simultaneously injecting deep features After two upsampling adjustments to the P3 scale, and combined with shallow injection features... Channel stitching is performed, followed by obtaining P3-scale fused detection features through the C2f layer. : in, This represents the P3-scale C2f feature recoding operation; After obtaining the P3 scale fusion detection features, a bottom-up downsampling path is entered. First, the P3 scale fusion detection features are processed... Downsampling is performed to adjust the feature size to the P4 scale, and then the feature is fused with the P4 intermediate features obtained from the top-down path. Channel stitching is performed, and then the final P4-scale fused detection features are obtained through the C2f module. : in, This indicates a 2x downsampling operation. This represents the P4-scale C2f feature recoding operation in the bottom-up path; fusion detection features at the P4 scale. Downsampling was performed, the data was adjusted to the P5 scale, and then combined with deep injection features. Channel stitching is performed, and then P5-scale fused detection features are obtained through the C2f module. : ; in, This indicates the P5-scale C2f feature recoding operation.

9. The wind turbine surface damage detection method based on segmentation prior as described in claim 8, characterized in that, The working processes of the first segmentation quality entropy-guided adaptive guidance mechanism module, the second segmentation quality entropy-guided adaptive guidance mechanism module, and the third segmentation quality entropy-guided adaptive guidance mechanism module are consistent. The first segmentation quality entropy-guided adaptive guidance mechanism module specifically includes: Let the current scale fusion detection features be The segmentation mask is ; By using the Sigmoid function After processing, the segmentation probability is obtained: ; For the segmentation probability Its pixel-wise entropy Defined as: ; Therefore, the confidence level of the segmentation quality will be... Defined as: ; Achieving an overall segmentation quality map by averaging the quality confidence maps of all segmentation categories along the channel dimension: ;in, Indicates the first The overall segmentation quality map is obtained by averaging the quality confidence scores of all category channels, corresponding to the quality confidence scores of each segmentation category. ; Subsequently, the overall segmentation quality map will be generated. Adaptive adjustment to the spatial resolution of the detected features: ;in, This indicates an adaptive average pooling operation; Simultaneously, a spatial-channel attention map is generated based on the detection features at the current scale. ;in, This represents the attention generation function composed of convolutional layers, and its output is... ; in, 、 and These represent the convolution mapping operation; Finally, the segmentation quality map With attention map detection Together they form an adaptive modulation factor, which enhances the detection features. ;in, This represents element-wise multiplication. These are learnable guidance strength parameters.

10. A wind turbine surface damage detection system based on segmentation prior, characterized in that, include: The acquisition module is configured to acquire an image of the wind turbine to be detected. The detection module is configured to: input the image to be detected into a trained multi-task multi-learning network model to obtain the surface damage detection result of the wind turbine; wherein, the trained multi-task multi-learning network model is used to extract multi-scale features from the image to be detected to obtain shallow features, global features, and deep features; perform multi-scale feature fusion on the shallow features to obtain shallow fused features; perform multi-scale feature fusion on the global features to obtain global fused features; perform multi-scale feature fusion on the deep features to obtain deep fused features; perform interactive feature fusion on the global fused features to obtain interactive fused features; perform selective global adaptive injection on the shallow fused features and interactive fused features to obtain shallow injected features; perform selective global adaptive injection on the deep fused features and interactive fused features to obtain deep injected features; concatenate the shallow injected features and the upsampled deep injected features along the channel dimension, perform segmentation on the concatenation result to obtain a segmentation mask; and obtain the final detection result based on the shallow injected features, deep injected features, and segmentation mask.