Wind power blade intelligent diagnosis method and system based on large model visual semantic fusion
Through the method of visual semantic fusion of large models, combined with image and text features, the problem of low efficiency and insufficient accuracy in the existing technology is solved, and efficient and accurate fault identification and diagnosis is achieved.
Patent Information
- Application Number
- CN202510216609.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-18
AI Technical Summary
The existing methods of wind turbine blade fault diagnosis rely on manual inspection and sensor monitoring, which are inefficient and inaccurate, making it difficult to accurately identify subtle faults on the blade surface.
A method based on visual semantic fusion of large models is adopted, and the blade historical images are obtained for text annotation and encoding, combined with image and text features for feature fusion, and a multi-head attention mechanism and classification model are used for fault diagnosis.
It improves the accuracy and efficiency of fault diagnosis, realizes efficient identification of blade cracks, corrosion, and fall off, reduces dependence on manual inspection, and reduces operation and maintenance costs.
Smart Images

Figure CN120339671A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wind power generation, and particularly to an intelligent diagnosis method and system for wind turbine blades based on large model visual-semantic fusion. Background Art
[0002] Wind turbine blades are an important part of wind power generation units. During operation, they are vulnerable to various environmental factors, resulting in faults such as cracks, corrosion, and peeling on the blade surface. Traditional fault diagnosis methods mainly rely on manual inspections and sensor monitoring. Manual inspections require frame-by-frame analysis of images, which is inefficient and highly susceptible to the experience and subjective factors of the analysts. Moreover, it is difficult to accurately identify the types of faults based on data such as blade vibration and stress monitored by sensors, especially surface faults of the blades. Therefore, how to achieve efficient diagnosis of wind turbine blade faults has become an urgent problem to be solved.
[0003] In recent years, with the development of artificial intelligence technology, fault diagnosis methods based on multi-modal large models have gradually become a research hotspot. Most existing visual diagnosis methods rely on single image analysis and lack comprehensive utilization of semantic understanding of blade faults and context information. Based on the above background, the present invention proposes an intelligent diagnosis method and system for wind turbine blades based on large model visual-semantic fusion using a multi-modal fusion model. This method utilizes the multi-modal fusion ability of the large model and combines blade surface images and fault text descriptions to achieve more accurate and intelligent fault detection and diagnosis. By introducing semantic information, the system can better understand the types of faults, thereby improving the accuracy and reliability of diagnosis and providing strong support for the health management of wind turbines. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an intelligent diagnosis method and system for wind turbine blades based on large model visual-semantic fusion, which can solve the problems of low diagnosis efficiency and insufficient diagnosis accuracy in wind turbine blade fault diagnosis.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides an intelligent diagnosis method for wind turbine blades based on large model visual-semantic fusion, including:
[0008] Obtaining historical images of a target wind turbine blade and performing a first text annotation operation on the historical images;
[0009] Performing a first encoding and a second encoding on the historical images of the target wind turbine blade;
[0010] The first encoding is the encoding for historical images, and the second encoding is the encoding for the text after the first text annotation operation;
[0011] Perform feature fusion on the first encoding and the second encoding result, and use the feature fusion result as the input of the first classification model to train the first classification model;
[0012] Obtain the real-time image of the target wind turbine blade, and combine it with the first classification model to perform fault diagnosis on the wind turbine blade.
[0013] As a preferred solution of the intelligent diagnosis method for wind turbine blades based on large model vision-semantic fusion according to the present invention, wherein: the first text annotation operation on the historical image includes:
[0014] Perform a first classification on the historical image, and the first classification is used to divide the image into a first state and a second state;
[0015] Perform a first text annotation operation on the image after the first classification respectively;
[0016] The first text annotation operation includes non-fault text annotation for the first state and fault text annotation for the second state.
[0017] As a preferred solution of the intelligent diagnosis method for wind turbine blades based on large model vision-semantic fusion according to the present invention, wherein: the first encoding and the second encoding of the historical image of the target wind turbine blade include:
[0018] The first encoding includes image vector encoding, image position encoding, and image embedding vector encoding of the historical image;
[0019] The second encoding includes text vector encoding of the text after the first text annotation operation.
[0020] As a preferred solution of the intelligent diagnosis method for wind turbine blades based on large model vision-semantic fusion according to the present invention, wherein: the fault text annotation for the second state includes at least one of the following: fault type, fault location, fault degree, and fault cause.
[0021] As a preferred solution of the intelligent diagnosis method for wind turbine blades based on large model vision-semantic fusion according to the present invention, wherein: the first classification model is any model with the feature fusion result as the input and the fault diagnosis result or relevant parameters that can directly or indirectly obtain the fault diagnosis result as the output.
[0022] As a preferred solution of the intelligent diagnosis method for wind turbine blades based on large model visual-semantic fusion according to the present invention, wherein: the feature fusion of the first encoding and the second encoding result includes:
[0023] Obtain the features of the first encoding and the second encoding result;
[0024] Project the features of the first encoding and the second encoding result into multiple subspaces respectively;
[0025] Perform feature fusion in combination with the multi-head attention mechanism.
[0026] As a preferred solution of the intelligent diagnosis method for wind turbine blades based on large model visual-semantic fusion according to the present invention, wherein: the feature fusion in combination with the multi-head attention mechanism includes:
[0027] Calculate the attention scores and weighted fusion for each head in the multi-head attention mechanism;
[0028] Concatenate and linearly transform the outputs of all heads, add residual connections and layer normalization;
[0029] The fused comprehensive feature vector contains the dual information of the first encoding and the second encoding;
[0030] Map the comprehensive feature vector to the category space through a fully connected layer, and convert the classification score into a probability distribution:
[0031] Use the cross-entropy loss function to calculate the difference between the model prediction value and the true label, and by minimizing the loss function, select the category with the highest probability as the prediction result.
[0032] In a second aspect, the present invention provides an intelligent diagnosis system for wind turbine blades based on large model visual-semantic fusion, which is characterized in that it includes:
[0033] An annotation module for obtaining historical images of target wind turbine blades and performing a first text annotation operation on the historical images;
[0034] An encoding module for performing a first encoding and a second encoding on the historical images of the target wind turbine blades;
[0035] The first encoding is the encoding for the historical image, and the second encoding is the encoding for the text after the first text annotation operation;
[0036] A model establishment module for performing feature fusion on the first encoding and the second encoding results, and using the feature fusion result as the input of the first classification model to train the first classification model;
[0037] A diagnostic module for acquiring real-time images of target wind turbine blades and performing fault diagnosis on the wind turbine blades in combination with the first classification model.
[0038] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described above are implemented.
[0039] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes an intelligent diagnosis method and system for wind turbine blades based on large model vision-semantic fusion, acquires historical images of target wind turbine blades, and performs a first text annotation operation on the historical images; performs a first encoding and a second encoding on the historical images of the target wind turbine blades; performs feature fusion on the results of the first encoding and the second encoding, and uses the feature fusion result as the input of the first classification model to train the first classification model; acquires real-time images of target wind turbine blades, and performs fault diagnosis on the wind turbine blades in combination with the first classification model. This method combines image and text information, improving the accuracy and efficiency of fault diagnosis. By introducing a multi-modal fusion model, the system can make full use of the semantic information of the blade surface image and the fault text description to achieve more refined fault detection. In addition, this method also improves the intelligent level of diagnosis through feature fusion and classification model training, providing strong support for the health management of wind turbines. Through the multi-modal feature mining of combining images and texts, efficient identification of common faults such as blade cracks, corrosion, and shedding is achieved. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of the method for the intelligent diagnosis method of wind turbine blades based on large model vision-semantic fusion provided by an embodiment of the present invention;
[0043] Figure 2 It is an internal structure diagram of a computer device for the intelligent diagnosis method of wind turbine blades based on large model vision-semantic fusion provided by an embodiment of the present invention. Detailed Embodiments
[0044] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0045] Example 1
[0046] Reference Figure 1 - Figure 2 , which is the first embodiment of the present invention, provides a wind turbine blade intelligent diagnosis method based on large model visual semantic fusion, including:
[0047] In the existing related technologies, wind turbine blade fault diagnosis mainly relies on manual inspection and sensor monitoring, but both methods have obvious limitations. Manual inspection is not only inefficient, but the diagnostic results are easily affected by the analyst's experience and subjective factors, resulting in insufficient diagnostic accuracy. Although sensor monitoring can obtain blade vibration, stress and other data in real time, these data are often difficult to accurately reflect the fault type on the blade surface, especially fine cracks, corrosion and other problems.
[0048] This application provides a method that can effectively solve the above-mentioned problems. Next, we will combine multiple embodiments to explain in detail how to implement the wind turbine blade intelligent diagnosis method based on large model visual semantic fusion;
[0049] Figure 1 The method flow chart of the wind turbine blade intelligent diagnosis method based on large model visual semantic fusion is shown, including:
[0050] S101, acquiring a historical image of a target wind turbine blade, and performing a first text annotation operation on the historical image;
[0051] In an optional embodiment, the historical images of the target wind turbine blades can be obtained through a monitoring system of the wind farm or a special image acquisition device. These historical images contain status information of the wind turbine blades in different time periods and different working environments, which are important basis for subsequent fault diagnosis.
[0052] In an optional embodiment, the historical images of the target wind turbine blades may be high-definition historical images acquired by a drone, a fixed camera or other acquisition equipment. These historical images have high resolution and clarity and can accurately reflect the detailed information of the blade surface.
[0053] In an alternative embodiment, the first text annotation operation on historical images can be performed by existing technologies, such as deep learning-based image recognition technologies, to automatically classify and annotate blade faults in the images. To improve the accuracy and efficiency of annotation, an image classification model can be pre-trained to enable it to recognize different fault types of blades, such as cracks, corrosion, detachment, etc. Then, this model is used to automatically classify and annotate the historical images to generate a text description containing fault information. In this way, the fault features in the images can be extracted in text form, providing a data basis for subsequent multimodal fusion. However, there may still be cases of misannotation or missed annotation in this method.
[0054] In an alternative embodiment, the first text annotation operation on historical images can also be performed manually. Experienced wind turbine blade detection experts carefully analyze the historical images and manually annotate the fault types, locations, degrees, and possible causes of the blades. Although this method takes a longer time, it can ensure the accuracy and integrity of the annotation, providing reliable data support for subsequent feature fusion and model training. After the annotation is completed, the generated fault text description is associated and stored with the historical images for convenient subsequent processing and analysis.
[0055] In the embodiment of the present application, the first text annotation operation on historical images includes:
[0056] Performing a first classification on the historical images, where the first classification is used to divide the images into a first state and a second state;
[0057] Performing a first text annotation operation on the images after the first classification respectively;
[0058] The first text annotation operation includes non-fault text annotation for the first state and fault text annotation for the second state.
[0059] In the embodiment of the present application, the first state can be an image in a non-fault state, and the second state can be an image in a fault state. When performing non-fault text annotation, it can be annotated as a non-fault type, such as "normal", "no abnormality", etc.; when performing fault text annotation, according to the actual fault situation of the blade, the fault type, location, degree, and possible cause are detailedly annotated, such as "crack: at the blade edge, length 10 cm, slight, possibly caused by long-term wear", etc. Such an annotation method not only provides the fault information of the blade but also provides rich semantic information for subsequent feature fusion and model training, helping to improve the accuracy and efficiency of fault diagnosis.
[0060] It should be noted that in the embodiments of the present application, the fault text annotation for the second state includes at least one of the following: fault type, fault location, fault degree, and fault cause. Therefore, it is necessary to annotate a certain fault image for one or more of the fault type, fault location, fault degree, and fault cause. For example, in the present application, "slight damage to the coating in the blade leading edge mold joint area", "oil stain corrosion on the SS surface of the blade", "slight lightning strike damage on the PS surface of the blade", and "no fault" are used.
[0061] In fact, in the present application, for more accurate identification, a method of annotating all four types is adopted, that is, annotating all of the fault type, fault location, fault degree, and fault cause, so as to ensure a comprehensive understanding and accurate description of the fault information. In this way, the system can more accurately capture the characteristics of blade faults, providing a solid foundation for subsequent feature fusion and model training.
[0062] It should be noted that obtaining the historical images of the target wind turbine blade and performing the first text annotation operation on the historical images can improve the accuracy and efficiency of model training. By performing detailed text annotation on the historical images, the system can learn various characteristics and patterns of blade faults, so as to more accurately identify faults in subsequent real-time image diagnosis. In addition, this annotation method also provides rich semantic information for the model, helping the model better understand the type, location, and degree of faults, and further improving the accuracy and reliability of diagnosis. Therefore, this step is a key link in the method of the present invention and is of great significance for realizing the efficient and accurate diagnosis of wind turbine blade faults.
[0063] S102, perform the first encoding and the second encoding on the historical images of the target wind turbine blade;
[0064] In the embodiments of the present application, the first encoding is the encoding for the historical images, and the second encoding is the encoding for the text after the first text annotation operation;
[0065] In the embodiments of the present application, performing the first encoding and the second encoding on the historical images of the target wind turbine blade includes:
[0066] The first encoding includes image vector encoding, image position encoding, and image embedding vector encoding for the historical images;
[0067] The second encoding includes text vector encoding for the text after the first text annotation operation.
[0068] In an alternative embodiment, the first encoding may include a process of extracting and encoding image features of historical images. This generally involves using deep learning algorithms, such as convolutional neural networks (CNNs), to automatically identify and extract key features in the historical images. These features may include information such as the shape, texture, and color of the blades, which have important reference value for subsequent fault diagnosis. Through encoding, these image features are converted into a numerical form that can be understood and processed by a computer, providing a data basis for subsequent feature fusion and model training.
[0069] In an alternative embodiment, the first encoding may also use pre-trained models in a deep learning framework, such as ResNet, VGG, etc., to extract high-level semantic features of images. These pre-trained models have been trained on a large number of image datasets and can learn rich image feature representations, which helps to improve the accuracy and efficiency of fault diagnosis.
[0070] It should be noted that the reason for not using the above algorithms in this application is to consider computational resources and time costs, as well as the customized requirements for the fault characteristics of wind turbine blades. In the present invention, this application adopts a more efficient and targeted encoding method, which reduces the computational complexity while ensuring the accuracy of feature extraction and improves the real-time performance of the model. In addition, considering the particularity of wind turbine blade faults, this application designs specific encoding rules that can more accurately capture the key features of blade faults and improve the accuracy of fault diagnosis.
[0071] In the embodiment of this application, to perform vector encoding on the dataset images, that is, the first encoding, image vector encoding is required, and the specific steps are as follows:
[0072] Adjust the size of the collected wind turbine blade images to a fixed size. Assume that the blade image is represented by the letter X, that is The height of the image is H, the width is W, and the number of channels is C. Then the image is divided into N fixed-size image patches, and the size of each image patch is P*P. It can be deduced that the number of image patches N is: Each image patch is flattened into a vector: Then, through linear projection, each image patch is mapped to a D-dimensional vector v n :
[0073] v n = Ex n + b
[0074] Where is the projection matrix, is the bias term, is the embedding vector of the nth image patch.
[0075] It is also necessary to encode the image positions, specifically as follows:
[0076] After segmenting the wind turbine blade image into N fixed-size image patches, this application also needs to add position encoding to them to retain the spatial position information of the image patches.
[0077] It is also necessary to encode the image embedding vectors, specifically as follows:
[0078] The sequence composed of the embedding vectors of an image patch can be expressed as:
[0079] The sequence composed of the embedding vectors of all image patches:
[0080] Add a learnable classification token before the sequence which can be used for the final classification task of wind turbine blade faults:
[0081] In an optional embodiment, the second encoding is a process of encoding the text description generated after the first text annotation operation. This usually involves natural language processing (NLP) techniques such as word embedding or sentence embedding to convert the text description into numerical vectors. These numerical vectors can capture semantic information in the text, such as fault type, location, degree, etc., providing rich semantic features for subsequent feature fusion. Through the second encoding, the fault information in the text description is effectively converted into a numerical form, facilitating computer processing and analysis.
[0082] In an optional embodiment, the second encoding can also utilize deep learning algorithms such as the Transformer model to encode the text description. The Transformer model can capture the context information and semantic relationships in the text through the self-attention mechanism, thereby generating more accurate and rich text feature vectors. These feature vectors combined with the image feature vectors can more comprehensively reflect the fault information of the blade, improving the accuracy and efficiency of fault diagnosis.
[0083] It should be noted that the reason for not using the above algorithms in this application is due to considerations of computing resources and time. Although these algorithms can theoretically provide more accurate and rich feature representations, in practical applications, their computational complexity is relatively high and may not be suitable for real-time fault diagnosis scenarios. In addition, considering the actual requirements of wind turbine blade fault diagnosis, this application selects a more concise and effective feature extraction and fusion method to ensure the real-time performance and reliability of the diagnosis system. Therefore, this application adopts a more practical technical path in the second coding stage to balance the diagnostic performance and computational efficiency.
[0084] In the embodiments of this application, for vector encoding of dataset text annotation, that is, the second coding, it is necessary to perform word segmentation on the text annotation, specifically as follows:
[0085] First, perform word segmentation on the dataset text annotation. The method adopted is to perform word segmentation according to predefined rules. For example, the text "Slight damage to the coating in the mold joint area at the leading edge of the blade" will be segmented into [blade / leading edge / mold joint area / coating / slight / damage]. The segmented text sequence is converted into a corresponding word segmentation sequence [t1, t2,..., t n , where n is the sequence length.
[0086] It is also necessary to perform text vector encoding, specifically as follows:
[0087] Each word segment can be mapped to a vector of a fixed dimension where D is the embedding dimension. And introduce the matrix where V is the size of the vocabulary. For each word segment t n the vector can be expressed as:
[0088] e n = E[t n .
[0089] Therefore, the word segmentation sequence [t1, t2,..., t n is mapped to a vector sequence [e1, e2,..., e n .
[0090] It is also necessary to perform final text encoding, specifically as follows:
[0091] Assume that the position encoding of each word segment is represented by The position encoding is usually generated using sine and cosine functions and can be expressed as: where k is the dimension index. Therefore, the final output vector encoding is: l n = e n + q n .
[0092] It should be noted that, according to the first coding and the second coding methods designed in this application, not only the visual features in the image are considered, but also the semantic information in the text description is integrated, so as to achieve a comprehensive capture of the fault features of the wind turbine blade. In this way, the system can learn a richer and more accurate fault feature representation, laying a solid foundation for subsequent feature fusion and model training.
[0093] It should also be noted that performing the first coding and the second coding on the historical images of the target wind turbine blade can enable the image and text information to be represented and processed in the same feature space, providing convenience for subsequent feature fusion. By encoding the image in vector form, the system can capture the visual features of the blade, such as shape, texture, etc., which are crucial for identifying blade faults. At the same time, by encoding the text description, the system can obtain semantic information such as fault type, location, degree, etc., which provides a more detailed and accurate context for fault diagnosis. In the subsequent feature fusion step, the image features and text features will be effectively combined and jointly act on the fault diagnosis model, thereby improving the accuracy and efficiency of diagnosis. Therefore, this step is a bridge connecting the image and text information and plays an indispensable role in achieving efficient and accurate diagnosis of wind turbine blade faults.
[0094] S103, perform feature fusion on the results of the first coding and the second coding, and use the feature fusion result as the input of the first classification model to train the first classification model;
[0095] In the embodiment of this application, the first classification model is any model with the feature fusion result as the input and the fault diagnosis result or relevant parameters that can directly or indirectly obtain the fault diagnosis result as the output.
[0096] In an optional embodiment, the relevant parameters that can directly or indirectly obtain the fault diagnosis result can be the fault category probability distribution, fault location coordinates, fault degree level, and fault cause description, etc. These parameters can comprehensively reflect the fault situation of the blade and provide strong support for subsequent diagnosis and maintenance.
[0097] In an alternative embodiment, the first classification model can be constructed using deep learning algorithms such as convolutional neural networks (CNNs) or Transformer models. These models have powerful feature extraction and classification capabilities and can learn the key information of blade faults from the fused features. During training, the model continuously adjusts its internal parameters to minimize the difference between the prediction results and the true labels, thereby gradually improving the accuracy of fault diagnosis. In addition, to improve the generalization ability of the model, techniques such as data augmentation and regularization can be adopted to increase the diversity of training data and reduce the risk of overfitting of the model. Through such a training process, the first classification model can learn the mapping relationship from image and text features to fault diagnosis results, providing a reliable guarantee for subsequent practical applications.
[0098] In an alternative embodiment, the first classification model can also use transfer learning techniques. A model that has been trained on a similar task is used as a pre-trained model and then fine-tuned on the wind turbine blade fault diagnosis task of the present invention. In this way, the existing knowledge and experience can be fully utilized to accelerate the convergence speed of the model and improve the accuracy and efficiency of fault diagnosis. Transfer learning techniques are particularly suitable for scenarios with scarce labeled data such as wind turbine blade fault diagnosis. It can effectively alleviate the challenges brought by insufficient data and improve the performance of the model. During the fine-tuning process, the model will make adaptive adjustments according to the specific characteristics of wind turbine blade faults to better adapt to the new diagnosis task.
[0099] It should be noted that the reason for not establishing the first classification model in the above manner is mainly due to considerations of model complexity and computing resources. Although deep learning algorithms and transfer learning techniques can theoretically provide powerful feature extraction and classification capabilities, in practical applications, their model complexity and computing requirements are relatively high and may not be suitable for all scenarios. Especially in specific applications such as wind turbine blade fault diagnosis, considering real-time performance and resource limitations, this application needs to select a more concise and effective model establishment method. Therefore, the present invention adopts a more targeted technical path to balance diagnostic performance and computing efficiency and ensure the practicality and reliability of the diagnostic system. In this way, this application can reduce the complexity and computing requirements of the model while ensuring diagnostic accuracy, so as to better meet the needs of the actual application scenario.
[0100] In the embodiment of this application, the feature fusion of the first encoding and the second encoding result includes:
[0101] Obtain the features of the first encoding and the second encoding result, denoted as the visual feature vector v and the language description vector l;
[0102] Project the features of the first encoding and the second encoding result into multiple subspaces respectively;
[0103] Feature fusion is performed by combining the multi - head attention mechanism.
[0104] In the embodiments of the present application, performing feature fusion by combining the multi - head attention mechanism includes:
[0105] Calculate the attention scores and weighted fusion for each head in the multi - head attention mechanism;
[0106] Concatenate the outputs of all heads and perform a linear transformation, add a residual connection and layer normalization;
[0107] The fused comprehensive feature vector contains dual information of the first encoding and the second encoding;
[0108] Map the comprehensive feature vector to the class space through a fully - connected layer, and convert the classification score into a probability distribution:
[0109] Use the cross - entropy loss function to calculate the difference between the model prediction value and the true label. By minimizing the loss function, select the class with the highest probability as the prediction result.
[0110] Specifically, embed the dataset vector of the present application into the model, that is, perform feature fusion on the visual feature vector v and the language description vector l. The method of the multi - head attention mechanism is adopted.
[0111] First, project to multiple sub - spaces, project the visual feature v and the text feature l to sub - spaces respectively: where m = 1, 2, …, h represents the m - th head, is a learnable weight matrix, D represents the feature dimension, and D h = D / h.
[0112] Secondly, calculate the attention heads, calculate the attention scores and weighted fusion for each head:
[0113]
[0114] where: S (m) is the attention score matrix, is the projection of the visual feature (Query), is the projection of the text feature key (Key), A (m) is the attention weight matrix, is the scaling factor, used to prevent the attention score from being too large, resulting in unstable gradients, v ′(m) is the visual feature after weighted fusion, is the projection of the text feature (Value).
[0115] Through the above steps, the multi-head attention mechanism can capture the complex relationships between visual and text features, achieving cross-modal feature fusion.
[0116] Thirdly, concatenation and linear transformation: Concatenate the outputs of all heads and perform a linear transformation:
[0117] v ′ = Concat(v ′(1) , v ′(2) , …, v ′(h) )W o
[0118] where is the output weight matrix.
[0119] Thirdly, residual connection and layer normalization: To stabilize training, residual connections and layer normalization are usually added:
[0120] v out = LayerNorm(v′ + v)
[0121] where v′ + v is the residual connection, which can retain the information of the original features and helps alleviate the vanishing gradient problem, making the deep network easier to train.
[0122] The fused comprehensive feature vector v out contains dual information of vision and language and can more comprehensively represent the state of the blade.
[0123] Thirdly, classify the comprehensive feature vector, using a fully connected neural network in machine learning as the fault diagnosis module. The comprehensive feature vector v out is input into the classification model, that is, use a fully connected layer to map v out to the category space:
[0124] y = Wv out + b
[0125] where W is the weight parameter, b is the bias term, and y is the classification score.
[0126] Thirdly, convert the classification score y into a probability distribution through the softmax function:
[0127]
[0128] where p k represents the probability that the sample belongs to the Kth class.
[0129] Thirdly, use the cross-entropy loss function to calculate the difference between the model prediction value and the true label:
[0130]
[0131] Among them, t k is the one-hot encoding of the true label of the wind turbine blade fault.
[0132] Finally, by minimizing the loss function The model will select the category with the highest probability as the prediction result according to the rules learned from the training data, so as to classify the state of the wind turbine blade, which is convenient for the operation and maintenance personnel to understand and take corresponding measures.
[0133] It should be noted that by performing feature fusion on the first encoding and the second encoding results and using the feature fusion result as the input of the first classification model, the first classification model can be trained to make full use of the information of two different modalities, namely images and texts, to achieve a more comprehensive and accurate fault diagnosis. By combining the visual features in the images with the semantic information in the text descriptions, the system can capture various aspects of the blade faults, such as shape, texture, location, degree, etc., thus providing a richer and more detailed context for fault diagnosis. This cross-modal feature fusion method not only improves the accuracy of diagnosis, but also enhances the robustness and generalization ability of the system. In addition, using the feature fusion result as the input of the first classification model enables the model to learn the direct mapping relationship from complex features to fault diagnosis results, simplifies the diagnosis process, and improves the diagnosis efficiency. Therefore, this step is of great significance for achieving efficient and accurate diagnosis of wind turbine blade faults.
[0134] S104. Obtain the real-time image of the target wind turbine blade and perform fault diagnosis on the wind turbine blade in combination with the first classification model.
[0135] In an optional embodiment, the real-time image of the target wind turbine blade can be obtained by an image acquisition device such as a drone or a camera. These images will be input into the pre-trained first classification model, and the model will diagnose the fault condition of the blade according to the visual features in the images and the text description information fused with them. The diagnosis results can include detailed information such as fault category, location, degree, etc., and these information will be fed back to the operation and maintenance personnel in real time so that they can take corresponding measures for repair or maintenance in a timely manner. In this way, the rapid response and efficient processing of wind turbine blade faults can be achieved, improving the stability and reliability of the wind power generation system. At the same time, this fault diagnosis method based on image and text information can also reduce the dependence on traditional manual inspections, reduce the operation and maintenance costs, and improve the overall operation efficiency.
[0136] In summary, the present invention proposes an intelligent diagnosis method for wind turbine blades based on the visual-semantic fusion of large models, which obtains the historical images of the target wind turbine blades and performs the first text annotation operation on the historical images; performs the first encoding and the second encoding on the historical images of the target wind turbine blades; performs feature fusion on the results of the first encoding and the second encoding, and uses the feature fusion result as the input of the first classification model to train the first classification model; obtains the real-time images of the target wind turbine blades and combines the first classification model to diagnose the faults of the wind turbine blades. This method combines image and text information, improving the accuracy and efficiency of fault diagnosis. By introducing a multimodal fusion model, the system can fully utilize the semantic information of the blade surface images and fault text descriptions to achieve more refined fault detection. In addition, this method also improves the intelligence level of diagnosis through feature fusion and classification model training, providing strong support for the health management of wind turbines. Through the multimodal feature mining that combines images and texts, the efficient identification of common faults such as blade cracks, corrosion, and shedding is realized.
[0137] Embodiment 2
[0138] In a preferred embodiment, it is assumed that the input image shows slight damage in the mold joint area of the blade leading edge, and the specific process is as follows:
[0139] Step 1: Visual feature extraction (i.e., the first encoding operation part): The image extracts the visual feature vector: v = [0.2, 0.8, 0.1,..., 0.5].
[0140] Step 2: Text description generation (i.e., the second encoding operation part): The language description text: "Slight damage to the coating in the mold joint area of the blade leading edge", etc. It is converted into a text vector through the model: l = [0.3, 0.7, 0.4,..., 0.6].
[0141] Step 3: Feature fusion (i.e., perform feature fusion according to the results of the first encoding and the second encoding): Use the method of multi-head attention mechanism to generate a comprehensive feature vector: v out = [v; l] = [0.2, 0.8, 0.1,..., 0.5, 0.3, 0.7, 0.4,..., 0.6].
[0142] Step 4: Fault classification (i.e., the part for diagnosis according to the first classification model): Input v out into the trained fully connected neural network model, and the model outputs the probability distribution:
[0143] Oil stain corrosion on the SS surface of the blade: 0.1;
[0144] Slight damage to the coating in the mold joint area of the blade leading edge: 0.85;
[0145] Slight lightning strike damage on the PS surface of the blade: 0.03;
[0146] No fault: 0.02;
[0147] Select the category with the highest probability, "Slight damage to the coating in the mold joint area of the blade leading edge", as the diagnostic result.
[0148] Step Five: Result Output: The multi-modal large model outputs the diagnostic result: "Slight damage to the coating in the mold joint area of the blade leading edge".
[0149] It should be noted that in practical applications, the above steps not only achieve efficient identification of wind turbine blade faults, but also provide detailed fault information, such as fault location, degree, etc., which is crucial for maintenance personnel. Through accurate diagnostic results, maintenance personnel can quickly locate the fault point and take corresponding repair measures, thus avoiding the further expansion of the fault and ensuring the stable operation of the wind power generation system. In addition, this method can also continuously monitor the health status of the blade, timely detect potential faults, and provide strong support for preventive maintenance.
[0150] Embodiment 3
[0151] In this embodiment, an intelligent diagnostic system for wind turbine blades based on large model vision-semantic fusion is also provided. Applying the method according to any one of claims 1 to 7, it is characterized in that it includes:
[0152] The annotation module is used to obtain historical images of the target wind turbine blade and perform the first text annotation operation on the historical images;
[0153] The encoding module is used to perform the first encoding and the second encoding on the historical images of the target wind turbine blade;
[0154] The first encoding is the encoding for the historical images, and the second encoding is the encoding for the text after the first text annotation operation;
[0155] The model establishment module is used to perform feature fusion on the results of the first encoding and the second encoding, and use the feature fusion result as the input of the first classification model to train the first classification model;
[0156] The diagnostic module is used to obtain real-time images of the target wind turbine blade and perform wind turbine blade fault diagnosis in combination with the first classification model.
[0157] The above-mentioned unit modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0158] This embodiment also provides a computer device, which can be a terminal, and its internal structure diagram can be as Figure 2As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it realizes an intelligent diagnosis method for wind turbine blades based on large model visual semantic fusion. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball, or touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0159] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are realized:
[0160] Obtain the historical image of the target wind turbine blade, and perform a first text annotation operation on the historical image;
[0161] Perform a first encoding and a second encoding on the historical image of the target wind turbine blade;
[0162] The first encoding is the encoding for the historical image, and the second encoding is the encoding for the text after the first text annotation operation;
[0163] Perform feature fusion on the results of the first encoding and the second encoding, and use the feature fusion result as the input of the first classification model to train the first classification model;
[0164] Obtain the real-time image of the target wind turbine blade, and perform a fault diagnosis on the wind turbine blade in combination with the first classification model.
[0165] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
[0166] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages.
[0167] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0168] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0170] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0171] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. An intelligent diagnosis method for wind turbine blades based on the visual semantic fusion of large models, characterized in that Including: Obtain historical images of the target wind turbine blade, and perform a first text annotation operation on the historical images; Perform a first encoding and a second encoding on the historical images of the target wind turbine blade; The first encoding is the encoding for the historical images, and the second encoding is the encoding for the text after the first text annotation operation; Perform feature fusion on the results of the first encoding and the second encoding, and use the feature fusion result as the input of the first classification model to train the first classification model; Obtain real-time images of the target wind turbine blade, and perform fault diagnosis on the wind turbine blade in combination with the first classification model.
2. The intelligent diagnosis method for wind turbine blades based on large model visual semantic fusion according to claim 1, wherein The performing a first text annotation operation on the historical images includes: Perform a first classification on the historical images, where the first classification is used to classify the images into a first state and a second state; Perform a first text annotation operation on the images after the first classification respectively; The first text annotation operation includes non-fault text annotation for the first state and fault text annotation for the second state.
3. The intelligent diagnosis method for wind turbine blades based on large model visual semantic fusion according to claim 2, wherein The performing a first encoding and a second encoding on the historical images of the target wind turbine blade includes: The first encoding includes image vector encoding, image position encoding, and image embedding vector encoding for the historical images; The second encoding includes text vector encoding for the text after the first text annotation operation.
4. The intelligent diagnosis method for wind turbine blades based on large model visual semantic fusion according to claim 3, wherein, The fault text annotation for the second state includes at least one of the following: fault type, fault location, fault degree, and fault cause.
5. The intelligent diagnosis method for wind turbine blades based on large model visual semantic fusion according to claim 4, wherein, The first classification model is any model with the feature fusion result as the input and the fault diagnosis result or relevant parameters that can directly or indirectly obtain the fault diagnosis result as the output.
6. The intelligent diagnosis method for wind turbine blades based on large model visual semantic fusion according to claim 5, wherein, The performing feature fusion on the results of the first encoding and the second encoding includes: Obtain the features of the results of the first encoding and the second encoding; Project the features of the results of the first encoding and the second encoding into multiple subspaces respectively; Perform feature fusion in combination with the multi-head attention mechanism.
7. The intelligent diagnosis method for wind turbine blades based on large model visual semantic fusion according to claim 6, wherein, The performing feature fusion in combination with the multi-head attention mechanism includes: Calculate the attention scores and weighted fusion for each head in the multi-head attention mechanism; Concatenate and linearly transform the outputs of all heads, add residual connections and layer normalization; The fused comprehensive feature vector contains dual information of the first encoding and the second encoding; Map the comprehensive feature vector to the category space through a fully connected layer, and convert the classification score into a probability distribution: Use the cross-entropy loss function to calculate the difference between the model prediction value and the true label, and by minimizing the loss function, select the category with the highest probability as the prediction result.
8. The intelligent diagnosis system for wind turbine blades based on the visual semantic fusion of large models, applying the method according to any one of claims 1 to 7, characterized in that, Including: An annotation module for obtaining historical images of the target wind turbine blade and performing a first text annotation operation on the historical images; An encoding module for performing a first encoding and a second encoding on the historical images of the target wind turbine blade; The first encoding is the encoding for the historical images, and the second encoding is the encoding for the text after the first text annotation operation; A model establishment module for performing feature fusion on the results of the first encoding and the second encoding, and using the feature fusion result as the input of the first classification model to train the first classification model; A diagnostic module for acquiring real-time images of a target wind turbine blade and performing fault diagnosis on the wind turbine blade in combination with the first classification model.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the intelligent diagnosis method for wind power blades based on large model vision-semantic fusion according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the intelligent diagnosis method for wind power blades based on large model vision-semantic fusion according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Fault analysis method, system and equipment for humanoid robot and medium
CN121132710A