A method and system for detecting the state of a wind turbine blade based on multi-modal data fusion
The multi-modal data fusion approach for wind turbine blade detection enhances accuracy and convenience by integrating visible light, infrared, and vibration data, addressing inefficiencies in current human-dependent and single-modal methods.
Patent Information
- Application Number
- CN202310073745.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-01-30
AI Technical Summary
In the prior art, the detection of fan blades mainly relies on human resources, single-modal data detection or artificial intelligence-based methods, and there are problems such as high human resources consumption, insufficient accuracy, and inability to reflect internal failures.
The multimodal data fusion method is adopted, combining feature level and decision-making level fusion, and using blade visible light images, infrared images, sound and vibration signals, data is obtained through drones, infrared imaging, sound collectors and vibration sensors, and feature extraction and model training are used for attention mechanism and convolutional neural network to achieve fusion detection of multimodal data.
It improves the accuracy, convenience and real-time performance of fan blade detection, reduces labor costs and detection cycles, and achieves more accurate and timely fault detection.
Smart Images

Figure CN116123040B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wind power generation, and particularly relates to a method and system for detecting the state of a wind turbine blade based on multi-modal data fusion. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] At present, the detection of wind turbine blades is largely based on manual detection, pattern recognition, or single-modal data of blades based on artificial intelligence for detection; among them, manual detection mainly relies on the experience of technicians for detection, which is simple and effective but consumes a large amount of human resources; pattern recognition mainly determines faults by establishing a fault database and comparing the blades to be detected with the database, with a clear idea but insufficient accuracy; while the single-modal prediction based on artificial intelligence mainly obtains one modal data of the blade, such as images or sounds, and realizes prediction through artificial intelligence, which is efficient and accurate, but for some faults of the blade, single-modal data cannot reflect them, for example, internal cracks cannot be reflected in the surface image, so there will be a phenomenon of not being suitable for all situations.
[0004] Therefore, there is an urgent need for a method for detecting the state of a wind turbine blade based on multi-modal data to achieve more accurate, timely, and convenient detection. Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for detecting the state of a wind turbine blade based on multi-modal data fusion. By combining the multi-modal data fusion methods of feature-level fusion and decision-level fusion, while reducing the task amount and redundant data, it enhances interpretability and improves data utilization rate, and generally improves the accuracy, convenience, and real-time performance of blade detection.
[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:
[0007] The first aspect of the present invention provides a method for detecting the state of a wind turbine blade based on multi-modal data fusion;
[0008] A method for detecting the state of a wind turbine blade based on multi-modal data fusion includes:
[0009] Obtain multiple modal data of the blade to be detected, extract the features of each modal data therefrom, perform feature-level fusion, and generate a new multi-modal fusion feature;
[0010] Input the multi-modal fusion feature into the trained modal models to obtain the detection results of each modal model;
[0011] Perform decision-level fusion on the detection results of each modal model to obtain the final detection result of the wind turbine blade state.
[0012] Further, the modal data includes visible light images of the blade, infrared images of the blade, blade sound, and blade vibration signals.
[0013] Further, the feature extraction of the visible light image of the blade is specifically as follows:
[0014] (1) Obtain the surface image of the blade to be detected through a drone;
[0015] (2) Perform defogging processing;
[0016] (3) Use a CNN model with an attention mechanism to extract features from the defogged image to obtain the visible light map feature map A(i, j), where i and j represent the pixel positions.
[0017] Further, the attention mechanism includes channel attention and spatial attention;
[0018] The channel attention is to perform max pooling and average pooling on the feature map to obtain two vectors, put the two vectors of the same dimension into the same perceptron for learning, add the output results one by one and put them into the sigmod function for activation to obtain the channel attention vector, and multiply it with the original vector to obtain the feature vector under the attention mechanism;
[0019] The spatial attention is to perform max pooling and average pooling on the output result of the channel attention, splice the two results and perform a convolution operation to obtain the spatial attention vector, and multiply it with the output result of the channel attention mechanism to obtain the final feature map.
[0020] Further, the feature-level fusion is to fuse the features of the visible light image of the blade and the infrared image of the blade, specifically as follows:
[0021] (1) Obtain the infrared image B(i, j) of the blade through infrared imaging, and obtain the grayscale image C(i, j) of the infrared image of the blade through grayscale processing, where i and j represent the pixel positions;
[0022] (2) Set a zero matrix D(i, j) to be coupled with the same size as the grayscale image C, and set a threshold t. When C(i, j) is greater than t, take the pixel value of the corresponding point of the infrared image of the blade, that is, make D(i, j) equal to B(i, j); when C(i, j) is less than or equal to t, take the pixel value of the corresponding point of the visible light image, that is, make D(i, j) equal to A(i, j).
[0023] Further, the modal model includes a decision-making model based on visible light images and infrared images, a decision-making model based on acoustic features, and a decision-making model based on vibration signals.
[0024] Further, in the decision-level fusion, the detection results of each modal model are weighted and fused through a trained decision perceptron to obtain the final detection result of the wind turbine blade state.
[0025] The second aspect of the present invention provides a wind turbine blade state detection system based on multi-modal data fusion.
[0026] A wind turbine blade state detection system based on multi-modal data fusion includes a feature generation module, a modal detection module, and a decision fusion module:
[0027] The feature generation module is configured to: obtain multiple modal data of the blade to be detected, extract the features of each modal data therefrom, perform feature-level fusion, and generate new multi-modal fusion features;
[0028] The modal detection module is configured to: input the multi-modal fusion features into the trained modal models to obtain the detection results of each modal model;
[0029] The decision fusion module is configured to: perform decision-level fusion on the detection results of each modal model to obtain the final detection result of the wind turbine blade state.
[0030] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in a method for detecting the state of a wind turbine blade based on multi-modal data fusion as described in the first aspect of the present invention are implemented.
[0031] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor, and when the processor executes the program, the steps in a method for detecting the state of a wind turbine blade based on multi-modal data fusion as described in the first aspect of the present invention are implemented.
[0032] The above one or more technical solutions have the following beneficial effects:
[0033] The multi-modal data fusion-based blade detection method of the present invention makes full use of the multi-modal data of blades to accurately detect blade faults in all aspects, greatly improving the detection accuracy. At the same time, traditional blade detection requires a large amount of manpower for on-site detection, has high requirements for technicians, high labor costs, long detection cycles, and large maintenance workloads. For this method, no on-site manpower detection is required, the technical requirements are low, the labor costs required are low, and real-time detection can be achieved with a low detection workload. Therefore, the blade detection method of the present invention realizes more accurate, timely, and convenient detection, with a huge improvement compared to traditional detection methods.
[0034] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become apparent from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0036] Figure 1 It is a flowchart of the method for the first embodiment.
[0037] Figure 2 It is a structural diagram of the CNN model introducing the attention mechanism CBAM for the first embodiment.
[0038] Figure 3 It is a structural diagram of the convolutional neural network for the first embodiment.
[0039] Figure 4 It is a flowchart of image detection for the first embodiment.
[0040] Figure 5 It is a flowchart of sound detection for the first embodiment.
[0041] Figure 6 It is a flowchart of vibration detection for the first embodiment.
[0042] Figure 7 It is a structural diagram of the decision perceptron for the first embodiment.
[0043] Figure 8 It is a system structural diagram for the second embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0045] There are two fusion methods involved in blade detection:
[0046] An early fusion, namely feature-level fusion, after the data of different modalities are subjected to feature extraction, the feature vectors are fused to obtain a fused feature vector that fuses multi-modal data, and then a fault detection model is trained based on the fused feature vector to realize the discrimination of blade faults;
[0047] Another is late fusion, namely decision-level fusion. For data of different modalities, features are extracted, and respective fault detection models are trained in this modality. Then, the output results of the respective fault detection models are calculated to obtain a comprehensive result in a weighted manner to give the final detection result.
[0048] Compared with the two methods, feature-level fusion has the characteristics of small task volume and less data redundancy; decision-level fusion has the characteristics of strong interpretability and high data utilization rate. Based on the two fusion methods, the present invention proposes a multi-modal data fusion method that combines feature-level fusion and decision-level fusion, and realizes intelligent early warning of wind turbine blades through data fusion and decision fusion from multiple angles; the physical state involves multiple parameters, including visible light images of blades, infrared images of blades, blade sounds, blade vibration signals, etc.; through drone shooting and infrared imaging technology, visible light images and infrared images of blades are obtained; through a sound collector, the blade sounds are comprehensively collected from multiple directions; through a vibration sensor, vibration-related data are obtained; then, combined with machine learning methods such as dehazing, attention mechanism, and convolutional neural network, the data are processed and the faults of the blades are predicted to realize real-time intelligent monitoring of wind turbine blades, which can not only reduce the task volume and redundant data, but also enhance the interpretability and improve the data utilization rate.
[0049] Embodiment 1
[0050] This embodiment discloses a method for detecting the state of a wind turbine blade based on multi-modal data fusion;
[0051] As Figure 1 shown, a method for detecting the state of a wind turbine blade based on multi-modal data fusion includes:
[0052] Step S1: Obtain multi-modal data of the blade to be detected, extract the features of each modal data therefrom, perform feature-level fusion, and generate a new multi-modal fusion feature;
[0053] The modal data includes visible light images of blades, infrared images of blades, blade sounds, and blade vibration signals.
[0054] The processing steps of the visible light image of the blade are:
[0055] (1) Obtain the surface picture of the blade to be detected through a drone;
[0056] (2) Perform defogging processing to obtain a defogged image: Since there is sufficient wind energy resources in the western and coastal regions, wind turbines are often distributed in such places, and foggy days are often accompanied. The obtained images have a great impact on subsequent processing. Therefore, a method of surface image defogging processing is adopted. According to the principle of atmospheric scattering, the fogging imaging model is as follows: The first part comes from the attenuated incident light source; the second part comes from the scattering of other light sources. The fogging imaging model, that is, the degradation mathematical expression of the foggy image is:
[0057] I(x) = J(x)t(x) + A(1 - t(x)) (1)
[0058] Among them, the first part is J(x)t(x), I(x) represents the foggy image collected by the drone, J(x) represents the target defogged image, the second part is A(1 - t(x)), t(x) represents the transmittance of the scene, and A represents the atmospheric light value.
[0059] From formula (1), the mathematical expression of the target defogged image J(x) can be obtained as:
[0060] J(x) = (I(x) - A(1 - t(x))) / t(x) (2)
[0061] The dark channel gray value of the foggy image is mainly determined by the atmospheric light. The transmittance t and the atmospheric light value A can be estimated. The estimation method of A is: Extract the pixel positions where the gray value of the dark channel image is in the top 0.1%, correspond to the corresponding positions in the foggy image, and then find the maximum pixel value at these positions. This pixel value is the atmospheric light value A.
[0062] The estimation method of t is: Use the guided filter method to refine the transmittance. The mathematical expression of the guided filter method is:
[0063] q = a*I + b (3)
[0064] Among them, I is the guidance map, q is the target defogged image obtained by its transformation, a and b are the fixed parameters of this window. This model believes that there is a linear relationship between the guidance map and the target map within a certain range.
[0065] The actual obtained original image is p. By calculating the difference between the input p and the output q and then performing a square operation, and using the least squares method to calculate the values of a and b, we can obtain:
[0066]
[0067]
[0068]
[0069] Among them, p is the actually obtained image, q is the image after output, and a and b are parameters defined by the guided filtering method.
[0070] Through the above method, dehazing can be completed quickly and accurately.
[0071] (3) Extract features from the dehazed image.
[0072] This embodiment adopts a feature extraction method introducing an attention mechanism: in actual detection, the main damage modes of fan blades are often cracks and corrosion. Therefore, the pictures obtained by the drone are divided into cracks, corrosion, and no defects. In the environment of the blades, there are often problems such as complex backgrounds, uneven illumination, and small defect ratios. Conventional convolutional neural networks often have difficulty extracting key features.
[0073] To address these problems, the attention mechanism CBAM is introduced, and a CNN model with the attention mechanism CBAM is used to extract image features. The specific structure of the model is as Figure 2 shown.
[0074] CNN consists of an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The input is multiplied by the convolutional kernel to obtain the projection of the input in this feature. After multiple calculations with the convolutional layer and the pooling layer, the feature vector of the output image is finally obtained. When processing the images of fan blades, to solve the problems of uneven illumination, complex backgrounds, and small defect ratios and achieve accurate feature extraction, an attention mechanism is introduced.
[0075] The attention mechanism CBAM is a way of introducing attention with relatively low computational cost, which can improve the quality of features almost without changing the computational cost. The attention mechanism includes channel attention and spatial attention.
[0076] To achieve channel attention: first, perform max pooling and average pooling on the feature map to obtain two vectors. Put the two vectors with the same dimension into the same perceptron for learning. Add the output results one by one and put them into the sigmod function for activation to obtain the channel attention vector, and multiply it with the original vector to obtain the feature vector under the attention mechanism. The feature map is the matrix storage after image dehazing. Channel attention is part of the attention mechanism. The attention mechanism is a special operation layer, including convolution, pooling, full connection, and some operations integrating the attention mechanism, which is equivalent to adding a special layer in the cnn.
[0077] To achieve spatial attention: perform max pooling and average pooling on the output result of the channel attention mechanism, splice the two results and perform convolution operation to obtain the spatial attention vector, and multiply it with the output result of the channel attention mechanism to obtain the feature map A(i,j) with enhanced spatial attention, where i and j represent the positions of pixel points.
[0078] Since drones can only obtain visible light images of the blade surface and cannot judge internal faults, while infrared images can judge internal faults, and there is a high degree of information redundancy between the surface feature map and the infrared feature map of the blade, the feature maps of the blade visible light image and the infrared image are fused at the feature level to obtain the internal and external fault features of the blade and reduce information redundancy. Specifically:
[0079] (1) By means of infrared imaging, obtain the blade infrared image B(i,j), and obtain the grayscale image C(i,j) of the infrared image through grayscale processing, where i and j represent the positions of pixel points.
[0080] (2) Set a coupling zero matrix D(i,j) of the same size as the grayscale image C, and set a threshold t. When C(i,j) is greater than t, take the pixel value of the corresponding point in the infrared image, that is, make D(i,j) equal to B(i,j); when C(i,j) is less than or equal to t, take the pixel value of the corresponding point in the visible light image, that is, make D(i,j) equal to A(i,j). It is expressed by the formula:
[0081]
[0082] In this way, an image feature D(i,j) that fuses the visible light image and the infrared image is obtained. Compared with the single visible light image, this image can obtain the internal anomalies of the blade. Compared with the single infrared image, this image is clearer, can better handle small faults and can obtain the fault location more accurately.
[0083] Regarding the processing of the blade sound, specifically:
[0084] (1) Obtain the relevant data of the fan blade: The sound field of the fan is obtained by using an acoustic array sensor to collect sound information. Through the arrangement of the sound sensors, the omnidirectional collection of the fan sound is realized, the influence of noise on the local sound collection is reduced, and the integrity of information collection is ensured at the same time.
[0085] (2) For the obtained sound signal, preprocess it to remove noise. Through experiments, it is measured that the sound signal collected in an environment without fan sound is a low-frequency signal compared with the fan sound signal collected. Therefore, the high-pass filtering method is used to remove noise.
[0086] (3) For the extraction of the sound signal, an MFCC feature extraction method is used to extract the acoustic features. Specifically: First, add the vi ocebox package in mt l ab, and then based on the relevant tools in the vi ocebox package, subsequent operations such as reading, pre-emphasis, framing, windowing, and Fourier transform are realized to extract the features of the sound.
[0087] Regarding the processing of the blade vibration signal, specifically:
[0088] (1) Obtain relevant data on the vibration of the blade through a vibration sensor;
[0089] (2) Since the vibration of the fan blade has a natural frequency and the frequency is relatively stable, Fourier transform is used for filtering. The frequency-domain distribution of the vibration signal is obtained through Fourier transform, and the noise frequency domain is removed to obtain the denoised vibration signal, which is used as the vibration signal feature of the blade.
[0090] Step S2: Input the multi-modal fusion features into the trained models of each modality to obtain the detection results of each modality model;
[0091] The modality models include a fault detection model based on images, a fault detection model based on acoustics, and a fault detection model based on vibration signals.
[0092] The fault detection model based on images is constructed based on the convolutional neural network as shown in Figure 3 and is trained by the image features extracted from the training set samples; the image features of the blade to be detected are input into the trained model to obtain the image detection result, and the process is as shown in Figure 4 shown.
[0093] The fault detection model based on acoustics is constructed based on the convolutional neural network as shown in Figure 3 and is trained by the acoustic features extracted from the training set samples to train the neural network model; the acoustic features of the blade to be detected are input into the trained model to obtain the acoustic detection result, and the process is as shown in Figure 5 shown.
[0094] The fault detection model based on vibration signals is constructed based on the convolutional neural network as shown in Figure 3 and, according to the denoised vibration signals of the training set samples, using the vibration signals of the blades in the normal state as positive samples and the vibration signals of the blades in the fault state as negative samples, trains the neural network model to obtain a fault detection model based on vibration signals; the vibration signals of the blade to be detected are filtered by Fourier transform, and the denoised signal obtained after filtering is input into the fault detection model based on vibration signals to obtain the output as the vibration signal detection result, and the process is as shown in Figure 6 shown.
[0095] Step S3: Perform decision-level fusion on the detection results of each modality model to obtain the final detection result of the fan blade state, specifically:
[0096] Construct a decision perceptron, the structure of which is as shown in Figure 7 and, based on the image detection results, acoustic detection results, and vibration signal detection results of the samples in the training set, train the decision perceptron to obtain the trained decision perceptron.
[0097] Based on the trained decision perceptron, the image detection result, acoustic detection result, and vibration signal detection result of the blade to be detected are weighted and fused to obtain the final detection result of the fan blade state.
[0098] Embodiment 2
[0099] This embodiment discloses a fan blade state detection system based on multimodal data fusion;
[0100] As Figure 8 shown, a fan blade state detection system based on multimodal data fusion includes a feature generation module, a modal detection module, and a decision fusion module:
[0101] The feature generation module is configured to: obtain multiple modal data of the blade to be detected, extract the features of each modal data therefrom, perform feature-level fusion, and generate new multimodal fusion features;
[0102] The modal detection module is configured to: input the multimodal fusion features into the trained modal models of each modality to obtain the detection results of each modal model;
[0103] The decision fusion module is configured to: perform decision-level fusion on the detection results of each modal model to obtain the final detection result of the fan blade state.
[0104] Embodiment 3
[0105] The purpose of this embodiment is to provide a computer-readable storage medium.
[0106] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the steps in a method for detecting the state of a fan blade based on multimodal data fusion as described in Embodiment 1 of the present disclosure.
[0107] Embodiment 4
[0108] The purpose of this embodiment is to provide an electronic device.
[0109] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for detecting the state of a fan blade based on multimodal data fusion as described in Embodiment 1 of the present disclosure.
[0110] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting the state of a wind turbine blade based on multi-modal data fusion, characterized in that, Including: Obtain multiple modal data of the blade to be detected, extract the features of each modal data, and perform feature-level fusion on the feature map of the visible light image of the blade and the infrared image of the blade to obtain the blade image feature; The modal data includes the visible light image of the blade, the infrared image of the blade, the blade sound, and the blade vibration signal; The feature extraction of the visible light image of the blade is specifically: (1) Obtain the surface image of the blade to be detected by an unmanned aerial vehicle; (2) Perform dehazing processing; (3) Use a CNN model with an attention mechanism to extract features from the dehazed image to obtain the visible light map feature map A(i,j), where i and j represent the pixel positions; The feature-level fusion is to fuse the features of the visible light image of the blade and the infrared image of the blade to obtain the image feature D(i,j) of the fused visible light image and infrared image, specifically: (1) Obtain the infrared image B(i,j) of the blade by infrared imaging, and obtain the grayscale image C(i,j) of the infrared image of the blade by grayscale processing, where i and j represent the pixel positions; (2) Set a zero matrix D(i,j) to be coupled with the same size as the grayscale image C, set a threshold t. When C(i,j) is greater than t, take the pixel value of the corresponding point of the infrared image of the blade, that is, let D(i,j) be equal to B(i,j); when C(i,j) is less than or equal to t, take the pixel value of the corresponding point of the visible light map, that is, let D(i,j) be equal to A(i,j). The formula is expressed as: ; Input the blade image feature, the blade acoustic feature, and the blade vibration signal feature into the trained modal models respectively to obtain the detection results of the modal models; Perform decision-level fusion on the detection results of the modal models to obtain the final detection result of the fan blade state.
2. The method for detecting the state of a wind turbine blade based on multi-modal data fusion according to claim 1, wherein, The attention mechanism includes channel attention and spatial attention; The channel attention is to perform max pooling and average pooling on the feature map to obtain two vectors, put the two vectors with the same dimension into the same perceptron for learning, add the output results one by one and put them into the Sigmoid function for activation to obtain the channel attention vector, and multiply it with the original vector to obtain the feature vector under the attention mechanism; The spatial attention is to perform max pooling and average pooling processing on the result output by the channel attention, splice the two results and perform a convolution operation to obtain the spatial attention vector, and multiply it with the result output by the channel attention mechanism to obtain the final feature map.
3. The method for detecting the state of a wind turbine blade based on multi-modal data fusion according to claim 1, wherein The modal models include: a fault detection model based on images, a fault detection model based on acoustics, and a fault detection model based on vibration signals.
4. The method for detecting the state of a fan blade based on multi-modal data fusion according to claim 1, wherein, The decision-level fusion, through a trained decision perceptron, weights and fuses the detection results of the modal models to obtain the final detection result of the fan blade state.
5. A fan blade condition detection system based on multi-modal data fusion, characterized in that, Using the method described in any one of claims 1-4 for fan blade state detection, Including a feature generation module, a modal detection module, and a decision fusion module: A feature generation module, configured to: obtain multiple modal data of a blade to be detected, extract features of each modal data therefrom, and perform feature-level fusion on the feature map of the visible light image of the blade and the infrared image of the blade to obtain blade image features; A modal detection module, configured to: respectively input the blade image features, the blade acoustic features, and the blade vibration signal features into the trained modal models to obtain the detection results of the modal models; A decision fusion module, configured to: perform decision-level fusion on the detection results of the modal models to obtain the final detection result of the fan blade state.
6. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a method for detecting the state of a fan blade based on multi-modal data fusion as described in any one of claims 1-4.
7. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a method for detecting the state of a fan blade based on multi-modal data fusion as described in any one of claims 1-4.
Citation Information
Patent Citations
Intelligent fan blade defect detection method based on bimodal fusion
CN114429457A
Fan blade crack detection device and method
CN115166032A
Multi-modal data fusion decision-making method based on image type intermediate state
CN115393678A