Fan blade monitoring method and device based on multi-modal analysis, equipment and medium
By combining image, vibration and sound data through multimodal analysis, multimodal vectors of wind turbine blades are generated, which solves the problem of low reliability of traditional single-mode monitoring and achieves more reliable fault prediction.
Patent Information
- Application Number
- CN202510550896.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Traditional wind turbine blade monitoring methods rely on a single mode, which cannot comprehensively and accurately reflect the true state of the blades, resulting in low monitoring reliability.
By employing a multimodal analysis method, multimodal vectors of wind turbine blades are generated through the acquisition and fusion of image, vibration, and sound modal data, thereby predicting blade fault data.
The reliability of wind turbine blade monitoring has been improved. By mining and fusing multimodal data, semantic representation capabilities have been enhanced, ensuring the accuracy of fault prediction.
Smart Images

Figure CN120312508B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of wind turbine blade monitoring, and in particular to a wind turbine blade monitoring method, apparatus, device, and medium based on multimodal analysis. Background Art
[0002] With the rapid development of the wind power industry, the operating efficiency and stability of wind turbines have become crucial factors in the energy industry. As one of the core components of a wind turbine, the performance of wind turbine blades directly affects the efficiency of wind power generation. Therefore, accurate and real-time monitoring of the health of wind turbine blades is of great significance for ensuring the stable operation of wind turbines and reducing downtime. Traditional wind turbine blade monitoring methods usually rely on a single monitoring method, such as monitoring the vibration state of the blades through vibration sensors to determine whether there are any abnormalities. Although these methods can effectively provide some fault diagnosis information, due to the complex and harsh working environment of wind turbine blades, single-mode monitoring often cannot fully and accurately reflect the true state of the blades. Summary of the Invention
[0003] In view of this, the purpose of the present application is to provide a method and device, equipment and medium for monitoring wind blades based on multimodal analysis, so as to improve the problem of relatively low reliability of wind blade monitoring in the prior art.
[0004] To achieve the above objectives, this application adopts the following technical solutions:
[0005] A wind turbine blade monitoring method based on multimodal analysis, comprising:
[0006] Based on fan blade multimodal data corresponding to a target fan blade, mining a corresponding blade image vector, a blade vibration vector, and a blade sound vector according to predetermined image modes, vibration modes, and sound modes, respectively, wherein the fan blade multimodal data is formed by collecting image, vibration, and sound information of the target fan blade;
[0007] Performing multimodal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector to output a fan blade multimodal vector;
[0008] Based on the wind blade multi-modal vector, wind blade fault data is predicted and output, wherein the wind blade fault data is used to characterize whether the target wind blade has a fault.
[0009] In a preferred embodiment of the present application, in the above-mentioned wind turbine blade monitoring method based on multimodal analysis, the step of performing multimodal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector to output the wind turbine blade multimodal vector includes:
[0010] According to the generated interference data, mining the corresponding interference image vector, interference vibration vector and interference sound vector according to the image mode, the vibration mode and the sound mode respectively;
[0011] Based on the blade image vector, the blade vibration vector, and the blade sound vector, semantically diffuse the interference image vector, the interference vibration vector, and the interference sound vector, and output corresponding diffuse image vectors, diffuse vibration vectors, and diffuse sound vectors;
[0012] The diffuse sound vector is semantically diffused according to the diffuse image vector and the diffuse vibration vector, and a fan blade multimodal vector is output.
[0013] In a preferred embodiment of the present application, in the above-mentioned wind turbine blade monitoring method based on multimodal analysis, the step of performing semantic diffusion on the interference image vector, the interference vibration vector, and the interference sound vector based on the blade image vector, the blade vibration vector, and the blade sound vector, and outputting the corresponding diffusion image vector, diffusion vibration vector, and diffusion sound vector includes:
[0014] Performing X stages of vector compression operations on the interference image vector, and determining an image compression vector formed by the vector compression operation of the Xth stage as a diffusion image vector, wherein for each stage of the vector compression operation, the leaf image vector is fused with the image compression vector formed by the vector compression operation of the previous stage to form an image compression vector of the vector compression operation of the current stage;
[0015] Performing X stages of vector compression operations on the interference vibration vector, and determining a vibration compression vector formed by the vector compression operation of the Xth stage as a diffusion vibration vector, wherein for each stage of the vector compression operation, the blade vibration vector is fused based on the vibration compression vector and the image compression vector formed by the vector compression operation of the previous stage to form a vibration compression vector of the vector compression operation of the current stage;
[0016] X stages of vector compression operations are performed based on the interference sound vector, and the sound compression vector formed by the vector compression operation of the Xth stage is determined as a diffuse sound vector, wherein for each stage of the vector compression operation, the blade sound vector is fused based on the sound compression vector formed by the vector compression operation of the previous stage and the vibration compression vector formed by the vector compression operation of the current stage to form the sound compression vector of the vector compression operation of the current stage.
[0017] In a preferred embodiment of the present application, in the above-mentioned wind turbine blade monitoring method based on multimodal analysis, the step of fusing the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage, and forming the vibration compression vector of the vector compression operation in the current stage, includes:
[0018] Adding the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage to form a vector to be processed in the vector compression operation in the current stage;
[0019] A vector compression operation is performed based on the to-be-processed vector of the vector compression operation in the current stage and the blade vibration vector to form a vibration compression vector of the vector compression operation in the current stage.
[0020] In a preferred embodiment of the present application, in the above-mentioned wind turbine blade monitoring method based on multimodal analysis, the step of performing semantic diffusion on the diffuse sound vector based on the diffuse image vector and the diffuse vibration vector and outputting the wind turbine blade multimodal vector includes:
[0021] Performing X stages of vector dilation operations on the diffusion image vector, and determining an image dilation vector formed by the vector dilation operation of the Xth stage as an image dilation target vector, wherein for each stage of the vector dilation operation, the leaf image vector is fused with the image dilation vector formed by the vector dilation operation of the previous stage to form an image dilation vector for the vector dilation operation of the current stage;
[0022] Performing X stages of vector expansion operations on the diffuse vibration vector, and determining a vibration expansion vector formed by the X-th stage of the vector expansion operation as a vibration expansion target vector, wherein for each stage of the vector expansion operation, the blade vibration vector is fused with the vibration expansion vector formed by the vector expansion operation of the previous stage to form a vibration expansion vector of the vector expansion operation of the current stage;
[0023] Performing X-stage vector expansion operations on the diffuse sound vector, and determining the sound expansion vector formed by the X-th stage vector expansion operation as a fan blade multimodal vector;
[0024] The sound expansion vector includes the sound mode vector, the sound mode internal focus vector, and the sound mode correlation focus vector formed during the vector expansion operation. For each stage of the vector expansion operation:
[0025] Deeply mine the sound modal correlation focus vector formed by the vector expansion operation in the previous stage to form the sound modal vector of the vector expansion operation in the current stage;
[0026] Performing modal internal focusing on the sound modal vector formed by the vector expansion operation at the current stage and the image modal vector in the image expansion vector formed by the vector expansion operation at the current stage, thereby forming a sound modal internal focusing vector of the vector expansion operation at the current stage;
[0027] Based on the sound modal internal focusing vector formed by the vector expansion operation of the current stage and the vibration modal internal focusing vector in the vibration expansion vector formed by the vector expansion operation of the current stage, modal correlation focusing is performed with the blade sound vector to form the sound modal correlation focusing vector of the vector expansion operation of the current stage.
[0028] In a preferred embodiment of the present application, in the above-mentioned wind turbine blade monitoring method based on multimodal analysis, the step of performing modal internal focusing on the sound modal vector formed by the vector expansion operation at the current stage and the image modal vector in the image expansion vector formed by the vector expansion operation at the current stage to form the sound modal internal focusing vector of the vector expansion operation at the current stage includes:
[0029] Determining, based on the sound mode vector formed by the vector expansion operation in the current stage, a first mapping vector, a second mapping vector, and a third mapping vector corresponding to the sound mode and applied to the modal internal focusing;
[0030] forming a new first mapping vector based on the first mapping vector corresponding to the sound modality and the first mapping vector corresponding to the image modality, wherein the first mapping vector corresponding to the image modality is determined based on the image modality vector formed by the vector expansion operation in the current stage;
[0031] The sound modal internal focusing vector of the vector expansion operation at the current stage is determined based on the second mapping vector, the third mapping vector and the new first mapping vector corresponding to the sound modal and applied to the internal focusing of the modal.
[0032] In a preferred embodiment of the present application, in the above-mentioned wind turbine blade monitoring method based on multimodal analysis, the sound mode internal focusing vector formed according to the vector expansion operation at the current stage, the vibration mode internal focusing vector in the vibration expansion vector formed by the vector expansion operation at the current stage, and the blade sound vector are modally correlated and focused to form the sound mode correlation focusing vector of the vector expansion operation at the current stage, including:
[0033] Determining, based on the sound modal internal focusing vector of the vector expansion operation in the current stage, a first mapping vector corresponding to the sound modal and applied to modal correlation focusing, and, based on the blade sound vector, determining a second mapping vector and a third mapping vector corresponding to the sound modal and applied to modal correlation focusing;
[0034] forming a new first mapping vector based on the first mapping vector corresponding to the sound mode and the first mapping vector corresponding to the vibration mode, wherein the first mapping vector corresponding to the vibration mode is determined based on the vibration mode vector formed by the vector expansion operation in the current stage;
[0035] forming a new third mapping vector according to the third mapping vector corresponding to the sound mode and applied to modal correlation focusing and the third mapping vector corresponding to the vibration mode, wherein the third mapping vector corresponding to the vibration mode is determined based on the blade vibration vector;
[0036] The sound modality-associated focusing vector of the vector expansion operation at the current stage is determined based on the second mapping vector corresponding to the sound modality and applied to the internal focusing of the modality, the new third mapping vector and the new first mapping vector.
[0037] The present application also provides a wind turbine blade monitoring device based on multimodal analysis, comprising:
[0038] a data mining module for mining, based on fan blade multimodal data corresponding to a target fan blade, corresponding blade image vectors, blade vibration vectors, and blade sound vectors according to predetermined image modes, vibration modes, and sound modes, respectively, wherein the fan blade multimodal data is generated by collecting image, vibration, and sound information of the target fan blade;
[0039] a semantic fusion module, configured to perform multimodal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and output a fan blade multimodal vector;
[0040] A fault prediction module is used to predict and output fan blade fault data based on the fan blade multi-modal vector, wherein the fan blade fault data is used to characterize whether the target fan blade has a fault.
[0041] Based on the above, the present application further provides an electronic device, including:
[0042] memory for storing computer programs;
[0043] The processor connected to the memory is used to execute the computer program stored in the memory to implement the above-mentioned wind turbine blade monitoring method based on multimodal analysis.
[0044] On the basis of the above, the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run, each step of the above-mentioned wind turbine blade monitoring method based on multimodal analysis is executed.
[0045] The wind turbine blade monitoring method, device, equipment and medium based on multimodal analysis provided in the present application mine the corresponding blade image vector, blade vibration vector and blade sound vector according to the predetermined image mode, vibration mode and sound mode based on the multimodal data of the wind turbine blade corresponding to the target wind turbine blade; secondly, the blade image vector, blade vibration vector and blade sound vector are subjected to multimodal semantic information fusion to output the wind turbine blade multimodal vector; then, based on the wind turbine blade multimodal vector, the wind turbine blade fault data is predicted and output. Based on the above content, since the multimodal data can be mined separately, the accuracy of semantic mining can be higher. In addition, by fusing the semantic vectors of multiple modes, a multimodal vector with stronger representation ability can be obtained, so that the semantic vector of a single mode can be strengthened. In this way, the reliability of prediction based on the multimodal vector can be guaranteed, thereby obtaining reliable wind turbine blade fault data. Therefore, the problem of relatively low reliability of wind turbine blade monitoring in the prior art can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings.
[0047] Figure 1 This is a structural block diagram of an electronic device provided in an embodiment of the present application.
[0048] Figure 2 A flow chart of a wind turbine blade monitoring method based on multimodal analysis provided in an embodiment of the present application.
[0049] Figure 3 A schematic diagram of semantic diffusion provided in an embodiment of the present application.
[0050] Figure 4 Another schematic diagram of semantic diffusion provided in an embodiment of the present application.
[0051] Figure 5 A block diagram of a wind turbine blade monitoring device based on multimodal analysis provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0053] like Figure 1 As shown, an embodiment of the present application provides an electronic device, wherein the electronic device may include a memory, a processor, and a wind turbine blade monitoring device based on multimodal analysis.
[0054] In detail, the memory and the processor are electrically connected directly or indirectly to realize data transmission or interaction. For example, the memory and the processor can be electrically connected through one or more communication buses or signal lines. The wind turbine blade monitoring device based on multimodal analysis includes at least one software function module stored in the memory in the form of software or firmware. The processor is used to execute the executable computer program stored in the memory, for example, the software function modules and computer programs included in the wind turbine blade monitoring device based on multimodal analysis, so as to realize the wind turbine blade monitoring method based on multimodal analysis provided in the embodiment of the present application.
[0055] Optionally, the memory may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. Furthermore, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on a chip (SoC), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0056] I understand. Figure 1 The structure shown is only for illustration, and the electronic device may also include Figure 1 More or fewer components than shown, or with Figure 1 The different configurations shown may, for example, further include a communication unit for exchanging information with other devices.
[0057] Combine Figure 2 , the embodiment of the present application also provides a wind turbine blade monitoring method based on multimodal analysis that can be applied to the above electronic device. Among them, the method steps defined in the process related to the wind turbine blade monitoring method based on multimodal analysis can be implemented by the electronic device. Figure 2 The specific process shown is explained in detail.
[0058] In step S110 , based on the multimodal data of the wind turbine blade corresponding to the target wind turbine blade, the corresponding blade image vector, blade vibration vector and blade sound vector are mined according to predetermined image mode, vibration mode and sound mode respectively.
[0059] In an embodiment of the present application, the electronic device can mine corresponding blade image vectors, blade vibration vectors, and blade sound vectors based on the multimodal data of a target wind turbine blade, according to predetermined image modes, vibration modes, and sound modes. The multimodal data of the wind turbine blade is generated by collecting image, vibration, and sound information from the target wind turbine blade, such as by an image acquisition device collecting data in the image mode, a vibration sensor collecting data in the vibration mode, and a sound collector collecting data in the sound mode. Thus, after collecting data in each modality, data can be mined separately to mine the corresponding blade image vectors, blade vibration vectors, and blade sound vectors, that is, semantic information in the data in each modality is mined and represented in the form of vectors. For example, semantic mining can be performed on the three modal data separately using a convolutional network in a trained neural network model. Since the data modalities are different, to ensure the accuracy of semantic mining, three different convolutional networks can be used. For the sound modality data, the data can be first converted into a spectrogram, and then the spectrogram is convolved to obtain the corresponding blade sound vector. Alternatively, the sound time series data may be directly convolved to form a corresponding leaf sound vector.
[0060] Step S120 , performing multimodal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and outputting a wind turbine blade multimodal vector.
[0061] In an embodiment of the present application, after mining the blade image vector, the blade vibration vector, and the blade sound vector, the electronic device can perform multimodal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector to output a multimodal vector of the fan blade. Based on this, since the image modality may not be able to capture the details of the blade during high-speed rotation or slight vibration, the vibration modality may not be able to effectively identify certain types of damage or cracks, and the sound modality may be affected by environmental noise, that is, the semantic vector of a single modality may not be able to fully characterize the fault of the target fan blade. Therefore, by performing multimodal semantic information fusion, the semantic vectors of each modality can be fully fused, thereby obtaining a multimodal vector of the fan blade with a more reliable and sufficient semantic representation.
[0062] Step S130 : predicting and outputting fan blade fault data based on the fan blade multi-modal vector.
[0063] In an embodiment of the present application, after obtaining the fan blade multimodal vector, the electronic device can predict and output fan blade fault data based on the fan blade multimodal vector. The fan blade fault data is used to characterize whether the target fan blade has a fault. Exemplarily, the above-mentioned neural network model can also include a fully connected network and an output network. In this way, after obtaining the fan blade multimodal vector, the fan blade multimodal vector can be processed by the fully connected network to output a 1*2 fully connected vector. Then, the fully connected vector can be mapped to a probability distribution (a, b) through the output function in the output network (such as softmax and other classification functions), such as a corresponds to the probability of a fault, and b corresponds to the probability of no fault. Then, the type corresponding to the larger probability is determined as the fan blade fault data.
[0064] Based on the above content, since multimodal data can be mined separately, the accuracy of semantic mining can be higher. In addition, by fusing the semantic vectors of multiple modalities, a multimodal vector with stronger representation ability can be obtained, so that the semantic vector of a single modality can be strengthened. In this way, the reliability of prediction based on multimodal vectors can be guaranteed, thereby obtaining reliable wind blade fault data. Therefore, the problem of relatively low reliability of wind blade monitoring in the existing technology can be improved.
[0065] Regarding the above-mentioned step S120, it should be noted that the specific method of performing multimodal semantic information fusion is not limited and can be selected according to actual needs.
[0066] For example, in an alternative embodiment, in order to reduce the amount of data processing, the blade image vector, the blade vibration vector and the blade sound vector can be spliced together, and then the obtained spliced vector can be subjected to self-attention processing, convolution processing, etc., so as to obtain a multimodal vector of the wind turbine blade.
[0067] For another example, in another alternative embodiment, in order to improve the reliability of multimodal semantic information fusion, the above-mentioned step S120 may further include step S121, step S122 and step S123, and the specific content of each step is described as follows.
[0068] Step S121 , based on the generated interference data, mining the corresponding interference image vector, interference vibration vector and interference sound vector according to the image mode, the vibration mode and the sound mode respectively.
[0069] In an embodiment of the present application, the generated interference data can be used to mine corresponding interference image vectors, interference vibration vectors, and interference sound vectors according to the image mode, the vibration mode, and the sound mode. For example, the interference data can be randomly generated data or multimodal data of the target wind turbine blade collected historically. Thus, referring to the method of step S110, the interference data can be subjected to multimodal data mining to thereby obtain corresponding interference image vectors, interference vibration vectors, and interference sound vectors.
[0070] Step S122, based on the blade image vector, the blade vibration vector and the blade sound vector, the interference image vector, the interference vibration vector and the interference sound vector are semantically diffused, and the corresponding diffuse image vector, diffuse vibration vector and diffuse sound vector are output.
[0071] In an embodiment of the present application, after obtaining the interference image vector, the interference vibration vector, and the interference sound vector, the interference image vector, the interference vibration vector, and the interference sound vector can be semantically diffused based on the blade image vector, the blade vibration vector, and the blade sound vector, and the corresponding diffuse image vector, diffuse vibration vector, and diffuse sound vector can be output. In other words, since the interference image vector, the interference vibration vector, and the interference sound vector are historical data and have a certain reference significance, but the validity of the semantic information represented is not very high, they can be used as interference data. While providing a certain reference significance, they can also avoid the overfitting problem that is easy to occur when directly fusing the blade image vector, the blade vibration vector, and the blade sound vector. Based on this, by performing semantic diffusion, the blade image vector, the blade vibration vector, and the blade sound vector can be fully fused to obtain a diffuse image vector, a diffuse vibration vector, and a diffuse sound vector with a higher semantic representation accuracy.
[0072] Step S123 , performing semantic diffusion on the diffuse sound vector based on the diffuse image vector and the diffuse vibration vector, and outputting a fan blade multimodal vector.
[0073] In an embodiment of the present application, after obtaining a diffuse image vector, a diffuse vibration vector, and a diffuse sound vector, semantic diffusion can be performed on the diffuse sound vector based on the diffuse image vector and the diffuse vibration vector to output a multimodal vector of the wind turbine blade. In other words, the diffuse image vector and the diffuse vibration vector can be further fused into the diffuse sound vector, so that the semantic information in the diffuse sound vector can be enhanced based on the semantic information in the diffuse image vector and the diffuse vibration vector, thereby obtaining a multimodal vector of the wind turbine blade with richer semantic information.
[0074] It can be understood that, in the above-mentioned step S122, the specific manner of semantically diffusing the interference image vector, the interference vibration vector, and the interference sound vector is not limited. For example, in an alternative embodiment, in order to fully integrate the blade image vector, the blade vibration vector, and the blade sound vector representing the real semantic information into the interference image vector, the interference vibration vector, and the interference sound vector during the process of semantic diffusion, so as to suppress and remove the interference information, the above-mentioned step S122 may further include the following steps S122a, S122b, and S122c. The specific contents of each step are as follows (combined with Figure 3 ).
[0075] Step S122a: performing X stages of vector compression operations on the interference image vector, and determining an image compression vector formed by the X-th stage of vector compression operations as a diffusion image vector.
[0076] In an embodiment of the present application, the interference image vector can be subjected to X stages of vector compression operations, and the image compression vector formed by the X-th stage of vector compression operations can be determined as the diffusion image vector. For each stage of vector compression operations, the leaf image vector is fused with the image compression vector formed by the previous stage of vector compression operations to form the image compression vector for the current stage of vector compression operations. In other words, the image compression vector formed by the first stage of vector compression operations can be used to form the image compression vector formed by the second stage of vector compression operations. The image compression vector formed by the second stage of vector compression operations can then be used to form the image compression vector formed by the third stage of vector compression operations. This process continues in this order, resulting in the image compression vector formed by the final stage of vector compression operations. The image compression vector formed by the final stage of vector compression operations can then be determined as the diffusion image vector. Since the leaf image vectors are fused in each stage of vector compression operations, the leaf image vectors can be gradually integrated, thereby gradually removing or suppressing interference information, thereby obtaining a reliable diffusion image vector. Furthermore, performing vector compression operations can reduce the size of the vectors, thereby reducing the amount of data required for subsequent processing. Exemplarily, the above-mentioned neural network model may further include a first compression network, which may include multiple convolutional layers, batch normalization (Batch Normalization) and an activation function (ReLU), so that the corresponding vector can be compressed to achieve a reduction in vector size. For example, during the vector compression operation of the first stage, the interference image vector is convolved, normalized and activated (i.e., deep mining) to obtain the image modality vector of the vector compression operation of the first stage, and the image modality vector of the vector compression operation of the first stage is modally focused (refer to the explanation of step b2 below) to obtain the image modality internal focus vector of the vector compression operation of the first stage, and, based on the leaf image vector, the image modality internal focus vector of the vector compression operation of the first stage is modally correlated focused (refer to the explanation of step c3 below) to obtain the image modality correlation focus vector of the vector compression operation of the first stage (i.e., the image compression vector formed by the vector compression operation of the first stage).During the vector compression operation of the second stage, the image modal association focusing vector of the vector compression operation of the first stage will be (deep mined) to obtain the image modal vector of the vector compression operation of the second stage, and the image modal vector of the vector compression operation of the second stage will be modally internally focused to obtain the image modal internal focusing vector of the vector compression operation of the second stage. Moreover, based on the leaf image vector, the image modal internal focusing vector of the vector compression operation of the second stage will be modally associated focused to obtain the image modal association focusing vector of the vector compression operation of the second stage (i.e., the image compression vector formed by the vector compression operation of the second stage).
[0077] Step S122b: performing X-stage vector compression operations on the interference vibration vector, and determining a vibration compression vector formed by the X-th stage vector compression operation as a diffusion vibration vector.
[0078] In an embodiment of the present application, the interference vibration vector can be subjected to X stages of vector compression operations, and the vibration compression vector formed by the vector compression operation in the Xth stage is determined as a diffusion vibration vector. For each stage of the vector compression operation, the blade vibration vector is fused based on the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage to form the vibration compression vector for the vector compression operation in the current stage. That is, the vibration compression vector formed by the vector compression operation in the first stage can be used to form the vibration compression vector formed by the vector compression operation in the second stage. Then, the vibration compression vector formed by the vector compression operation in the second stage can be used to form the vibration compression vector formed by the vector compression operation in the third stage. And so on, the vibration compression vector formed by the vector compression operation in the last stage can be obtained. The vibration compression vector formed by the vector compression operation in the last stage can then be determined as the diffusion vibration vector. Since the blade vibration vector is fused in each stage of the vector compression operation, the blade vibration vector can be gradually integrated, i.e., the interference information can be gradually removed or suppressed, thereby obtaining a reliable diffusion vibration vector. In addition, the image compression vector of the corresponding stage can be fused in, so that the semantic information of the image modality and the vibration modality can be fused to further improve the suppression of interference information, thereby ensuring the accuracy of the semantic representation of the obtained diffuse vibration vector.
[0079] Step S122c: performing X stages of vector compression operations based on the interfering sound vector, and determining the sound compression vector formed by the X-th stage of vector compression operations as the diffuse sound vector.
[0080] In an embodiment of the present application, X stages of vector compression operations can be performed based on the interfering sound vector, and the sound compression vector formed by the X-th stage of vector compression operations can be determined as a diffuse sound vector. For each stage of vector compression operations, the blade sound vector is fused based on the sound compression vector formed by the previous stage of vector compression operations and the vibration compression vector formed by the current stage of vector compression operations to form the sound compression vector for the current stage of vector compression operations. In other words, the sound compression vector formed by the first stage of vector compression operations can be used to form the sound compression vector formed by the second stage of vector compression operations. Then, the sound compression vector formed by the second stage of vector compression operations can be used to form the sound compression vector formed by the third stage of vector compression operations. This process continues in this order, resulting in a sound compression vector formed by the final stage of vector compression operations. The sound compression vector formed by the final stage of vector compression operations can then be determined as the diffuse sound vector. Because the blade sound vectors are fused in each stage of vector compression operations, the blade sound vectors can be gradually integrated, i.e., interference information can be gradually removed or suppressed, thereby obtaining a reliable diffuse sound vector. In addition, the vibration compression vector of the corresponding stage can be integrated (the vibration compression vector also integrates the semantic information in the image compression vector). In this way, the semantic information of the image modality, vibration modality and sound modality can be integrated to further improve the suppression of interference information, thereby ensuring the accuracy of the semantic representation of the obtained diffuse sound vector.
[0081] It can be understood that, in the above-mentioned step S122b, the specific method of forming the vibration compression vector of the vector compression operation in the current stage is not limited. For example, in an alternative embodiment, in order to achieve full fusion of semantic information, the above-mentioned step S122b may further include step b1 and step b2, and the specific content of each step is described below.
[0082] Step b1: Add the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage to form a vector to be processed in the vector compression operation in the current stage.
[0083] In an embodiment of the present application, the vibration compression vector and the image compression vector formed by the vector compression operation of the previous stage can be added together to form a vector to be processed for the vector compression operation of the current stage. For example, in other embodiments, the result of the addition calculation can also be normalized to obtain the corresponding vector to be processed. Specifically, for the vector compression operation of the first stage, the interference vibration vector and the interference image vector can be added together to form the vector to be processed for the vector compression operation of the first stage. For the vector compression operation of the second stage, the vibration compression vector formed by the vector compression operation of the first stage and the image compression vector formed by the vector compression operation of the first stage can be added together.
[0084] Step b2: performing a vector compression operation based on the to-be-processed vector of the vector compression operation in the current stage and the blade vibration vector to form a vibration compression vector of the vector compression operation in the current stage.
[0085] In this embodiment of the present application, after obtaining the pending vector for the current phase of the vector compression operation, a vector compression operation can be performed based on the pending vector for the current phase of the vector compression operation and the blade vibration vector to form a vibration compression vector for the current phase of the vector compression operation. In this way, a blade vibration vector that represents true semantic information within interference information can be obtained, thereby continuously integrating true semantic information at each stage and obtaining a reliable vibration compression vector.
[0086] It is understandable that, in the above step b2, the specific manner of performing the vector compression operation based on the to-be-processed vector of the vector compression operation at the current stage and the blade vibration vector is not limited. For example, in an alternative embodiment, in order to achieve reliable fusion of real semantic information during the compression process, the above step b2 may further include the following:
[0087] First, the vector to be processed of the vector compression operation at the current stage can be deeply mined (as described above) to form the vibration modal vector in the vibration compression vector of the vector compression operation at the current stage, and, based on the vibration modal vector of the vector compression operation at the current stage, the modal internal focusing is performed to form the vibration modal internal focusing vector in the vibration compression vector of the vector compression operation at the current stage; illustratively, the transposed vector of the vibration modal vector of the vector compression operation at the current stage can be dot-producted with itself to obtain the corresponding dot product parameter distribution, and, based on the dot product parameter distribution, the vibration modal vector of the vector compression operation at the current stage is weightedly summed to obtain the vibration modal internal focusing vector of the vector compression operation at the current stage. In this way, the correlation between the two vectors can be characterized by the dot product parameter distribution, thereby achieving modal internal focusing;
[0088] Secondly, based on the vibration mode internal focusing vector of the vector compression operation at the current stage, modal correlation focusing is performed with the blade vibration vector to form a vibration mode correlation focusing vector in the vibration compression vector of the vector compression operation at the current stage, that is, the correlation information between the vibration mode internal focusing vector and the blade vibration vector is mined, so that irrelevant interference information is discarded, thereby obtaining a reliable vibration mode correlation focusing vector. Please refer to the explanation of step c3 in the following text.
[0089] It can be understood that in the above-mentioned step S122c, the specific method of forming the sound compression vector of the vector compression operation in the current stage is not limited. For example, in an alternative embodiment, in order to achieve full fusion of semantic information, and the sound compression vector includes the sound modal vector, the sound modal internal focus vector, and the sound modal association focus vector during the vector compression operation, based on this, the above-mentioned step S122c can further include step c1, step c2 and step c3, and the specific content of each step is described as follows.
[0090] In step c1, the sound modal association focus vector formed by the vector compression operation in the previous stage is deeply mined to form the sound modal vector of the vector compression operation in the current stage.
[0091] In an embodiment of the present application, the sound modal association focus vector formed by the vector compression operation in the previous stage is deeply mined to form the sound modal vector of the vector compression operation in the current stage. For example, the sound modal association focus vector formed by the vector compression operation in the first stage can be deeply mined to form the sound modal vector of the vector compression operation in the second stage, and the sound modal association focus vector formed by the vector compression operation in the second stage can be deeply mined to form the sound modal vector of the vector compression operation in the third stage. Exemplarily, for the first stage, the interfering sound vector can be deeply mined (as described above).
[0092] Step c2, based on the vibration mode internal focusing vector formed by the vector compression operation of the current stage, and the sound mode vector of the vector compression operation of the current stage, perform modal internal focusing to form the sound mode internal focusing vector of the vector compression operation of the current stage.
[0093] In an embodiment of the present application, after obtaining the vibration modal internal focusing vector formed by the vector compression operation of the current stage, the vibration modal internal focusing vector formed by the vector compression operation of the current stage can be modally internally focused with the sound modal vector of the vector compression operation of the current stage to form the sound modal internal focusing vector of the vector compression operation of the current stage. Among them, the vibration modal internal focusing vector formed by the vector compression operation of the current stage is formed based on the vibration compression vector formed by the vector compression operation of the previous stage (as described above). For example, the vibration modal internal focusing vector formed by the vector compression operation of the first stage can be modally internally focused with the sound modal vector of the vector compression operation of the first stage to form the sound modal internal focusing vector of the vector compression operation of the first stage.
[0094] Step c3, based on the sound modal internal focusing vector of the vector compression operation at the current stage, modal correlation focusing is performed with the blade sound vector to form the sound modal correlation focusing vector of the vector compression operation at the current stage.
[0095] In an embodiment of the present application, after obtaining the sound modal internal focusing vector of the vector compression operation at the current stage, the sound modal internal focusing vector of the vector compression operation at the current stage can be modally associated with the blade sound vector to form the sound modal associated focusing vector of the vector compression operation at the current stage. In this way, the semantic information in the blade sound vector can be further fused at the current stage so that the interference information is suppressed, thereby obtaining a sound modal associated focusing vector with higher semantic representation accuracy. Exemplarily, in order to reduce the amount of calculation, the transposed vector of the sound modal internal focusing vector of the vector compression operation at the current stage can be dot-producted with the blade sound vector to obtain the corresponding dot product parameter distribution, and based on the dot product parameter distribution, the sound modal internal focusing vector of the vector compression operation at the current stage can be weighted summed to obtain the sound modal associated focusing vector of the vector compression operation at the current stage. In this way, the association relationship between the two vectors can be characterized by the dot product parameter distribution, thereby achieving modal association focusing and further realizing the mining of associated information.
[0096] It is understood that, in the above step c2, the specific manner of forming the sound modal internal focusing vector of the vector compression operation at the current stage is not limited. For example, in an alternative embodiment, in order to achieve semantic fusion between the vibration modality and the sound modality during the process of internal modal focusing, the above step c2 may further include the following:
[0097] First, based on the internal focusing vector of the vibration mode formed by the vector compression operation in the current stage, the first mapping vector, the second mapping vector, and the third mapping vector corresponding to the vibration mode can be determined; illustratively, the corresponding neural network model carries a mapping matrix, which includes a first mapping sub-matrix, a second mapping sub-matrix, and a third mapping sub-matrix. In this way, the first mapping sub-matrix, the second mapping sub-matrix, and the third mapping sub-matrix can be multiplied by the internal focusing vector of the vibration mode respectively to obtain the first mapping vector, the second mapping vector, and the third mapping vector;
[0098] Secondly, the first mapping vector, the second mapping vector, and the third mapping vector corresponding to the sound mode can be determined based on the sound mode vector of the vector compression operation in the current stage; illustratively, the corresponding neural network model carries a mapping matrix (which may be different from the previous mapping matrix), and the mapping matrix may also include a first mapping sub-matrix, a second mapping sub-matrix, and a third mapping sub-matrix. In this way, the first mapping sub-matrix, the second mapping sub-matrix, and the third mapping sub-matrix can be multiplied with the sound mode vector respectively to obtain the first mapping vector, the second mapping vector, and the third mapping vector;
[0099] Then, the first mapping vector, the second mapping vector and the third mapping vector corresponding to the vibration mode can be subjected to a vector fusion operation with the first mapping vector, the second mapping vector and the third mapping vector corresponding to the sound mode to form a new vector to be processed, and the new vector to be processed can be subjected to modal internal focusing to form the sound modal internal focusing vector of the vector compression operation in the current stage; exemplarily, the first mapping vector corresponding to the vibration mode can be spliced with the first mapping vector corresponding to the sound mode to form a new first mapping vector, the second mapping vector corresponding to the vibration mode can be spliced with the second mapping vector corresponding to the sound mode to form a new second mapping vector, and the third mapping vector corresponding to the vibration mode can be spliced with the third mapping vector corresponding to the sound mode to form a new third mapping vector (that is, the new vector to be processed includes the new first mapping vector, the new second mapping vector and the new third mapping vector), and then, the dot product between the transposed vectors of the new first mapping vector and the new second mapping vector can be calculated to obtain the corresponding dot product parameter distribution, and then, based on the dot product parameter distribution, the new third mapping vector can be weighted summed to obtain the sound modal internal focusing vector.
[0100] It is understood that in the above step S123, the specific manner of semantically diffusing the diffuse sound vector is not limited. For example, in an alternative embodiment, in order to fully integrate the semantic information of the three modalities again during the semantic diffusion process, the above step S123 may further include step S123a, step S123b and step S123c. The specific contents of each step are as follows (combined with Figure 4 ).
[0101] Step S123a: performing X stages of vector expansion operations on the diffusion image vector, and determining the image expansion vector formed by the X-th stage of vector expansion operations as the image expansion target vector.
[0102] In an embodiment of the present application, the diffused image vector can be subjected to X stages of vector expansion operations, and the image expansion vector formed by the vector expansion operation of the Xth stage is determined as the image expansion target vector. For each stage of the vector expansion operation, the leaf image vector is fused based on the image expansion vector formed by the vector expansion operation of the previous stage to form the image expansion vector of the vector expansion operation of the current stage. That is, based on the image expansion vector formed by the vector expansion operation of the first stage, the leaf image vector is fused to form the image expansion vector of the vector expansion operation of the second stage; based on the image expansion vector formed by the vector expansion operation of the second stage, the leaf image vector is fused to form the image expansion vector of the vector expansion operation of the third stage, and so on and so forth, the image expansion vector formed by the vector expansion operation of the last stage can be obtained, and the image expansion vector is determined as the image expansion target vector. Based on this, the leaf image vectors can be fused through multiple stages, so that interference information is further suppressed, thereby obtaining a more reliable image expansion target vector. For example, in the vector expansion operation of each stage, the image expansion vector formed by the vector expansion operation of the previous stage (the first stage is the diffused image vector) can be deeply mined to achieve a size increase (for example, it can be achieved through a first expansion network, and the first expansion network can include multiple convolutional layers, batch normalization (Batch Normalization) and activation function (ReLU). It should be noted that the first expansion network and the aforementioned first compression network have different network parameters. In this way, the expansion and compression of the vector size can be achieved respectively. For example, the expansion of the vector size can be achieved through deconvolution or transposed convolution, and the compression of the vector size can be achieved through convolution with a step size greater than 1). Then, the image modality vector formed by deep mining can be modally focused (as described above) to obtain the image modality internal focusing vector. Furthermore, based on the leaf image vector, the image modality internal focusing vector can be modally associated focused (as described above) to obtain the image expansion vector of the vector expansion operation of the current stage.
[0103] Step S123b: performing X-stage vector expansion operations on the diffusion vibration vector, and determining the vibration expansion vector formed by the X-th stage vector expansion operation as the vibration expansion target vector.
[0104] In an embodiment of the present application, the diffuse vibration vector can be subjected to X stages of vector expansion operations, and the vibration expansion vector formed by the vector expansion operation of the Xth stage is determined as the vibration expansion target vector. For each stage of the vector expansion operation, the blade vibration vector is fused based on the vibration expansion vector formed by the vector expansion operation of the previous stage to form the vibration expansion vector of the vector expansion operation of the current stage. That is, based on the vibration expansion vector formed by the vector expansion operation of the first stage, the blade vibration vector is fused to form the vibration expansion vector of the vector expansion operation of the second stage; based on the vibration expansion vector formed by the vector expansion operation of the second stage, the blade vibration vector is fused to form the vibration expansion vector of the vector expansion operation of the third stage, and so on and so forth, the vibration expansion vector formed by the vector expansion operation of the last stage can be obtained, and this vibration expansion vector is determined as the vibration expansion target vector. Based on this, the blade vibration vectors can be fused through multiple stages, so that interference information is further suppressed, thereby obtaining a more reliable vibration expansion target vector. For example, in the vector expansion operation of each stage, the vibration expansion vector formed by the vector expansion operation of the previous stage (the first stage is the diffusion vibration vector) can be deeply mined to achieve a size increase (for example, this can be achieved through a second expansion network, which can include multiple convolutional layers, batch normalization (Batch Normalization) and activation function (ReLU)). Then, the vibration modal vector formed by the deep mining can be modally internally focused (as described above) to obtain a vibration modal internal focused vector. Furthermore, based on the blade vibration vector, the vibration modal internal focused vector can be modally correlated focused (as described above) to obtain the vibration expansion vector of the vector expansion operation of the current stage.
[0105] Step S123c: performing X-stage vector expansion operations on the diffuse sound vector, and determining the sound expansion vector formed by the X-th stage vector expansion operation as the fan blade multimodal vector.
[0106] In an embodiment of the present application, the diffuse sound vector can be subjected to X stages of vector expansion operations, and the sound expansion vector formed by the vector expansion operation of the Xth stage is determined as the fan blade multimodal vector. That is, based on the sound expansion vector formed by the vector expansion operation of the first stage, the blade sound vector is fused to form the sound expansion vector of the vector expansion operation of the second stage; based on the sound expansion vector formed by the vector expansion operation of the second stage, the blade sound vector is fused to form the sound expansion vector of the vector expansion operation of the third stage, and so on and so forth, the sound expansion vector formed by the vector expansion operation of the last stage can be obtained, and the sound expansion vector is determined as the fan blade multimodal vector. It should be noted that in each of the aforementioned stages, it is also necessary to fuse the semantic information in the image expansion vector formed by the vector expansion operation and the semantic information in the vibration expansion vector formed by the vector expansion operation, so that the fan blade multimodal vector can effectively represent the semantic information of all three modalities.
[0107] It can be understood that in the above-mentioned step S123c, the specific method of performing X-stage vector expansion operations on the diffuse sound vector is not limited. For example, in an alternative embodiment, in order to fully integrate the semantic information of the three modalities, and the sound expansion vector includes the sound modal vector formed during the vector expansion operation, the sound modal internal focus vector and the sound modal association focus vector, based on this, the above-mentioned step S123c can further include step c11, step c12 and step c13, the specific contents of which are described below.
[0108] In step c11, the sound modal association focus vector formed by the vector expansion operation in the previous stage is deeply mined to form the sound modal vector of the vector expansion operation in the current stage.
[0109] In an embodiment of the present application, the sound modal association focus vector formed by the vector expansion operation in the previous stage can be deeply mined to form the sound modal vector of the vector expansion operation in the current stage. For example, for the second stage of the vector expansion operation, the sound modal association focus vector formed by the vector expansion operation in the first stage can be deeply mined (as described above) to form the sound modal vector of the vector expansion operation in the first stage. For the first stage of the vector expansion operation, the diffuse sound vector can be deeply mined.
[0110] Step c12, performing modal internal focusing on the sound modal vector formed by the vector expansion operation of the current stage and the image modal vector in the image expansion vector formed by the vector expansion operation of the current stage, to form the sound modal internal focusing vector of the vector expansion operation of the current stage.
[0111] In an embodiment of the present application, after forming the sound modality vector for the current phase of the vector expansion operation, intra-modal focusing can be performed based on the sound modality vector formed by the current phase of the vector expansion operation and the image modality vector in the image expansion vector formed by the current phase of the vector expansion operation, thereby forming the sound modality intra-focus vector for the current phase of the vector expansion operation. In other words, the semantic information of the image modality can be pre-fused with the semantic information of the sound modality.
[0112] Step c13, based on the sound modal internal focusing vector formed by the vector expansion operation of the current stage, the vibration modal internal focusing vector in the vibration expansion vector formed by the vector expansion operation of the current stage, and the blade sound vector are modally associated and focused to form the sound modal associated focusing vector of the vector expansion operation of the current stage.
[0113] In an embodiment of the present application, after forming the sound modal internal focus vector for the current phase of the vector expansion operation, a modal correlation focus can be performed with the blade sound vector based on the sound modal internal focus vector formed by the current phase of the vector expansion operation and the vibration modal internal focus vector within the vibration expansion vector formed by the current phase of the vector expansion operation, thereby forming the sound modal correlation focus vector for the current phase of the vector expansion operation. In other words, semantic information of the vibration modality can be further fused into the vector that fuses semantic information of the image modality and semantic information of the sound modality, thereby achieving a gradual fusion of semantic information of different modalities.
[0114] It is understood that in the above-mentioned step c12, the specific method of forming the sound modal internal focus vector of the vector expansion operation in the current stage is not limited. For example, in an alternative embodiment, in order to fully mine the association information between different modal data in the intra-modal focus and make the semantic representation of the mined sound modal internal focus vector more accurate, the above-mentioned step c12 may further include the following:
[0115] First, based on the sound modal vector formed by the vector expansion operation in the current stage, the first mapping vector, the second mapping vector, and the third mapping vector corresponding to the sound modal and applied to the modal internal focusing can be determined; illustratively, the first mapping sub-matrix, the second mapping sub-matrix, and the third mapping sub-matrix (which may be different from the aforementioned sub-matrices, and the specific parameters may be formed during the training process) in the corresponding neural network model can be multiplied by the sound modal vector to obtain the corresponding first mapping vector, the second mapping vector, and the third mapping vector;
[0116] Secondly, a new first mapping vector can be formed based on the first mapping vector corresponding to the sound modality and applied to the intra-modal focusing and the first mapping vector corresponding to the image modality (for example, the two first mapping vectors can be added or averaged so that the semantic information between the sound modality and the image modality can be preliminarily fused), wherein the first mapping vector corresponding to the image modality is determined based on the image modality vector formed by the vector expansion operation in the current stage (e.g., multiplying the image modality vector by the corresponding mapping submatrix);
[0117] Then, the sound modal internal focusing vector of the vector expansion operation at the current stage can be determined based on the second mapping vector, the third mapping vector and the new first mapping vector corresponding to the sound modal and applied for internal modal focusing; exemplarily, the dot product between the transposed vectors of the new first mapping vector and the second mapping vector can be used to obtain the corresponding dot product parameter distribution, and then, based on the dot product parameter distribution, the third mapping vector can be weighted summed to obtain the sound modal internal focusing vector of the vector expansion operation at the current stage.
[0118] It is understood that in the above-mentioned step c13, the specific method of forming the sound modal internal focus vector of the vector expansion operation in the current stage is not limited. For example, in an alternative embodiment, in order to fully mine the association information between different modal data in the modal association focus, so that the semantic representation accuracy of the mined sound modal association focus vector is higher, the above-mentioned step c13 may further include the following:
[0119] First, based on the sound mode internal focusing vector of the vector expansion operation in the current stage, a first mapping vector corresponding to the sound mode and applied to modal correlation focusing can be determined. Also, based on the blade sound vector, a second mapping vector and a third mapping vector corresponding to the sound mode and applied to modal correlation focusing can be determined. As described above, the corresponding mapping vectors can be obtained by multiplying the mapping submatrices obtained by training in the neural network model.
[0120] Secondly, a new first mapping vector can be formed based on the first mapping vector corresponding to the sound mode and applied to modal association focusing and the first mapping vector corresponding to the vibration mode (e.g., by adding or averaging the two first mapping vectors), wherein the first mapping vector corresponding to the vibration mode is determined based on the vibration mode vector formed by the vector expansion operation in the current stage (i.e., multiplying the vibration mode vector by the corresponding mapping submatrix);
[0121] Then, a new third mapping vector can be formed based on the third mapping vector corresponding to the sound mode and applied to modal correlation focusing and the third mapping vector corresponding to the vibration mode (e.g., adding or averaging the two third mapping vectors), wherein the third mapping vector corresponding to the vibration mode is determined based on the blade vibration vector (i.e., multiplying the vibration mode vector by the corresponding mapping submatrix);
[0122] Finally, the sound modal associated focusing vector of the vector expansion operation at the current stage can be determined based on the second mapping vector applied to the modal internal focusing corresponding to the sound modal, the new third mapping vector and the new first mapping vector; illustratively, a dot product calculation can be performed on the new first mapping vector and the second mapping vector to obtain the corresponding dot product parameter distribution, and, based on the dot product parameter distribution, a weighted sum calculation can be performed on the new third mapping vector to obtain the corresponding sound modal associated focusing vector. Based on this, a dot product calculation and a weighted sum calculation can be used to fully fuse the three vectors: the sound modal internal focusing vector formed by the vector expansion operation at the current stage, the vibration modal internal focusing vector in the vibration expansion vector formed by the vector expansion operation at the current stage, and the blade sound vector. This can reduce the amount of computation required for fusing semantic information in the vectors to a certain extent, thereby improving computational efficiency.
[0123] Combine Figure 5 The present application also provides a wind turbine blade monitoring device based on multimodal analysis applicable to the electronic device described above. The wind turbine blade monitoring device based on multimodal analysis may include a data mining module, a semantic fusion module, and a fault prediction module.
[0124] The data mining module is used to mine the corresponding blade image vector, blade vibration vector and blade sound vector according to the predetermined image mode, vibration mode and sound mode based on the fan blade multimodal data corresponding to the target fan blade, wherein the fan blade multimodal data is formed by collecting image, vibration and sound information of the target fan blade. In the embodiment of the present application, the data mining module can be used to perform Figure 2 As shown in step S110, for the relevant content of the data mining module, reference may be made to the above description of step S110.
[0125] The semantic fusion module is used to fuse the blade image vector, the blade vibration vector and the blade sound vector into multimodal semantic information and output a fan blade multimodal vector. Figure 2 As shown in step S120, for the relevant content of the semantic fusion module, reference may be made to the above description of step S120.
[0126] The fault prediction module is used to predict and output fan blade fault data based on the fan blade multi-modal vector, wherein the fan blade fault data is used to characterize whether the target fan blade has a fault. In the embodiment of the present application, the fault prediction module can be used to perform Figure 2 For the step S130 shown, the relevant content can refer to the above description of step S130.
[0127] In an embodiment of the present application, corresponding to the above-mentioned wind blade monitoring method based on multimodal analysis applied to the electronic device, a computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is run, the various steps of the wind blade monitoring method based on multimodal analysis are executed.
[0128] The steps executed when the aforementioned computer program is running will not be described in detail here, and reference may be made to the above explanation of the wind turbine blade monitoring method based on multimodal analysis.
[0129] In summary, the wind turbine blade monitoring method, device, equipment and medium based on multimodal analysis provided by the present application mines the corresponding blade image vector, blade vibration vector and blade sound vector according to the predetermined image mode, vibration mode and sound mode based on the wind turbine blade multimodal data corresponding to the target wind turbine blade; secondly, the blade image vector, blade vibration vector and blade sound vector are subjected to multimodal semantic information fusion to output the wind turbine blade multimodal vector; then, based on the wind turbine blade multimodal vector, the wind turbine blade fault data is predicted and output. Based on the above content, since the multimodal data can be mined separately, the accuracy of semantic mining can be higher. In addition, by fusing the semantic vectors of multiple modes, a multimodal vector with stronger representation ability can be obtained, so that the semantic vector of a single mode can be strengthened. In this way, the reliability of prediction based on the multimodal vector can be guaranteed, thereby obtaining reliable wind turbine blade fault data. Therefore, the problem of relatively low reliability of wind turbine blade monitoring in the prior art can be improved.
[0130] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions. In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk. It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0131] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A wind turbine blade monitoring method based on multimodal analysis, characterized in that: include: Based on fan blade multimodal data corresponding to a target fan blade, mining a corresponding blade image vector, a blade vibration vector, and a blade sound vector according to predetermined image modes, vibration modes, and sound modes, respectively, wherein the fan blade multimodal data is formed by collecting image, vibration, and sound information of the target fan blade; According to the generated interference data, the corresponding interference image vector, interference vibration vector and interference sound vector are mined according to the image mode, the vibration mode and the sound mode respectively; according to the leaf image vector, the leaf vibration vector and the leaf sound vector, the interference image vector, the interference vibration vector and the interference sound vector are semantically diffused to output corresponding diffusion image vectors, diffusion vibration vectors and diffusion sound vectors; the diffusion image vector is subjected to X stages of vector expansion operations, wherein, for each stage of vector expansion operation, the leaf image vector is fused with the image expansion vector formed by the vector expansion operation of the previous stage to form the image expansion vector of the vector expansion operation of the current stage; The diffuse vibration vector is subjected to X stages of vector expansion operations, wherein, for each stage of the vector expansion operation, the blade vibration vector is fused based on the vibration expansion vector formed by the vector expansion operation in the previous stage to form a vibration expansion vector for the vector expansion operation in the current stage; the diffuse sound vector is subjected to X stages of vector expansion operations, and the sound expansion vector formed by the vector expansion operation in the Xth stage is determined as a wind blade multimodal vector, wherein, in the process of forming the wind blade multimodal vector, semantic information in the image expansion vector formed by the vector expansion operation and semantic information in the vibration expansion vector formed by the vector expansion operation are also fused, so that the wind blade multimodal vector represents semantic information of all three modalities; Based on the wind blade multi-modal vector, wind blade fault data is predicted and output, wherein the wind blade fault data is used to characterize whether the target wind blade has a fault.
2. The wind turbine blade monitoring method based on multimodal analysis according to claim 1, characterized in that: The step of performing semantic diffusion on the interference image vector, the interference vibration vector, and the interference sound vector based on the blade image vector, the blade vibration vector, and the blade sound vector, and outputting corresponding diffusion image vectors, diffusion vibration vectors, and diffusion sound vectors includes: Performing X stages of vector compression operations on the interference image vector, and determining an image compression vector formed by the vector compression operation of the Xth stage as a diffusion image vector, wherein for each stage of the vector compression operation, the leaf image vector is fused with the image compression vector formed by the vector compression operation of the previous stage to form an image compression vector of the vector compression operation of the current stage; Performing X stages of vector compression operations on the interference vibration vector, and determining a vibration compression vector formed by the vector compression operation of the Xth stage as a diffusion vibration vector, wherein for each stage of the vector compression operation, the blade vibration vector is fused based on the vibration compression vector and the image compression vector formed by the vector compression operation of the previous stage to form a vibration compression vector of the vector compression operation of the current stage; X stages of vector compression operations are performed based on the interference sound vector, and the sound compression vector formed by the vector compression operation of the Xth stage is determined as a diffuse sound vector, wherein for each stage of the vector compression operation, the blade sound vector is fused based on the sound compression vector formed by the vector compression operation of the previous stage and the vibration compression vector formed by the vector compression operation of the current stage to form the sound compression vector of the vector compression operation of the current stage.
3. The wind turbine blade monitoring method based on multimodal analysis according to claim 2, characterized in that: The step of fusing the blade vibration vector with the image compression vector formed by the vector compression operation in the previous stage to form the vibration compression vector of the vector compression operation in the current stage includes: Adding the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage to form a vector to be processed in the vector compression operation in the current stage; A vector compression operation is performed based on the to-be-processed vector of the vector compression operation in the current stage and the blade vibration vector to form a vibration compression vector of the vector compression operation in the current stage.
4. The wind turbine blade monitoring method based on multimodal analysis according to any one of claims 1 to 3, characterized in that: The sound expansion vector includes the sound mode vector, the sound mode internal focus vector and the sound mode correlation focus vector formed during the vector expansion operation. For each stage of the vector expansion operation: Deeply mine the sound modal correlation focus vector formed by the vector expansion operation in the previous stage to form the sound modal vector of the vector expansion operation in the current stage; Performing modal internal focusing on the sound modal vector formed by the vector expansion operation at the current stage and the image modal vector in the image expansion vector formed by the vector expansion operation at the current stage, thereby forming a sound modal internal focusing vector of the vector expansion operation at the current stage; Based on the sound modal internal focusing vector formed by the vector expansion operation of the current stage and the vibration modal internal focusing vector in the vibration expansion vector formed by the vector expansion operation of the current stage, modal correlation focusing is performed with the blade sound vector to form the sound modal correlation focusing vector of the vector expansion operation of the current stage.
5. The wind turbine blade monitoring method based on multimodal analysis according to claim 4, characterized in that: The step of performing modal internal focusing on the sound modal vector formed by the vector expansion operation at the current stage and the image modal vector in the image expansion vector formed by the vector expansion operation at the current stage to form the sound modal internal focusing vector of the vector expansion operation at the current stage includes: Determining, based on the sound mode vector formed by the vector expansion operation in the current stage, a first mapping vector, a second mapping vector, and a third mapping vector corresponding to the sound mode and applied to the modal internal focusing; forming a new first mapping vector based on the first mapping vector corresponding to the sound modality and the first mapping vector corresponding to the image modality, wherein the first mapping vector corresponding to the image modality is determined based on the image modality vector formed by the vector expansion operation in the current stage; The sound modal internal focusing vector of the vector expansion operation at the current stage is determined based on the second mapping vector, the third mapping vector and the new first mapping vector corresponding to the sound modal and applied to the internal focusing of the modal.
6. The wind turbine blade monitoring method based on multimodal analysis according to claim 4, characterized in that: The step of performing modal correlation focusing on the sound modal internal focusing vector formed according to the vector expansion operation at the current stage, the vibration modal internal focusing vector in the vibration expansion vector formed by the vector expansion operation at the current stage, and the blade sound vector to form the sound modal correlation focusing vector of the vector expansion operation at the current stage includes: Determining, based on the sound modal internal focusing vector of the vector expansion operation in the current stage, a first mapping vector corresponding to the sound modal and applied to modal correlation focusing, and, based on the blade sound vector, determining a second mapping vector and a third mapping vector corresponding to the sound modal and applied to modal correlation focusing; forming a new first mapping vector based on the first mapping vector corresponding to the sound mode and the first mapping vector corresponding to the vibration mode, wherein the first mapping vector corresponding to the vibration mode is determined based on the vibration mode vector formed by the vector expansion operation in the current stage; forming a new third mapping vector according to the third mapping vector corresponding to the sound mode and applied to modal correlation focusing and the third mapping vector corresponding to the vibration mode, wherein the third mapping vector corresponding to the vibration mode is determined based on the blade vibration vector; The sound modality-associated focusing vector of the vector expansion operation at the current stage is determined based on the second mapping vector corresponding to the sound modality and applied to the internal focusing of the modality, the new third mapping vector and the new first mapping vector.
7. A wind turbine blade monitoring device based on multimodal analysis, characterized in that: include: a data mining module for mining, based on fan blade multimodal data corresponding to a target fan blade, corresponding blade image vectors, blade vibration vectors, and blade sound vectors according to predetermined image modes, vibration modes, and sound modes, respectively, wherein the fan blade multimodal data is generated by collecting image, vibration, and sound information of the target fan blade; The semantic fusion module is used to mine the corresponding interference image vector, interference vibration vector and interference sound vector according to the generated interference data, respectively, according to the image mode, the vibration mode and the sound mode; perform semantic diffusion on the interference image vector, the interference vibration vector and the interference sound vector according to the leaf image vector, the leaf vibration vector and the leaf sound vector, and output the corresponding diffusion image vector, diffusion vibration vector and diffusion sound vector; perform X stages of vector expansion operations on the diffusion image vector, wherein, for each stage of vector expansion operation, the leaf image vector is fused with the image expansion vector formed by the vector expansion operation of the previous stage to form the image expansion vector of the vector expansion operation of the current stage vector; performing X stages of vector expansion operations on the diffuse vibration vector, wherein for each stage of the vector expansion operation, the blade vibration vector is fused based on the vibration expansion vector formed by the vector expansion operation of the previous stage to form a vibration expansion vector of the vector expansion operation of the current stage; performing X stages of vector expansion operations on the diffuse sound vector, and determining the sound expansion vector formed by the vector expansion operation of the Xth stage as a wind turbine blade multimodal vector, wherein, in the process of forming the wind turbine blade multimodal vector, semantic information in the image expansion vector formed by the vector expansion operation and semantic information in the vibration expansion vector formed by the vector expansion operation are also fused, so that the wind turbine blade multimodal vector represents the semantic information of all three modes; A fault prediction module is used to predict and output fan blade fault data based on the fan blade multi-modal vector, wherein the fan blade fault data is used to characterize whether the target fan blade has a fault.
8. An electronic device, characterized in that: include: memory for storing computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the wind turbine blade monitoring method based on multimodal analysis according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when running, executes the wind turbine blade monitoring method based on multimodal analysis according to any one of claims 1 to 6.
Citation Information
Patent Citations
High-frequency welded pipe weld defect detection method based on image data mining
CN120163777A
Fan blade state analysis method, device and equipment and storage medium
CN120251458A