Fan blade monitoring method and device based on multi-modal analysis, equipment and medium

The multi-modal analysis of wind turbine blades using image, vibration, and sound data improves fault detection reliability by integrating diverse data sources for a more accurate health assessment.

CN120312508AActive Publication Date: 2025-07-15HUANDIAN (FUJIAN) WIND POWER CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510550896.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-15
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Traditional fan blade monitoring methods rely on a single mode and cannot fully and accurately reflect the true status of the blades, resulting in low monitoring reliability.

Method used

The multimodal analysis method is adopted, combining image mode, vibration mode and sound mode, and through data mining and semantic information fusion, the fan blade multimodal vector is generated to predict blade fault data.

Benefits of technology

It improves the reliability and accuracy of fan blade monitoring, can more accurately identify blade faults, and reduces fault downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120312508A_ABST
    Figure CN120312508A_ABST
Patent Text Reader

Abstract

The invention provides a fan blade monitoring method and device based on multi-modal analysis, equipment and a medium, and relates to the technical field of fan blade monitoring. In the application, firstly, according to fan blade multi-modal data corresponding to a target fan blade, a corresponding blade image vector, a corresponding blade vibration vector and a corresponding blade sound vector are mined according to a predetermined image mode, a predetermined vibration mode and a predetermined sound mode; secondly, performing multi-modal semantic information fusion on the blade image vector, the blade vibration vector and the blade sound vector, and outputting a fan blade multi-modal vector; and then, according to the multi-modal vector of the fan blade, predicting and outputting fan blade fault data. On the basis of the content, the problem that in the prior art, the reliability of fan blade monitoring is relatively low can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of fan blade monitoring. Specifically, it relates to a fan blade monitoring method, device, equipment, and medium based on multimodal analysis. Background Art

[0002] With the rapid development of the wind power generation industry, the operating efficiency and stability of wind turbines have become crucial factors in the energy industry. As one of the core components of wind turbines, the performance of fan blades directly affects the efficiency of wind power generation. Therefore, accurately and real-time monitoring the health status of fan blades is of great significance for ensuring the stable operation of fans and reducing the fault downtime. Traditional fan blade monitoring methods usually rely on a single monitoring method, such as monitoring the vibration state of the blade through a vibration sensor to determine whether there is an abnormality. Although these methods can effectively provide some information for fault diagnosis, due to the complex and harsh working environment of fan blades, single-modal monitoring often cannot comprehensively and accurately reflect the true state of the blades. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a fan blade monitoring method, device, equipment, and medium based on multimodal analysis to improve the relatively low reliability of fan blade monitoring in the existing technology.

[0004] To achieve the above purpose, this application adopts the following technical solutions:

[0005] A fan blade monitoring method based on multimodal analysis, including:

[0006] According to the fan blade multimodal data corresponding to the target fan blade, respectively, according to the pre-determined image modality, vibration modality, and sound modality, extract the corresponding blade image vector, blade vibration vector, and blade sound vector, where the fan blade multimodal data is formed by collecting image, vibration, and sound information of the target fan blade;

[0007] Perform multimodal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and output a fan blade multimodal vector;

[0008] According to the fan blade multimodal vector, predict and output fan blade fault data, where the fan blade fault data is used to characterize whether there is a fault in the target fan blade.

[0009] In a preferred option of the present application, in the above-mentioned fan blade monitoring method based on multimodal analysis, the step of performing multimodal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector and outputting a multimodal vector of the fan blade includes:

[0010] According to the generated interference data, respectively, according to the image modality, the vibration modality, and the sound modality, extract the corresponding interference image vector, interference vibration vector, and interference sound vector;

[0011] According to the blade image vector, the blade vibration vector, and the blade sound vector, perform semantic diffusion on the interference image vector, the interference vibration vector, and the interference sound vector, and output the corresponding diffusion image vector, diffusion vibration vector, and diffusion sound vector;

[0012] According to the diffusion image vector and the diffusion vibration vector, perform semantic diffusion on the diffusion sound vector, and output a multimodal vector of the fan blade.

[0013] In a preferred option of the present application, in the above-mentioned fan blade monitoring method based on multimodal analysis, the step of performing semantic diffusion on the interference image vector, the interference vibration vector, and the interference sound vector according to the blade image vector, the blade vibration vector, and the blade sound vector, and outputting the corresponding diffusion image vector, diffusion vibration vector, and diffusion sound vector includes:

[0014] Perform vector compression operations on the interference image vector in X stages, and determine the image compression vector formed by the vector compression operation in the Xth stage as the diffusion image vector. For each stage of vector compression operation, according to the image compression vector formed by the vector compression operation in the previous stage, fuse the blade image vector to form the image compression vector of the current stage of vector compression operation;

[0015] Perform vector compression operations on the interference vibration vector in X stages, and determine the vibration compression vector formed by the vector compression operation in the Xth stage as the diffusion vibration vector. For each stage of vector compression operation, according to the vibration compression vector and image compression vector formed by the vector compression operation in the previous stage, fuse the blade vibration vector to form the vibration compression vector of the current stage of vector compression operation;

[0016] Perform vector compression operations in X stages based on the interference sound vector, and determine the sound compression vector formed by the vector compression operation in the Xth stage as the diffusion sound vector. For each stage of vector compression operation, based on the sound compression vector formed by the vector compression operation in the previous stage and the vibration compression vector formed by the vector compression operation in the current stage, fuse the blade sound vector to form the sound compression vector for the vector compression operation in the current stage.

[0017] In a preferred selection of this application, in the above-mentioned fan blade monitoring method based on multi-modal analysis, the step of fusing the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage with the blade vibration vector to form the vibration compression vector for the vector compression operation in the current stage includes:

[0018] Add the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage to calculate and form the vector to be processed for the vector compression operation in the current stage;

[0019] Perform a vector compression operation based on the vector to be processed for the vector compression operation in the current stage and the blade vibration vector to form the vibration compression vector for the vector compression operation in the current stage.

[0020] In a preferred selection of this application, in the above-mentioned fan blade monitoring method based on multi-modal analysis, the step of semantically diffusing the diffusion sound vector based on the diffusion image vector and the diffusion vibration vector and outputting the multi-modal vector of the fan blade includes:

[0021] Perform vector expansion operations on the diffusion image vector in X stages, and determine the image expansion vector formed by the vector expansion operation in the Xth stage as the image expansion target vector. For each stage of vector expansion operation, based on the image expansion vector formed by the vector expansion operation in the previous stage, fuse the blade image vector to form the image expansion vector for the vector expansion operation in the current stage;

[0022] Perform vector expansion operations on the diffusion vibration vector in X stages, and determine the vibration expansion vector formed by the vector expansion operation in the Xth stage as the vibration expansion target vector. For each stage of vector expansion operation, based on the vibration expansion vector formed by the vector expansion operation in the previous stage, fuse the blade vibration vector to form the vibration expansion vector for the vector expansion operation in the current stage;

[0023] Perform vector expansion operations on the diffusion sound vector in X stages, and determine the sound expansion vector formed by the vector expansion operation in the Xth stage as the multi-modal vector of the fan blade;

[0024] Among them, the sound expansion vector includes the sound modality vector, the internal focus vector of the sound modality, and the associated focus vector of the sound modality formed during the vector expansion operation. For the vector expansion operation at each stage:

[0025] Deeply mine the associated focus vector of the sound modality formed by the vector expansion operation in the previous stage to form the sound modality vector of the vector expansion operation in the current stage;

[0026] Based on the sound modality vector formed by the vector expansion operation in the current stage, perform internal focus within the modality with the image modality vector in the image expansion vector formed by the vector expansion operation in the current stage to form the internal focus vector of the sound modality of the vector expansion operation in the current stage;

[0027] Based on the internal focus vector of the sound modality formed by the vector expansion operation in the current stage, the internal focus vector of the vibration modality in the vibration expansion vector formed by the vector expansion operation in the current stage, and the blade sound vector, perform associated focus of modalities to form the associated focus vector of the sound modality of the vector expansion operation in the current stage.

[0028] In a preferred selection of the present application, in the above-mentioned fan blade monitoring method based on multi-modal analysis, the step of performing internal focus within the modality with the sound modality vector formed by the vector expansion operation in the current stage and the image modality vector in the image expansion vector formed by the vector expansion operation in the current stage to form the internal focus vector of the sound modality of the vector expansion operation in the current stage includes:

[0029] Based on the sound modality vector formed by the vector expansion operation in the current stage, determine the first mapping vector, the second mapping vector, and the third mapping vector applied to internal focus within the modality corresponding to the sound modality;

[0030] Based on the first mapping vector applied to internal focus within the modality corresponding to the sound modality and the first mapping vector corresponding to the image modality, form a new first mapping vector, where the first mapping vector corresponding to the image modality is determined based on the image modality vector formed by the vector expansion operation in the current stage;

[0031] Based on the second mapping vector, the third mapping vector applied to internal focus within the modality corresponding to the sound modality, and the new first mapping vector, determine the internal focus vector of the sound modality of the vector expansion operation in the current stage.

[0032] In a preferred option of the present application, in the above-mentioned fan blade monitoring method based on multi-modal analysis, the step of performing modal association focusing on the blade sound vector with the internal focusing vector of the sound modality formed by the vector expansion operation in the current stage and the internal focusing vector of the vibration modality in the vibration expansion vector formed by the vector expansion operation in the current stage to form the sound modality association focusing vector of the vector expansion operation in the current stage includes:

[0033] Determine a first mapping vector corresponding to the sound modality for modal association focusing according to the internal focusing vector of the sound modality of the vector expansion operation in the current stage, and determine a second mapping vector and a third mapping vector corresponding to the sound modality for modal association focusing according to the blade sound vector;

[0034] Form a new first mapping vector according to the first mapping vector corresponding to the sound modality for modal association focusing and the first mapping vector corresponding to the vibration modality, wherein the first mapping vector corresponding to the vibration modality is determined based on the vibration modality vector formed by the vector expansion operation in the current stage;

[0035] Form a new third mapping vector according to the third mapping vector corresponding to the sound modality for modal association focusing and the third mapping vector corresponding to the vibration modality, wherein the third mapping vector corresponding to the vibration modality is determined based on the blade vibration vector;

[0036] Determine the sound modality association focusing vector of the vector expansion operation in the current stage according to the second mapping vector corresponding to the sound modality for internal focusing, the new third mapping vector, and the new first mapping vector.

[0037] The present application also provides a fan blade monitoring device based on multi-modal analysis, including:

[0038] A data mining module, configured to respectively mine the corresponding blade image vector, blade vibration vector, and blade sound vector according to the fan blade multi-modal data corresponding to the target fan blade in the pre-determined image modality, vibration modality, and sound modality, wherein the fan blade multi-modal data is formed by collecting image, vibration, and sound information of the target fan blade;

[0039] A semantic fusion module, configured to perform multi-modal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and output a fan blade multi-modal vector;

[0040] A fault prediction module, configured to predict and output fan blade fault data according to the fan blade multi-modal vector, wherein the fan blade fault data is used to characterize whether the target fan blade has a fault.

[0041] Based on the above, the present application further provides an electronic device, including:

[0042] A memory for storing a computer program;

[0043] A processor connected to the memory for executing the computer program stored in the memory to implement the above-mentioned method for monitoring a wind turbine blade based on multimodal analysis.

[0044] Based on the above, the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program runs, it executes each step of the above-mentioned method for monitoring a wind turbine blade based on multimodal analysis.

[0045] The method, device, equipment and medium for monitoring a wind turbine blade based on multimodal analysis provided by the present application respectively extract corresponding blade image vectors, blade vibration vectors and blade sound vectors according to the pre-determined image modality, vibration modality and sound modality based on the multimodal data of the target wind turbine blade; secondly, perform multimodal semantic information fusion on the blade image vectors, blade vibration vectors and blade sound vectors to output a multimodal vector of the wind turbine blade; then, predict and output the fault data of the wind turbine blade based on the multimodal vector of the wind turbine blade. Based on the above, since multimodal data can be respectively mined, the accuracy of semantic mining can be higher. In addition, by fusing semantic vectors of multiple modalities, a multimodal vector with stronger representation ability can be obtained, so that the semantic vectors of a single modality can be strengthened. In this way, the reliability of prediction based on the multimodal vector can be ensured, and thus reliable fault data of the wind turbine blade can be obtained. Therefore, the problem of relatively low reliability in monitoring wind turbine blades existing in the prior art can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to make the above-mentioned objects, features and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and cooperates with the attached drawings for detailed description as follows.

[0047] Figure 1 It is a structural block diagram of the electronic device provided by the embodiment of the present application.

[0048] Figure 2 It is a schematic flowchart of the method for monitoring a wind turbine blade based on multimodal analysis provided by the embodiment of the present application.

[0049] Figure 3 It is a schematic diagram of semantic diffusion provided by the embodiment of the present application.

[0050] Figure 4 It is another schematic diagram of semantic diffusion provided by the embodiment of the present application.

[0051] Figure 5 It is a block diagram of the fan blade monitoring device based on multimodal analysis provided by the embodiments of the present application. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. Components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.

[0053] As Figure 1 shown, the embodiments of the present application provide an electronic device. Among them, the electronic device may include a memory, a processor, and a fan blade monitoring device based on multimodal analysis.

[0054] Specifically, the memory and the processor are directly or indirectly electrically connected to achieve data transmission or interaction. For example, the memory and the processor may be electrically connected through one or more communication buses or signal lines. The fan blade monitoring device based on multimodal analysis includes at least one software function module stored in the memory in the form of software or firmware. The processor is used to execute the executable computer programs stored in the memory, for example, the software function modules and computer programs included in the fan blade monitoring device based on multimodal analysis, etc., to implement the fan blade monitoring method based on multimodal analysis provided by the embodiments of the present application.

[0055] Optionally, the memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. And the processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a System on Chip (SoC), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0056] It can be understood that Figure 1 The structure shown is only for illustration, and the electronic device may also include more or fewer components than those shown Figure 1 herein, or have a different configuration from that shown Figure 1 herein. For example, it may also include a communication unit for information interaction with other devices.

[0057] In combination with Figure 2 , an embodiment of the present application further provides a multi-modal analysis-based fan blade monitoring method applicable to the above-mentioned electronic device. Among them, the method steps defined by the processes related to the multi-modal analysis-based fan blade monitoring method can be implemented by the electronic device. The following will elaborate in detail on Figure 2 the specific process shown.

[0058] Step S110: According to the fan blade multi-modal data corresponding to the target fan blade, respectively extract the corresponding blade image vector, blade vibration vector, and blade sound vector according to the pre-determined image mode, vibration mode, and sound mode.

[0059] In an embodiment of the present application, the electronic device may, according to the multi-modal data of the target wind turbine blade, respectively extract the corresponding blade image vector, blade vibration vector, and blade sound vector according to the pre-determined image modality, vibration modality, and sound modality. Among them, the multi-modal data of the wind turbine blade is formed by collecting image, vibration, and sound information of the target wind turbine blade. For example, an image acquisition device collects data in the image modality, a vibration sensor collects data in the vibration modality, and a sound collector collects data in the sound modality. Thus, after collecting data in each modality, extraction can be performed separately to extract the corresponding blade image vector, blade vibration vector, and blade sound vector, that is, extract the semantic information in the data of each modality and represent it in the form of a vector. Exemplarily, a convolutional network in a trained neural network model can be used to perform semantic extraction on the data of the three modalities respectively. Since the modalities of the data are different, in order to ensure the accuracy of semantic extraction, it can be implemented through three different convolutional networks. Among them, for the data in the sound modality, it can be first converted into a spectrogram, and then the spectrogram is subjected to convolutional processing to obtain the corresponding blade sound vector. Alternatively, convolutional processing can also be directly performed on the sound time series data to form the corresponding blade sound vector.

[0060] Step S120: Perform multi-modal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and output a multi-modal vector of the wind turbine blade.

[0061] In an embodiment of the present application, after extracting the blade image vector, the blade vibration vector, and the blade sound vector, the electronic device may perform multi-modal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and output a multi-modal vector of the wind turbine blade. Based on this, since the image modality may not be able to capture the details of the blade during high-speed rotation or micro-vibration, the vibration modality may not be able to effectively identify certain types of damage or cracks, and the sound modality may be affected by environmental noise, that is, the semantic vector of a single modality may not be able to fully characterize the faults of the target wind turbine blade. Therefore, through multi-modal semantic information fusion, the semantic vectors of each modality can be fully fused, so as to obtain a multi-modal vector of the wind turbine blade with a more reliable and sufficient semantic representation.

[0062] Step S130: Predict and output the fault data of the wind turbine blade according to the multi-modal vector of the wind turbine blade.

[0063] In the embodiment of the present application, after obtaining the multi-modal vector of the fan blade, the electronic device can predict and output the fan blade fault data according to the multi-modal vector of the fan blade. Wherein, the fan blade fault data is used to characterize whether the target fan blade has a fault. Exemplarily, the above neural network model may further include a fully connected network and an output network. Thus, after obtaining the multi-modal vector of the fan blade, the fully connected network can be used to process the multi-modal vector of the fan blade, so as to output a 1*2 fully connected vector. Then, through the output function (such as a classification function such as softmax) in the output network, the fully connected vector can be mapped into a probability distribution (a, b). For example, a corresponds to the probability of having a fault, and b corresponds to the probability of not having a fault. Then, the type corresponding to the larger probability is determined as the fan blade fault data.

[0064] Based on the above content, since multi-modal data can be mined separately, the accuracy of semantic mining can be higher. In addition, by fusing the semantic vectors of multiple modalities, a multi-modal vector with stronger representation ability can be obtained, so that the semantic vectors of a single modality can be strengthened. Thus, the reliability of prediction based on the multi-modal vector can be guaranteed, and reliable fan blade fault data can be obtained. Therefore, the problem of relatively low reliability in fan blade monitoring existing in the prior art can be improved.

[0065] Regarding the above step S120, it should be noted that the specific manner of performing multi-modal semantic information fusion is not limited and can be selected according to actual needs.

[0066] For example, in an alternative embodiment, in order to reduce the data processing volume, the blade image vector, the blade vibration vector, and the blade sound vector can be concatenated, and then self-attention processing, convolutional processing, etc. can be performed on the obtained concatenated vector to obtain the multi-modal vector of the fan blade.

[0067] For another example, in another alternative embodiment, in order to improve the reliability of multi-modal semantic information fusion, the above step S120 may further include step S121, step S122, and step S123. The specific content of each step is described as follows.

[0068] Step S121, according to the generated interference data, respectively dig out the corresponding interference image vector, interference vibration vector, and interference sound vector according to the image modality, the vibration modality, and the sound modality.

[0069] In the embodiments of the present application, according to the generated interference data, the corresponding interference image vectors, interference vibration vectors, and interference sound vectors can be mined respectively according to the image modality, the vibration modality, and the sound modality. Exemplarily, the interference data can be randomly generated data, or can be the multi-modal data of the fan blade collected in the history of the target fan blade. Thus, referring to the method of step S110, the interference data can be respectively mined for multi-modal data, so as to interfere and obtain the corresponding interference image vectors, interference vibration vectors, and interference sound vectors.

[0070] Step S122, according to the blade image vector, the blade vibration vector, and the blade sound vector, perform semantic diffusion on the interference image vector, the interference vibration vector, and the interference sound vector, and output the corresponding diffused image vector, diffused vibration vector, and diffused sound vector.

[0071] In the embodiments of the present application, after obtaining the interference image vector, the interference vibration vector, and the interference sound vector, according to the blade image vector, the blade vibration vector, and the blade sound vector, perform semantic diffusion on the interference image vector, the interference vibration vector, and the interference sound vector, and output the corresponding diffused image vector, diffused vibration vector, and diffused sound vector. That is to say, since the interference image vector, the interference vibration vector, and the interference sound vector are historical data and have a certain reference significance, however, the effectiveness of the semantic information represented will not be too high. Therefore, as a kind of interference data, while providing a certain reference significance, it can also avoid the overfitting problem that is likely to occur when directly fusing the blade image vector, the blade vibration vector, and the blade sound vector. Based on this, through semantic diffusion, the blade image vector, the blade vibration vector, and the blade sound vector can be fully fused to obtain diffused image vectors, diffused vibration vectors, and diffused sound vectors with higher semantic representation accuracy.

[0072] Step S123, according to the diffused image vector and the diffused vibration vector, perform semantic diffusion on the diffused sound vector, and output the multi-modal vector of the fan blade.

[0073] In the embodiments of the present application, after obtaining the diffusion image vector, the diffusion vibration vector, and the diffusion sound vector, the diffusion sound vector can be semantically diffused based on the diffusion image vector and the diffusion vibration vector to output the multi-modal vector of the fan blade. That is to say, the diffusion image vector and the diffusion vibration vector can be further fused into the diffusion sound vector, so that the semantic information in the diffusion sound vector can be strengthened based on the semantic information in the diffusion image vector and the diffusion vibration vector, thereby obtaining a multi-modal vector of the fan blade with richer semantic information.

[0074] It can be understood that in the above step S122, the specific manner of semantically diffusing the interference image vector, the interference vibration vector, and the interference sound vector is not limited. For example, in an alternative embodiment, in order to fully fuse the blade image vector, the blade vibration vector, and the blade sound vector representing the real semantic information into the interference image vector, the interference vibration vector, and the interference sound vector during the semantic diffusion process to achieve the suppression and removal of interference information, the above step S122 can further include the following steps S122a, S122b, and S122c, and the specific content of each step is as follows (in combination with Figure 3 ).

[0075] Step S122a: Perform a vector compression operation on the interference image vector for X stages, and determine the image compression vector formed by the vector compression operation in the Xth stage as the diffusion image vector.

[0076] In the embodiments of the present application, the interference image vector can be subjected to vector compression operations in X stages, and the image compression vector formed by the vector compression operation in the Xth stage is determined as the diffusion image vector. Among them, for the vector compression operation in each stage, based on the image compression vector formed by the vector compression operation in the previous stage, the blade image vector is fused to form the image compression vector of the vector compression operation in the current stage. That is to say, the image compression vector of the vector compression operation in the second stage can be formed based on the image compression vector formed by the vector compression operation in the first stage, and then, based on the image compression vector formed by the vector compression operation in the second stage, the image compression vector of the vector compression operation in the third stage is formed, and so on. In this way, the image compression vector formed by the vector compression operation in the last stage can be obtained, and then, the image compression vector formed by the vector compression operation in the last stage can be determined as the diffusion image vector. Among them, since the blade image vector is fused in the vector compression operation in each stage, in this way, the gradual integration of the blade image vector can be realized, that is, the gradual removal or suppression of interference information can be realized, so as to obtain a reliable diffusion image vector. In addition, by performing the vector compression operation, the size of the vector can also be reduced, so that the amount of data for subsequent processing is reduced. Exemplarily, the above neural network model may further include a first compression network, and the first compression network may include multiple convolutional layers, batch normalization (Batch Normalization), and activation function (ReLU). In this way, the corresponding vector can be compressed to reduce the size of the vector. For example, in the process of the vector compression operation in the first stage, the interference image vector is subjected to convolution, normalization, and activation processing (i.e., in-depth mining) to obtain the image modal vector of the vector compression operation in the first stage, and the image modal internal focusing is performed on the image modal vector of the vector compression operation in the first stage (reference can be made to the explanation of step b2 in the following text) to obtain the image modal internal focusing vector of the vector compression operation in the first stage, and, based on the blade image vector, the image modal internal focusing vector of the vector compression operation in the first stage is subjected to modal correlation focusing (reference can be made to the explanation of step c3 in the following text) to obtain the image modal correlation focusing vector of the vector compression operation in the first stage (i.e., the image compression vector formed by the vector compression operation in the first stage).During the vector compression operation in the second stage, the image modality correlation focus vector of the vector compression operation in the first stage will be (deeply mined) to obtain the image modality vector of the vector compression operation in the second stage, and the image modality vector of the vector compression operation in the second stage will be focused internally within the modality to obtain the image modality internal focus vector of the vector compression operation in the second stage. Moreover, based on the blade image vector, the image modality internal focus vector of the vector compression operation in the second stage will be subjected to modality correlation focus to obtain the image modality correlation focus vector of the vector compression operation in the second stage (i.e., the image compression vector formed by the vector compression operation in the second stage).

[0077] Step S122b: Perform vector compression operations on the interference vibration vector in X stages, and determine the vibration compression vector formed by the vector compression operation in the Xth stage as the diffusion vibration vector.

[0078] In the embodiment of the present application, the interference vibration vector can be subjected to vector compression operations in X stages, and the vibration compression vector formed by the vector compression operation in the Xth stage can be determined as the diffusion vibration vector. Among them, for the vector compression operation in each stage, based on the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage, the blade vibration vector is fused to form the vibration compression vector of the vector compression operation in the current stage. That is to say, the vibration compression vector formed by the vector compression operation in the second stage can be formed based on the vibration compression vector formed by the vector compression operation in the first stage, and then, the vibration compression vector formed by the vector compression operation in the third stage can be formed based on the vibration compression vector formed by the vector compression operation in the second stage, and so on. In this way, the vibration compression vector formed by the vector compression operation in the last stage can be obtained, and then, the vibration compression vector formed by the vector compression operation in the last stage can be determined as the diffusion vibration vector. Among them, since the blade vibration vector is fused in the vector compression operation in each stage, in this way, the gradual integration of the blade vibration vector can be realized, that is, the gradual removal or suppression of interference information can be realized, so as to obtain a reliable diffusion vibration vector. In addition, the image compression vector of the corresponding stage can also be fused, so that the semantic information of the image modality and the vibration modality can be fused, further improving the suppression effect on interference information, thereby ensuring the semantic representation accuracy of the obtained diffusion vibration vector.

[0079] Step S122c: Perform vector compression operations on the interference sound vector in X stages, and determine the sound compression vector formed by the vector compression operation in the Xth stage as the diffusion sound vector.

[0080] In the embodiments of the present application, X stages of vector compression operations can be performed based on the interference sound vector, and the sound compression vector formed by the vector compression operation in the Xth stage is determined as the diffusion sound vector. Among them, for the vector compression operation in each stage, based on the sound compression vector formed by the vector compression operation in the previous stage and the vibration compression vector formed by the vector compression operation in the current stage, the blade sound vector is fused to form the sound compression vector of the vector compression operation in the current stage. That is to say, the sound compression vector formed by the vector compression operation in the second stage can be formed based on the sound compression vector formed by the vector compression operation in the first stage, and then, based on the sound compression vector formed by the vector compression operation in the second stage, the sound compression vector formed by the vector compression operation in the third stage can be formed, and so on. In this way, the sound compression vector formed by the vector compression operation in the last stage can be obtained, and then, the sound compression vector formed by the vector compression operation in the last stage can be determined as the diffusion sound vector. Among them, since the blade sound vector is fused in the vector compression operation in each stage, in this way, the gradual integration of the blade sound vector can be realized, that is, the gradual removal or suppression of interference information can be realized, so as to obtain a reliable diffusion sound vector. Moreover, the vibration compression vector in the corresponding stage can also be fused (the semantic information in the image compression vector is also fused in the vibration compression vector), so that the semantic information of the image modality, vibration modality and sound modality can be fused, further improving the suppression effect on interference information, so as to ensure the semantic representation accuracy of the obtained diffusion sound vector.

[0081] It can be understood that in the above step S122b, the specific manner of forming the vibration compression vector of the vector compression operation in the current stage is not limited. For example, in an alternative embodiment, in order to achieve the full fusion of semantic information, the above step S122b can further include step b1 and step b2, and the specific content of each step is described as follows.

[0082] Step b1, add the vibration compression vector formed by the vector compression operation in the previous stage and the image compression vector to calculate and form the vector to be processed of the vector compression operation in the current stage.

[0083] In the embodiments of the present application, the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage can be added together to calculate and form the vector to be processed for the vector compression operation in the current stage. Exemplarily, in other embodiments, normalization processing can also be performed on the result of the addition calculation to obtain the corresponding vector to be processed. Among them, for the vector compression operation in the first stage, the interference vibration vector and the interference image vector can be added together to form the vector to be processed for the vector compression operation in the first stage. For the vector compression operation in the second stage, the vibration compression vector formed by the vector compression operation in the first stage and the image compression vector formed by the vector compression operation in the first stage can be added together for calculation.

[0084] Step b2: Perform a vector compression operation based on the vector to be processed for the vector compression operation in the current stage and the blade vibration vector to form the vibration compression vector for the vector compression operation in the current stage.

[0085] In the embodiments of the present application, after obtaining the vector to be processed for the vector compression operation in the current stage, a vector compression operation can be performed based on the vector to be processed for the vector compression operation in the current stage and the blade vibration vector to form the vibration compression vector for the vector compression operation in the current stage. In this way, the blade vibration vector representing the true semantic information in the interference information can be realized, so that the true semantic information can be continuously fused in each stage to obtain a reliable vibration compression vector.

[0086] It can be understood that in the above step b2, the specific manner of performing the vector compression operation based on the vector to be processed for the vector compression operation in the current stage and the blade vibration vector is not limited. For example, in an alternative embodiment, in order to achieve reliable fusion of the true semantic information during the compression process, the above step b2 can further include the following content:

[0087] First, the vector to be processed for the vector compression operation in the current stage can be deeply mined (as described above) to form the vibration mode vector in the vibration compression vector for the vector compression operation in the current stage, and, based on the vibration mode vector for the vector compression operation in the current stage, internal focusing within the mode is performed to form the internal focus vector of the vibration mode in the vibration compression vector for the vector compression operation in the current stage. Exemplarily, the transposed vector of the vibration mode vector for the vector compression operation in the current stage can be dot-product calculated with itself to obtain the corresponding dot-product parameter distribution, and, based on this dot-product parameter distribution, a weighted sum calculation is performed on the vibration mode vector for the vector compression operation in the current stage to obtain the internal focus vector of the vibration mode for the vector compression operation in the current stage. In this way, the correlation relationship between the two vectors can be characterized by the dot-product parameter distribution, thereby realizing internal focusing within the mode;

[0088] Secondly, based on the vibration mode internal focusing vector of the vector compression operation in the current stage, modal correlation focusing is performed with the blade vibration vector to form the vibration mode correlation focusing vector in the vibration compression vector of the vector compression operation in the current stage, that is, the correlation information between the vibration mode internal focusing vector and the blade vibration vector is mined, so that uncorrelated interference information is discarded, thereby obtaining a reliable vibration mode correlation focusing vector. The explanation in step c3 of the following text can be referred to.

[0089] It can be understood that in step S122c above, the specific manner of forming the sound compression vector of the vector compression operation in the current stage is not limited. For example, in an alternative embodiment, in order to achieve the full integration of semantic information, and the sound compression vector includes the sound mode vector, the sound mode internal focusing vector, and the sound mode correlation focusing vector in the process of the vector compression operation. Based on this, step S122c above can further include step c1, step c2, and step c3. The specific content of each step is described as follows.

[0090] Step c1: Deeply mine the sound mode correlation focusing vector formed by the vector compression operation in the previous stage to form the sound mode vector of the vector compression operation in the current stage.

[0091] In the embodiment of the present application, the sound mode correlation focusing vector formed by the vector compression operation in the previous stage is deeply mined to form the sound mode vector of the vector compression operation in the current stage. For example, the sound mode correlation focusing vector formed by the vector compression operation in the first stage can be deeply mined to form the sound mode vector of the vector compression operation in the second stage, and the sound mode correlation focusing vector formed by the vector compression operation in the second stage can be deeply mined to form the sound mode vector of the vector compression operation in the third stage. Exemplarily, for the first stage, the interference sound vector can be deeply mined (as described above).

[0092] Step c2: Based on the vibration mode internal focusing vector formed by the vector compression operation in the current stage, perform modal internal focusing with the sound mode vector of the vector compression operation in the current stage to form the sound mode internal focusing vector of the vector compression operation in the current stage.

[0093] In the embodiment of the present application, after obtaining the vibration mode internal focusing vector formed by the vector compression operation in the current stage, the vibration mode internal focusing vector formed by the vector compression operation in the current stage can be used to perform internal focusing on the sound mode vectors of the vector compression operation in the current stage to form the sound mode internal focusing vector of the vector compression operation in the current stage. Among them, the vibration mode internal focusing vector formed by the vector compression operation in the current stage is formed based on the vibration compression vector formed by the vector compression operation in the previous stage (as described in the relevant content above). For example, the vibration mode internal focusing vector formed by the vector compression operation in the first stage can be used to perform internal focusing on the sound mode vectors of the vector compression operation in the first stage to form the sound mode internal focusing vector of the vector compression operation in the first stage.

[0094] Step c3: Based on the sound mode internal focusing vector of the vector compression operation in the current stage, perform modal correlation focusing with the blade sound vector to form the sound mode correlation focusing vector of the vector compression operation in the current stage.

[0095] In the embodiment of the present application, after obtaining the sound mode internal focusing vector of the vector compression operation in the current stage, the sound mode internal focusing vector of the vector compression operation in the current stage can be used to perform modal correlation focusing with the blade sound vector to form the sound mode correlation focusing vector of the vector compression operation in the current stage. In this way, in the current stage, the semantic information in the blade sound vector can be further fused, so that the interference information is suppressed, and thus a sound mode correlation focusing vector with higher semantic representation accuracy can be obtained. Exemplarily, in order to reduce the computational complexity, the dot product calculation can be performed on the transposed vector of the sound mode internal focusing vector of the vector compression operation in the current stage and the blade sound vector to obtain the corresponding dot product parameter distribution, and based on this dot product parameter distribution, the weighted sum calculation is performed on the sound mode internal focusing vector of the vector compression operation in the current stage to obtain the sound mode correlation focusing vector of the vector compression operation in the current stage. In this way, the correlation relationship between the two vectors can be characterized by the dot product parameter distribution, so as to achieve modal correlation focusing and further realize the mining of correlation information.

[0096] It can be understood that in step c2 above, the specific manner of forming the sound mode internal focusing vector of the vector compression operation in the current stage is not limited. For example, in an alternative embodiment, in order to achieve semantic fusion between the vibration mode and the sound mode during the internal focusing of the mode, step c2 above can further include the following content:

[0097] First, the first mapping vector, the second mapping vector, and the third mapping vector corresponding to the vibration mode can be determined based on the internal focusing vector of the vibration mode formed by the vector compression operation in the current stage. Exemplarily, a mapping matrix is carried in the corresponding neural network model, and the mapping matrix includes a first mapping sub-matrix, a second mapping sub-matrix, and a third mapping sub-matrix. Thus, the first mapping sub-matrix, the second mapping sub-matrix, and the third mapping sub-matrix can be multiplied by the internal focusing vector of the vibration mode respectively to obtain the first mapping vector, the second mapping vector, and the third mapping vector.

[0098] Second, the first mapping vector, the second mapping vector, and the third mapping vector corresponding to the sound mode can be determined based on the sound mode vector of the vector compression operation in the current stage. Exemplarily, a mapping matrix (which can be different from the previous mapping matrix) is carried in the corresponding neural network model, and the mapping matrix can also include a first mapping sub-matrix, a second mapping sub-matrix, and a third mapping sub-matrix. Thus, the first mapping sub-matrix, the second mapping sub-matrix, and the third mapping sub-matrix can be multiplied by the sound mode vector respectively to obtain the first mapping vector, the second mapping vector, and the third mapping vector.

[0099] Then, the first mapping vector, the second mapping vector, and the third mapping vector corresponding to the vibration mode can be subjected to a vector fusion operation with the first mapping vector, the second mapping vector, and the third mapping vector corresponding to the sound mode to form a new vector to be processed. Moreover, the new vector to be processed is internally focused on the mode to form the internal focusing vector of the sound mode of the vector compression operation in the current stage. Exemplarily, the first mapping vector corresponding to the vibration mode can be concatenated with the first mapping vector corresponding to the sound mode to form a new first mapping vector, the second mapping vector corresponding to the vibration mode can be concatenated with the second mapping vector corresponding to the sound mode to form a new second mapping vector, and the third mapping vector corresponding to the vibration mode can be concatenated with the third mapping vector corresponding to the sound mode to form a new third mapping vector (that is to say, the new vector to be processed includes the new first mapping vector, the new second mapping vector, and the new third mapping vector). Then, the dot product between the new first mapping vector and the transposed vector of the new second mapping vector can be calculated to obtain the corresponding dot product parameter distribution. After that, based on the dot product parameter distribution, a weighted sum calculation is performed on the new third mapping vector to obtain the internal focusing vector of the sound mode.

[0100] It can be understood that in the above step S123, the specific manner of performing semantic diffusion on the diffusion sound vector is not limited. For example, in an alternative implementation manner, in order to fully fuse the semantic information of the three modalities again during the semantic diffusion process, the above step S123 may further include step S123a, step S123b, and step S123c. The specific content of each step is described as follows (in combination with Figure 4 ).

[0101] Step S123a: Perform vector expansion operations on the diffusion image vector for X stages, and determine the image expansion vector formed by the vector expansion operation in the Xth stage as the image expansion target vector.

[0102] In an embodiment of the present application, the diffusion image vector may be subjected to vector expansion operations in X stages, and the image expansion vector formed by the vector expansion operation in the Xth stage is determined as the image expansion target vector. Among them, for the vector expansion operation in each stage, based on the image expansion vector formed by the vector expansion operation in the previous stage, the blade image vector is fused to form the image expansion vector of the vector expansion operation in the current stage. That is to say, based on the image expansion vector formed by the vector expansion operation in the first stage, the blade image vector is fused to form the image expansion vector of the vector expansion operation in the second stage; based on the image expansion vector formed by the vector expansion operation in the second stage, the blade image vector is fused to form the image expansion vector of the vector expansion operation in the third stage, and so on. In this way, the image expansion vector formed by the vector expansion operation in the last stage can be obtained and determined as the image expansion target vector. Based on this, the blade image vector can be fused through multiple stages, so that interference information is further suppressed, and a more reliable image expansion target vector can be obtained. For example, in the vector expansion operation in each stage, the image expansion vector formed by the vector expansion operation in the previous stage (the diffusion image vector in the first stage) can be deeply mined to increase the size (for example, it can be achieved through a first expansion network, which may include multiple convolutional layers, batch normalization (Batch Normalization), and activation function (ReLU). It should be noted that the first expansion network has different network parameters from the aforementioned first compression network, so that the expansion and compression of the vector size can be realized respectively. For example, the expansion of the vector size can be achieved through deconvolution or transposed convolution, and the compression of the vector size can be achieved through convolution with a stride greater than 1). Then, the image modality vector formed by the deep mining can be subjected to in-modal focusing (as described above), so as to obtain the in-modal focusing vector of the image modality. Further, based on the blade image vector, the in-modal focusing vector of the image modality can be subjected to modal correlation focusing (as described above), so as to obtain the image expansion vector of the vector expansion operation in the current stage.

[0103] Step S123b: Perform vector expansion operations on the diffusion vibration vector in X stages, and determine the vibration expansion vector formed by the vector expansion operation in the Xth stage as the vibration expansion target vector.

[0104] In the embodiment of the present application, the diffusion vibration vector can be subjected to a vector expansion operation in X stages, and the vibration expansion vector formed by the vector expansion operation in the Xth stage is determined as the vibration expansion target vector. Among them, for the vector expansion operation in each stage, based on the vibration expansion vector formed by the vector expansion operation in the previous stage, the blade vibration vector is fused to form the vibration expansion vector of the vector expansion operation in the current stage. That is to say, based on the vibration expansion vector formed by the vector expansion operation in the first stage, the blade vibration vector is fused to form the vibration expansion vector of the vector expansion operation in the second stage; based on the vibration expansion vector formed by the vector expansion operation in the second stage, the blade vibration vector is fused to form the vibration expansion vector of the vector expansion operation in the third stage, and so on. In this way, the vibration expansion vector formed by the vector expansion operation in the last stage can be obtained and determined as the vibration expansion target vector. Based on this, the blade vibration vector can be fused through multiple stages, so that the interference information is further suppressed, and a more reliable vibration expansion target vector can be obtained. For example, in the vector expansion operation in each stage, the vibration expansion vector formed by the vector expansion operation in the previous stage (the diffusion vibration vector in the first stage) can be deeply mined to achieve an increase in size (for example, it can be realized through a second expansion network, which can include multiple convolutional layers, batch normalization (Batch Normalization), and activation function (ReLU)). Then, the vibration mode vector formed by the deep mining can be subjected to in-mode focusing (as described above), so as to obtain the in-mode focused vibration vector. Further, based on the blade vibration vector, the in-mode focused vibration vector can be subjected to mode correlation focusing (as described above), so as to obtain the vibration expansion vector of the vector expansion operation in the current stage.

[0105] Step S123c: Perform a vector expansion operation on the diffusion sound vector in X stages, and determine the sound expansion vector formed by the vector expansion operation in the Xth stage as the multi-modal vector of the fan blade.

[0106] In the embodiment of the present application, the diffused sound vector can be subjected to vector expansion operations in X stages, and the sound expansion vector formed by the vector expansion operation in the Xth stage is determined as the multi-modal vector of the fan blade. That is to say, based on the sound expansion vector formed by the vector expansion operation in the first stage, the blade sound vector is fused to form the sound expansion vector of the vector expansion operation in the second stage; based on the sound expansion vector formed by the vector expansion operation in the second stage, the blade sound vector is fused to form the sound expansion vector of the vector expansion operation in the third stage, and so on. In this way, the sound expansion vector formed by the vector expansion operation in the last stage can be obtained, and this sound expansion vector is determined as the multi-modal vector of the fan blade. It should be noted that in each of the foregoing stages, it is also necessary to fuse the semantic information in the image expansion vector formed by the vector expansion operation and the semantic information in the vibration expansion vector formed by the vector expansion operation, so that the multi-modal vector of the fan blade can effectively represent the semantic information of the three modalities.

[0107] It can be understood that in step S123c described above, the specific manner of performing vector expansion operations on the diffused sound vector in X stages is not limited. For example, in an alternative embodiment, in order to fully fuse the semantic information of the three modalities, and the sound expansion vector includes the sound modality vector, the internal focus vector of the sound modality, and the associated focus vector of the sound modality formed during the vector expansion operation. Based on this, step S123c described above can further include step c11, step c12, and step c13, and the specific content is as follows.

[0108] Step c11: Deeply mine the associated focus vector of the sound modality formed by the vector expansion operation in the previous stage to form the sound modality vector of the vector expansion operation in the current stage.

[0109] In the embodiment of the present application, the associated focus vector of the sound modality formed by the vector expansion operation in the previous stage can be deeply mined to form the sound modality vector of the vector expansion operation in the current stage. For example, for the vector expansion operation in the second stage, the associated focus vector of the sound modality formed by the vector expansion operation in the first stage can be deeply mined (as described above) to form the sound modality vector of the vector expansion operation in the first stage. For the vector expansion operation in the first stage, the diffused sound vector can be deeply mined.

[0110] Step c12: Perform internal focus within the modality on the sound modality vector formed by the vector expansion operation in the current stage and the image modality vector in the image expansion vector formed by the vector expansion operation in the current stage to form the internal focus vector of the sound modality of the vector expansion operation in the current stage.

[0111] In the embodiment of the present application, after the sound modal vector of the vector expansion operation at the current stage is formed, the sound modal vector formed by the vector expansion operation at the current stage and the image modal vector in the image expansion vector formed by the vector expansion operation at the current stage can be used for modal internal focusing to form the sound modal internal focusing vector of the vector expansion operation at the current stage. In other words, the semantic information of the image modality can be fused with the semantic information of the sound modality first.

[0112] Step c13, based on the sound modal internal focusing vector formed by the vector expansion operation of the current stage, the vibration modal internal focusing vector in the vibration expansion vector formed by the vector expansion operation of the current stage, perform modal association focusing with the blade sound vector to form the sound modal association focusing vector of the vector expansion operation of the current stage.

[0113] In the embodiment of the present application, after forming the sound modal internal focusing vector of the vector expansion operation at the current stage, the sound modal internal focusing vector formed by the vector expansion operation at the current stage and the vibration modal internal focusing vector in the vibration expansion vector formed by the vector expansion operation at the current stage can be modally associated with the blade sound vector to form the sound modal associated focusing vector of the vector expansion operation at the current stage. In other words, the semantic information of the vibration modality can be further integrated into the vector that integrates the semantic information of the image modality and the semantic information of the sound modality, so that the gradual integration of the semantic information of different modalities can be achieved.

[0114] It can be understood that, in the above step c12, the specific manner of forming the sound modal internal focus vector of the vector expansion operation at the current stage is not limited. For example, in an alternative implementation, in order to fully mine the association information between different modal data in the modal internal focus, so that the semantic representation accuracy of the mined sound modal internal focus vector is higher, the above step c12 may further include the following contents:

[0115] First, according to the sound modal vector formed by the vector expansion operation at the current stage, the first mapping vector, the second mapping vector and the third mapping vector corresponding to the sound modal and applied to the modal internal focusing can be determined; illustratively, the first mapping sub-matrix, the second mapping sub-matrix and the third mapping sub-matrix (which may be different from the aforementioned sub-matrix, and the specific parameters may be formed during the training process) in the corresponding neural network model may be multiplied with the sound modal vector to obtain the corresponding first mapping vector, the second mapping vector and the third mapping vector;

[0116] Second, a new first mapping vector can be formed based on the first mapping vector corresponding to the sound modality applied to within-modal focusing and the first mapping vector corresponding to the image modality (exemplarily, the two first mapping vectors can be added or averaged so that the semantic information between the sound modality and the image modality can be preliminarily fused). Among them, the first mapping vector corresponding to the image modality is determined based on the image modality vector formed by the vector expansion operation at the current stage (such as multiplying the image modality vector by the corresponding mapping submatrix);

[0117] Then, based on the second mapping vector, the third mapping vector corresponding to the sound modality applied to within-modal focusing, and the new first mapping vector, the within-modal focusing vector of the sound modality for the vector expansion operation at the current stage can be determined; exemplarily, the dot product between the new first mapping vector and the transposed vector of the second mapping vector can be calculated to obtain the corresponding dot product parameter distribution, and then, based on this dot product parameter distribution, the third mapping vector can be weighted and summed to obtain the within-modal focusing vector of the sound modality for the vector expansion operation at the current stage.

[0118] It can be understood that in the above step c13, the specific manner of forming the within-modal focusing vector of the sound modality for the vector expansion operation at the current stage is not limited. For example, in an alternative implementation, in order to fully extract the correlation information between different modality data in the modality association focusing and make the semantic representation accuracy of the extracted sound modality association focusing vector higher, the above step c13 can further include the following content:

[0119] First, based on the within-modal focusing vector of the sound modality for the vector expansion operation at the current stage, the first mapping vector corresponding to the sound modality applied to modality association focusing can be determined, and based on the blade sound vector, the second mapping vector and the third mapping vector corresponding to the sound modality applied to modality association focusing can be determined. As described above, the corresponding mapping vectors can be obtained through multiplication operations with the mapping submatrices trained in the neural network model;

[0120] Second, a new first mapping vector can be formed based on the first mapping vector corresponding to the sound modality applied to modality association focusing and the first mapping vector corresponding to the vibration modality (such as adding or averaging the two first mapping vectors). Among them, the first mapping vector corresponding to the vibration modality is determined based on the vibration modality vector formed by the vector expansion operation at the current stage (that is, multiplying the vibration modality vector by the corresponding mapping submatrix);

[0121] Then, a new third mapping vector can be formed based on the third mapping vector corresponding to the sound modality applied to modality correlation focusing and the third mapping vector corresponding to the vibration modality (such as adding the two third mapping vectors or calculating the mean value). Among them, the third mapping vector corresponding to the vibration modality is determined based on the blade vibration vector (that is, multiplying the vibration modality vector by the corresponding mapping sub-matrix).

[0122] Finally, based on the second mapping vector corresponding to the sound modality applied to modality internal focusing, the new third mapping vector, and the new first mapping vector, the sound modality correlation focusing vector for the vector expansion operation at the current stage can be determined. Exemplarily, the dot product of the new first mapping vector and the second mapping vector can be calculated to obtain the corresponding dot product parameter distribution, and based on this dot product parameter distribution, the weighted sum of the new third mapping vector can be calculated to obtain the corresponding sound modality correlation focusing vector. Based on this, through one dot product calculation and one weighted sum calculation, it is possible to fully fuse the three vectors: the sound modality internal focusing vector formed by the vector expansion operation at the current stage, the vibration modality internal focusing vector in the vibration expansion vector formed by the vector expansion operation at the current stage, and the blade sound vector, which can reduce the computational complexity when fusing the semantic information in the vectors to a certain extent and improve the computational efficiency.

[0123] Combined with Figure 5 , the embodiment of the present application also provides a fan blade monitoring device based on multimodal analysis that can be applied to the above-mentioned electronic device. Among them, the fan blade monitoring device based on multimodal analysis can include a data mining module, a semantic fusion module, and a fault prediction module.

[0124] The data mining module is used to respectively extract the corresponding blade image vector, blade vibration vector, and blade sound vector according to the fan blade multimodal data corresponding to the target fan blade in the pre-determined image modality, vibration modality, and sound modality. Among them, the fan blade multimodal data is formed by collecting image, vibration, and sound information of the target fan blade. In the embodiment of the present application, the data mining module can be used to execute Figure 2 the steps S110 shown, and the relevant content of the data mining module can be referred to the description of step S110 above.

[0125] The semantic fusion module is used to perform multimodal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and output the fan blade multimodal vector. In the embodiment of the present application, the semantic fusion module can be used to execute Figure 2 the steps S120 shown, and the relevant content of the semantic fusion module can be referred to the description of step S120 above.

[0126] The fault prediction module is configured to predict and output the fault data of the wind turbine blade based on the multi-modal vector of the wind turbine blade, wherein the fault data of the wind turbine blade is used to characterize whether there is a fault in the target wind turbine blade. In the embodiment of the present application, the fault prediction module can be used to execute Figure 2 the steps S130 shown, and the relevant content can be referred to the description of step S130 in the foregoing text.

[0127] In the embodiment of the present application, corresponding to the above-mentioned method for monitoring a wind turbine blade based on multi-modal analysis applied to the electronic device, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, and when the computer program runs, it executes each step of the method for monitoring a wind turbine blade based on multi-modal analysis.

[0128] Among them, the steps executed when the foregoing computer program runs will not be elaborated one by one here, and reference can be made to the foregoing explanation of the method for monitoring a wind turbine blade based on multi-modal analysis.

[0129] In summary, the method, device, equipment and medium for monitoring a wind turbine blade based on multi-modal analysis provided by the present application respectively extract the corresponding blade image vector, blade vibration vector and blade sound vector according to the image modality, vibration modality and sound modality determined in advance based on the multi-modal data of the target wind turbine blade; secondly, perform multi-modal semantic information fusion on the blade image vector, blade vibration vector and blade sound vector to output the multi-modal vector of the wind turbine blade; then, predict and output the fault data of the wind turbine blade based on the multi-modal vector of the wind turbine blade. Based on the above content, since multi-modal data can be respectively mined, the accuracy of semantic mining can be higher. In addition, by fusing the semantic vectors of multiple modalities, a multi-modal vector with stronger representation ability can be obtained, so that the semantic vector of a single modality can be strengthened. In this way, the reliability of prediction based on the multi-modal vector can be ensured, and thus reliable fault data of the wind turbine blade can be obtained. Therefore, the problem of relatively low reliability of wind turbine blade monitoring existing in the prior art can be improved.

[0130] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions. In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs. It should be noted that in this document, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0131] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A monitoring method for wind turbine blades based on multimodal analysis, characterized in that, Including: According to the multi-modal data of the target wind turbine blade, respectively extract the corresponding blade image vector, blade vibration vector, and blade sound vector according to the pre-determined image modality, vibration modality, and sound modality, where the multi-modal data of the wind turbine blade is formed by collecting image, vibration, and sound information of the target wind turbine blade; Perform multi-modal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and output a multi-modal vector of the wind turbine blade; Predict and output wind turbine blade fault data according to the multi-modal vector of the wind turbine blade, where the wind turbine blade fault data is used to characterize whether the target wind turbine blade has a fault.

2. The method for monitoring a wind turbine blade based on multimodal analysis according to claim 1, wherein The step of performing multi-modal semantic information fusion on the blade image vector, the blade vibration vector, and the blade sound vector, and outputting a multi-modal vector of the wind turbine blade includes: According to the generated interference data, respectively extract the corresponding interference image vector, interference vibration vector, and interference sound vector according to the image modality, the vibration modality, and the sound modality; According to the blade image vector, the blade vibration vector, and the blade sound vector, perform semantic diffusion on the interference image vector, the interference vibration vector, and the interference sound vector, and output the corresponding diffused image vector, diffused vibration vector, and diffused sound vector; According to the diffused image vector and the diffused vibration vector, perform semantic diffusion on the diffused sound vector, and output a multi-modal vector of the wind turbine blade.

3. The method for monitoring a wind turbine blade based on multimodal analysis according to claim 2, wherein, The step of performing semantic diffusion on the interference image vector, the interference vibration vector, and the interference sound vector according to the blade image vector, the blade vibration vector, and the blade sound vector, and outputting the corresponding diffused image vector, diffused vibration vector, and diffused sound vector includes: Perform vector compression operations on the interference image vector in X stages, and determine the image compression vector formed by the vector compression operation in the Xth stage as the diffused image vector. For each stage of vector compression operation, according to the image compression vector formed by the previous stage of vector compression operation, fuse the blade image vector to form the image compression vector of the current stage of vector compression operation; Perform vector compression operations on the interference vibration vector in X stages, and determine the vibration compression vector formed by the vector compression operation in the Xth stage as the diffused vibration vector. For each stage of vector compression operation, according to the vibration compression vector and the image compression vector formed by the previous stage of vector compression operation, fuse the blade vibration vector to form the vibration compression vector of the current stage of vector compression operation; Perform vector compression operations in X stages based on the interference sound vector, and determine the sound compression vector formed by the vector compression operation in the Xth stage as the diffusion sound vector. For each stage of the vector compression operation, based on the sound compression vector formed by the vector compression operation in the previous stage and the vibration compression vector formed by the vector compression operation in the current stage, fuse the blade sound vector to form the sound compression vector of the vector compression operation in the current stage.

4. The method for monitoring the fan blade based on multimodal analysis according to claim 3, wherein The step of fusing the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage with the blade vibration vector to form the vibration compression vector of the vector compression operation in the current stage includes: Add the vibration compression vector and the image compression vector formed by the vector compression operation in the previous stage to calculate and form the vector to be processed of the vector compression operation in the current stage; Perform a vector compression operation based on the vector to be processed of the vector compression operation in the current stage and the blade vibration vector to form the vibration compression vector of the vector compression operation in the current stage.

5. The method for monitoring a wind turbine blade based on multimodal analysis according to any one of claims 2-4, wherein The step of semantically diffusing the diffusion sound vector based on the diffusion image vector and the diffusion vibration vector to output the multi-modal vector of the fan blade includes: Perform vector expansion operations in X stages on the diffusion image vector, and determine the image expansion vector formed by the vector expansion operation in the Xth stage as the image expansion target vector. For each stage of the vector expansion operation, based on the image expansion vector formed by the vector expansion operation in the previous stage, fuse the blade image vector to form the image expansion vector of the vector expansion operation in the current stage; Perform vector expansion operations in X stages on the diffusion vibration vector, and determine the vibration expansion vector formed by the vector expansion operation in the Xth stage as the vibration expansion target vector. For each stage of the vector expansion operation, based on the vibration expansion vector formed by the vector expansion operation in the previous stage, fuse the blade vibration vector to form the vibration expansion vector of the vector expansion operation in the current stage; Perform vector expansion operations in X stages on the diffusion sound vector, and determine the sound expansion vector formed by the vector expansion operation in the Xth stage as the multi-modal vector of the fan blade; Among them, the sound expansion vector includes the sound modality vector, the internal focus vector of the sound modality, and the associated focus vector of the sound modality formed during the vector expansion operation. For each stage of the vector expansion operation: Deeply mine the associated focus vector of the sound modality formed by the vector expansion operation in the previous stage to form the sound modality vector of the vector expansion operation in the current stage; Based on the sound modality vector formed by the vector expansion operation in the current stage, perform internal focus of the modality with the image modality vector in the image expansion vector formed by the vector expansion operation in the current stage to form the internal focus vector of the sound modality of the vector expansion operation in the current stage; The internal focusing vector of the sound modality formed according to the vector expansion operation in the current stage, and the internal focusing vector of the vibration modality in the vibration expansion vector formed by the vector expansion operation in the current stage are associated and focused with the blade sound vector to form the sound modality associated focusing vector of the vector expansion operation in the current stage.

6. The method for monitoring a wind turbine blade based on multimodal analysis according to claim 5, wherein, The step of performing internal focusing of modalities on the sound modality vector formed according to the vector expansion operation in the current stage and the image modality vector in the image expansion vector formed by the vector expansion operation in the current stage to form the internal focusing vector of the sound modality of the vector expansion operation in the current stage includes: Based on the sound modality vector formed according to the vector expansion operation in the current stage, determine the first mapping vector, the second mapping vector, and the third mapping vector applied to internal focusing of modalities corresponding to the sound modality; Based on the first mapping vector applied to internal focusing of modalities corresponding to the sound modality and the first mapping vector corresponding to the image modality, form a new first mapping vector, where the first mapping vector corresponding to the image modality is determined based on the image modality vector formed according to the vector expansion operation in the current stage; Based on the second mapping vector, the third mapping vector applied to internal focusing of modalities corresponding to the sound modality, and the new first mapping vector, determine the internal focusing vector of the sound modality of the vector expansion operation in the current stage.

7. The method for monitoring a wind turbine blade based on multimodal analysis according to claim 5, characterized in that The step of associating and focusing the internal focusing vector of the sound modality formed according to the vector expansion operation in the current stage, and the internal focusing vector of the vibration modality in the vibration expansion vector formed by the vector expansion operation in the current stage with the blade sound vector to form the sound modality associated focusing vector of the vector expansion operation in the current stage includes: Based on the internal focusing vector of the sound modality of the vector expansion operation in the current stage, determine the first mapping vector applied to modal association and focusing corresponding to the sound modality, and based on the blade sound vector, determine the second mapping vector and the third mapping vector applied to modal association and focusing corresponding to the sound modality; Based on the first mapping vector applied to modal association and focusing corresponding to the sound modality and the first mapping vector corresponding to the vibration modality, form a new first mapping vector, where the first mapping vector corresponding to the vibration modality is determined based on the vibration modality vector formed according to the vector expansion operation in the current stage; Based on the third mapping vector applied to modal association and focusing corresponding to the sound modality and the third mapping vector corresponding to the vibration modality, form a new third mapping vector, where the third mapping vector corresponding to the vibration modality is determined based on the blade vibration vector; Based on the second mapping vector applied to internal focusing of modalities corresponding to the sound modality, the new third mapping vector, and the new first mapping vector, determine the sound modality associated focusing vector of the vector expansion operation in the current stage.

8. A monitoring device for fan blades based on multimodal analysis, characterized in that, Including: A data mining module, configured to respectively mine corresponding blade image vectors, blade vibration vectors, and blade sound vectors according to a pre-determined image modality, vibration modality, and sound modality based on the multi-modal data of the target wind turbine blade, wherein the multi-modal data of the wind turbine blade is formed by collecting image, vibration, and sound information of the target wind turbine blade; A semantic fusion module, configured to perform multi-modal semantic information fusion on the blade image vectors, the blade vibration vectors, and the blade sound vectors, and output a multi-modal vector of the wind turbine blade; A fault prediction module, configured to predict and output wind turbine blade fault data based on the multi-modal vector of the wind turbine blade, wherein the wind turbine blade fault data is used to characterize whether there is a fault in the target wind turbine blade.

9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor connected to the memory, configured to execute the computer program stored in the memory to implement the wind turbine blade monitoring method based on multi-modal analysis according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program runs, it executes the wind turbine blade monitoring method based on multi-modal analysis according to any one of claims 1-7.

Citation Information

Patent Citations

  • Fan blade state detection method and system based on multi-modal data fusion

    CN116123040A

  • Multi-sensor fusion and image recognition combined blade fault detection method and system

    CN119534636A

  • High-frequency welded pipe weld defect detection method based on image data mining

    CN120163777A

  • Fan blade state analysis method, device and equipment and storage medium

    CN120251458A

  • Workpiece surface morphology generation method and apparatus based on multimodal image generation

    WO2025060317A1