Abnormal data recovery and risk early warning method and system for offshore wind turbine
By integrating multimodal monitoring data across different modes and utilizing a large visual language model to simultaneously perform data recovery and risk warning, the problems of inaccurate data recovery and imprecise warning in offshore wind power systems have been solved, achieving high-precision data repair and risk warning.
Patent Information
- Application Number
- CN202511202125.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, data recovery methods for offshore wind power systems rely on single-mode time-series modeling, ignoring frequency domain characteristics. This leads to inaccurate data recovery and imprecise risk warnings. Furthermore, the separation between recovery and warning results in low monitoring reliability, large errors, and delayed or false alarm rates in warnings.
Multimodal monitoring data, including time-domain sensor data and spectral images, is used to perform cross-modal fusion through a large visual language model. Frequency domain and temporal features are extracted, physical constraints are embedded, and abnormal data recovery and risk warning are performed simultaneously. The cross-modal fusion module of the large visual language model is used for data recovery and risk warning.
It achieves high-precision data repair and intelligent risk early warning, improves the accuracy of data recovery and the precision of early warning, reduces error accumulation and false alarm rate, and meets the needs of online real-time monitoring of offshore wind turbines.
Smart Images

Figure CN121034129A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent recovery and early warning, and particularly relates to a method and system for abnormal data recovery and risk early warning of offshore wind turbines. BACKGROUND
[0002] Offshore wind power systems are long-term exposed to harsh marine environments such as salt spray corrosion, seabed scouring, strong wind, wave impact, etc. The structural health monitoring system of the offshore wind power system often leads to the loss or abnormality of key time series data such as vibration, inclination, strain, etc. due to sensor failure, data transmission interruption, etc. In the prior art, data recovery mainly relies on numerical interpolation or pure time series prediction models such as LSTM, ARIMA, but such methods only use time domain sensor signals, ignoring the significant performance of fault features in the frequency domain such as resonance peak shift and harmonic anomaly, and cannot capture the frequency domain features, resulting in inaccurate data recovery.
[0003] The existing scheme usually separates data recovery and risk assessment into independent links, which may cause the recovered data to deviate from the real working condition due to the lack of consideration of physical constraints such as spectral continuity; the risk judgment is only based on threshold rules, which is difficult to mine potential associated risks from multi-modal data, resulting in inaccurate early warning.
[0004] These limitations result in low reliability of offshore wind turbine monitoring, large data recovery error, and high risk early warning lag or false positive rate, especially under harsh marine environments such as salt spray corrosion, seabed scouring, strong wind, wave impact, etc. SUMMARY
[0005] The present application provides a method and system for abnormal data recovery and risk early warning of offshore wind turbines to solve the problem of low reliability of offshore wind turbine monitoring, large data recovery error, and high risk early warning lag or false positive rate caused by the separation of single-modal time series modeling, risk early warning, and data recovery in the prior art, and to realize high-precision data repair and intelligent risk early warning.
[0006] The present application provides a method for abnormal data recovery and risk early warning of offshore wind turbines, comprising the following steps:
[0007] Real-time receiving of multi-modal monitoring data of offshore wind turbines, the multi-modal monitoring data comprising time domain sensor data, frequency spectrum images, and environmental parameters, the frequency spectrum images being obtained based on the time domain sensor data;
[0008] inputting the multi-modal monitoring data into a pre-trained visual language large model, wherein the spectral image is input into a visual encoder of the visual language large model for extracting frequency domain features; the time domain sensor data and the environmental parameters are input into a language encoder of the visual language large model for generating time sequence features and semantic features, the time sequence features being used for characterizing time domain dynamic modes, and the semantic features being used for characterizing environmental states and potential influences of the environmental states on device risks;
[0009] cross-modal fusion of the frequency domain features, the time sequence features and the semantic features is performed through a cross-modal fusion module of the visual language large model to obtain cross-modal fusion features, and based on the cross-modal fusion features, synchronous execution of abnormal data recovery and risk warning is performed to obtain recovered complete monitoring data and risk warning results.
[0010] According to the offshore wind turbine abnormal data recovery and risk warning method provided by the application, the synchronous execution of abnormal data recovery and risk warning specifically includes:
[0011] Synchronous execution S1-S2:
[0012] S1, based on a cross-modal attention mechanism to recover abnormal data and embed physical constraints;
[0013] S2, generating a risk warning signal based on multi-modal association.
[0014] According to the offshore wind turbine abnormal data recovery and risk warning method provided by the application, the embedding of the physical constraints includes a spectral continuity constraint, and the embedding mode of the physical constraints is:
[0015] In the visual language large model training stage, a preset loss function is added as a spectral continuity loss term.
[0016] According to the offshore wind turbine abnormal data recovery and risk warning method provided by the application, the extraction of the frequency domain features includes:
[0017] The spectral image is divided into N×N pixel blocks;
[0018] Each pixel block is linearly projected into a feature vector;
[0019] The salient regions in the spectral image are identified through a self-attention mechanism, and the salient regions include a frequency band energy mutation region, a resonance peak shift region and a harmonic abnormal region
[0020] The feature vectors of the salient regions are weighted to generate a spectral feature vector.
[0021] According to the offshore wind turbine abnormal data recovery and risk warning method provided by the application, the generation of the time sequence features includes:
[0022] The time-domain sensor data is sorted by time to generate original time series data;
[0023] The original time series data is segmented into fixed-length segments, each data point is converted into an embedding vector, and a position code is added to each embedding vector;
[0024] The long-range dependency is captured by the multi-layer self-attention mechanism of the language encoder, and the time series feature vector is output.
[0025] According to the offshore wind turbine abnormal data recovery and risk early warning method provided by the application, the generation of the semantic feature includes:
[0026] The environmental parameters are mapped into semantic embedding vectors through the fully connected layer and the word embedding layer of the language encoder;
[0027] The semantic embedding vector and the time series feature vector are spliced and input into the upper Transformer layer of the language encoder, and the semantic embedding vector and the time series feature are interacted through the cross-attention mechanism to obtain a semantic feature vector.
[0028] According to the offshore wind turbine abnormal data recovery and risk early warning method provided by the application, the generation of the risk early warning signal includes:
[0029] The cross-modal fusion feature is mapped to a risk level label, and the risk level label includes three levels of normal, low risk and high risk.
[0030] According to the offshore wind turbine abnormal data recovery and risk early warning method provided by the application, the pre-training of the visual language large model includes:
[0031] The general visual language large model is fine-tuned using the offshore wind turbine historical monitoring data set, and the historical monitoring data set contains labeled abnormal data segments and corresponding spectral image abnormal region masks; environmental parameters and risk level association labels.
[0032] According to the offshore wind turbine abnormal data recovery and risk early warning method provided by the application, the fine-tuning adopts a multi-task joint loss function:
[0033]
[0034] Wherein, is the data recovery loss, is the spectral continuity loss, is the risk classification loss, and α, β, γ are weight coefficients.
[0035] The application also provides an offshore wind turbine abnormal data recovery and risk early warning system, comprising the following modules:
[0036] A multi-modal acquisition unit: real-time acquisition of multi-modal monitoring data of the offshore wind turbine, the multi-modal monitoring data comprising time-domain sensor data, a frequency spectrum image and environmental parameters, the frequency spectrum image being converted based on the time-domain sensor data;
[0037] A visual language large model: comprising a visual encoder, a language encoder and a cross-modal fusion module, configured to execute the offshore wind turbine abnormal data recovery and risk early warning method according to any one of the above;
[0038] A joint output unit: for synchronously transmitting the recovered complete monitoring data and the risk early warning result to the wind turbine monitoring system.
[0039] The present application also provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the offshore wind turbine abnormal data recovery and risk early warning method according to any one of the above.
[0040] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the offshore wind turbine abnormal data recovery and risk early warning method according to any one of the above.
[0041] The present application also provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the offshore wind turbine abnormal data recovery and risk early warning method according to any one of the above.
[0042] The offshore wind turbine monitoring abnormal data recovery and risk early warning method and system provided by the present application, by real-time receiving multi-modal monitoring data of the offshore wind turbine, the multi-modal monitoring data comprising time-domain sensor data, a frequency spectrum image and environmental parameters, the frequency spectrum image being converted based on the time-domain sensor data; inputting the multi-modal monitoring data into a pre-trained visual language large model, wherein the frequency spectrum image is input into a visual encoder of the visual language large model for extracting frequency domain features; the time-domain sensor data and the environmental parameters are input into a language encoder of the visual language large model for generating time sequence features and semantic features, the time sequence features being used for characterizing time-domain dynamic modes, and the semantic features being used for characterizing environmental states and potential influences of the environmental states on equipment risks; performing cross-modal fusion of the frequency domain features, the time sequence features and the semantic features through a cross-modal fusion module of the visual language large model to obtain cross-modal fusion features, and synchronously executing abnormal data recovery and risk early warning based on the cross-modal fusion features to obtain recovered complete monitoring data and risk early warning results. The present application further expands time-domain signals of multiple sensors to the frequency domain, and utilizes a visual language large model to solve multi-modal data fusion of different sources, so as to realize high-precision data repair and risk intelligent early warning. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative effort.
[0044] Figure 1 is a flowchart of the offshore wind turbine abnormal data recovery and risk early warning method provided by the present application.
[0045] Figure 2 is a schematic diagram of time domain data converted into a spectrum diagram.
[0046] Figure 3 is a model structure of the visual language large model provided by the present application.
[0047] Figure 4 is a structural schematic diagram of the offshore wind turbine abnormal data recovery and risk early warning device provided by the present application.
[0048] Figure 5 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0049] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, but not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort belong to the scope of protection of the present application.
[0050] The present application will be described in detail below in combination with the drawings in the specification. The specific operation methods in the method embodiments can also be applied to the device embodiments or system embodiments. In the description of the present application, unless otherwise specified, "at least one" includes one or more. "Multiple" refers to two or more. For example, at least one of A, B and C includes: A alone, B alone, A and B together, A and C together, B and C together, and A, B and C together. In the present application, " / " means or, for example, A / B can mean A or B; "and / or" in this document is only a description of the association between the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases: A alone, A and B together, and B alone.
[0051] The present application will be described in detail below in combination with the specific embodiments.
[0052] In some embodiments of the present application, as shown in Figure 1 The present application provides an offshore wind turbine abnormal data recovery and risk early warning method, comprising:
[0053] Step 100, real-time receiving multi-modal monitoring data of offshore wind turbine, the multi-modal monitoring data including time domain sensor data, frequency spectrum image and environmental parameters, the frequency spectrum image being converted based on the time domain sensor data;
[0054] Step 200, inputting the multi-modal monitoring data into a pre-trained visual language large model, wherein the frequency spectrum image is input into a visual encoder of the visual language large model for extracting frequency domain features; the time domain sensor data and environmental parameters are input into a language encoder of the visual language large model for generating time sequence features and semantic features, the time sequence features being used for characterizing time domain dynamic mode, and the semantic features being used for characterizing environmental state and potential influence of the environmental state on device risk;
[0055] Step 300, performing cross-modal fusion of the frequency domain features, time sequence features and semantic features through a cross-modal fusion module of the visual language large model to obtain cross-modal fusion features, and simultaneously performing abnormal data recovery and risk early warning based on the cross-modal fusion features to obtain recovered complete monitoring data and risk early warning results.
[0056] It should be noted that the existing abnormal data recovery and risk early warning scheme only processes a single mode (time sequence or image), cannot simultaneously utilize time domain and frequency domain information, separates data recovery and risk assessment, and causes error accumulation.
[0057] Specifically, the existing scheme has insufficient utilization of frequency domain information, the existing data recovery method only processes time domain signals and does not fuse frequency spectrum images generated by FFT transformation, resulting in neglect of frequency domain fault features (such as edge frequency band caused by structural damage); there is a defect of multi-modal cooperation, the traditional risk early warning system relies on manual definition of time / frequency domain feature rules and cannot adaptively learn deep correlation between time domain waveform and frequency spectrum image (such as corresponding relationship between waveform distortion and frequency spectrum asymmetry); there is also a lack of dynamic risk prediction, the existing method does not introduce a risk feedback mechanism in the data recovery stage, for example, when the frequency spectrum features of the recovered data meet the tower tube connecting bolt loosening mode, the risk level adjustment is not triggered simultaneously, resulting in disconnection between recovery and early warning; and small sample adaptation is poor, offshore wind turbine fault samples are scarce, and traditional neural networks need a large amount of labeled data for training, while multi-modal large models can realize few-shot transfer learning through pre-training knowledge (such as general structural vibration frequency spectrum library).
[0058] Therefore, the application receives multi-modal data (time domain / frequency spectrum graph / environmental parameters) in real time, processes the data through a visual language large model VLM in a double-channel mode (a visual encoder extracts frequency domain features, and a language encoder generates time sequence / semantic features), and outputs data recovery and risk warning results synchronously through cross-modal fusion features. Through an end-to-end framework, error accumulation in traditional serial processes such as false data recovery leading to false warning is avoided, and the task fragmentation problem is solved. Through a spectrum image, missing information in the time domain is supplemented (for example, when a sensor fails, data can still be recovered through frequency domain features), and multi-modal complementarity enhances robustness. Moreover, the application can realize full-process coverage from "receiving data" to "outputting results", and meets the online real-time monitoring needs of offshore wind turbines.
[0059] Specifically, visual feature extraction of a spectrum graph and cross-modal joint reasoning are performed through a multi-modal large model (such as GPT-4V and qwenvl). The convolutional neural network (CNN) branch can automatically identify abnormal frequency band distribution in the spectrum graph (such as sudden increase in wideband noise and disappearance of characteristic frequency), which makes up for the frequency domain blind area of a pure time sequence model. The multi-modal alignment technology based on the Transformer (such as CLIP) can establish semantic association between the time domain signal and the spectrum graph, for example, associating sudden increase in time domain vibration amplitude with occurrence of high-order harmonics in the spectrum as a bolt loosening risk.
[0060] In some possible embodiments of the application, the synchronous execution of abnormal data recovery and risk warning specifically includes:
[0061] Synchronous execution S1-S2:
[0062] S1, based on a cross-modal attention mechanism, abnormal data is recovered and physical constraints are embedded;
[0063] S2, based on multi-modal association, a risk warning signal is generated.
[0064] Specifically, the embodiment provides an implementation of synchronous execution of abnormal data recovery and risk warning. Through cross-modal attention and physical constraint embedding, data recovery is realized at the same time, and through multi-modal association, a signal is generated to realize risk warning. Moreover, physical constraints (such as spectrum continuity) are embedded when data is recovered, which ensures that the output conforms to the real working condition, improves the warning accuracy, and realizes collaborative optimization. Single inference completes two tasks, which is faster than traditional step-by-step processing and has efficiency advantages.
[0065] Specifically, in the prior art, sensor data is usually received first, and then data recovery (such as interpolation / LSTM) is performed based on the data, and risk warning of threshold rules is performed by using the recovered data. However, in the above process, the recovery link ignores the physical constraints (such as spectral continuity), so that the recovered data may not conform to the real physical state (for example, the interpolated vibration data is not continuous in the frequency domain). The warning link only relies on the recovered data, and uses simple rules, and cannot utilize the deep correlation between the original multi-modal information (such as the frequency domain features in the image, the environmental text). Therefore, the recovered data is distorted, the risk judgment is wrong, and the warning is inaccurate or lagging.
[0066] To solve the defect of the prior art that the processing is fragmented, the embodiment synchronously performs abnormal data recovery and risk warning, forming two tasks that are strongly associated and work cooperatively. The association design of the embodiment directly aims at the defect, and optimizes the recovery and the warning cooperatively in the same framework.
[0067] In a possible embodiment, the unified features extracted by the cross-modal fusion module of the VLM contain information for data recovery (such as the context of missing points, spectral continuity constraints) and information for risk warning (such as fault feature patterns, multi-modal associated risks). The recovery process embeds physical constraints (such as the spectral continuity loss of weight 2), ensuring that the recovered data conforms to the physical laws of the real device (for example, the energy mutation of a certain frequency band caused by a loose bolt still retains the feature after recovery). This provides a real and reliable data basis for risk warning. The potential fault patterns (such as “tower bolt loosening”) recognized by the risk warning module can guide the data recovery in reverse.
[0068] For example, if the VLM detects a certain frequency band energy mutation (bolt loosening feature) in the spectrum graph, when recovering the missing vibration data, it will tend to generate reasonable values that conform to the fault pattern, rather than general interpolation. When the environmental parameter (such as strong wind) is identified as a high-risk factor, the recovered data will pay more attention to the abnormal patterns related to the strong wind in that period.
[0069] The VLM minimizes the loss of the two tasks (data recovery error and risk classification error) at the same time during training. This joint training forces the model to learn that:
[0070] The recovery needs to serve accurate warning (the recovered data needs to contain effective risk information);
[0071] The warning needs to rely on physically real recovered data (cannot generate “good-looking but useless” data that deviates from physical constraints).
[0072] The risk information (failure mode) as the "semantic guidance" of recovery, combined with physical constraints, makes the recovery data closer to the failure performance in the real physical state, rather than pure mathematical fitting, and the recovery accuracy is improved.
[0073] The early warning is based on the deep association (VLM mining) of the physically real recovery data and the original multi-modal features, and can identify:
[0074] Implicit risks that cannot be captured by a single modality (such as slight fluctuations in the time domain + harmonic abnormalities in a specific frequency band = early bolt loosening);
[0075] Coupling risks of environmental factors and structure state (such as high salt fog + specific spectral features = accelerated corrosion risk).
[0076] The early warning accuracy is improved.
[0077] End-to-end processing avoids error accumulation and delay in traditional serial processes, and the recovery result and risk state can be output in one inference, improving efficiency and real-time performance.
[0078] Multi-modal information complements each other, and even if a sensor completely fails (all data is missing), VLM can still use other modalities (such as images, environmental text) to infer recovery data and evaluate risks, enhancing robustness.
[0079] Under the unified VLM framework, by sharing features, embedding physical constraints, feedback of risk information, and joint optimization, recovery serves early warning, and early warning guides recovery, forming a closed loop to overcome the defects of "single modality modeling deficiency" and "task fragmentation".
[0080] In summary, the visual language large model fuses frequency domain features (spectrum graph), captures failure modes (such as frequency band energy mutation caused by bolt loosening) ignored by traditional time series models, improves the authenticity of recovery data, and further improves the accuracy of data recovery. Deep associations are mined from multi-modal data (such as the synergistic effect of time domain anomalies, frequency domain features, and environmental parameters), and semantic-based early warning is achieved, rather than simple threshold rules, reducing false positives and enhancing risk early warning accuracy. Through end-to-end design, the processing steps are reduced, and the real-time performance is improved; and the generalization ability of the visual language large model is enhanced, which has strong robustness in harsh environments (data missing / noise).
[0081] In some possible embodiments of the present application, the embedding of physical constraints includes a spectrum continuity constraint, and the embedding method of the physical constraints is:
[0082] In the visual language large model training phase, a preset loss function is added to a spectrum continuity loss term.
[0083] Specifically, the embodiment provides an implementation of embedding physical constraints, spectral continuity constraints are realized by adding a spectral continuity loss term to a loss function in a training phase, avoiding distortion in a frequency domain (such as splitting of a formant) caused by mathematical interpolation, and the constraints are internalized into a trained model, and no additional calculation is required during inference.
[0084] In some possible implementation manners of the present application, the extraction of the frequency domain feature comprises:
[0085] segmenting the spectral image into N*N pixel blocks;
[0086] linearly projecting each pixel block into a feature vector;
[0087] identifying a salient region in the spectral image through a self-attention mechanism, the salient region comprising a frequency band energy mutation region, a formant shift region and a harmonic anomaly region
[0088] weighting the feature vector of the salient region to generate a spectral feature vector.
[0089] Specifically, the embodiment provides an implementation of extraction of a frequency domain feature, which realizes interpretable fault diagnosis by locating a fault frequency band through attention weights, and enhances the noise resistance of a model by weighting and aggregating to suppress noise in a non-salient region (such as low-frequency clutter caused by wave interference).
[0090] In some possible implementation manners of the present application, the generation of the time sequence feature comprises:
[0091] sorting time domain sensor data according to time to generate original time sequence data;
[0092] segmenting the original time sequence data into fixed-length segments, converting each data point into an embedding vector, and adding position encoding to each embedding vector;
[0093] capturing long-range dependencies through a multi-layer self-attention mechanism of the language encoder to output a time sequence feature vector.
[0094] Specifically, the embodiment provides an implementation of generation of a time sequence feature, which solves long-term dependencies by expanding a vibration period through self-attention modeling, and restores missing points more accurately by retaining a time sequence context through position encoding, and adapts to a data missing situation.
[0095] In some possible implementation manners of the present application, the generation of the semantic feature comprises:
[0096] mapping an environment parameter into a semantic embedding vector through a fully connected layer and a word embedding layer of the language encoder;
[0097] The semantic embedding vector is spliced with the time sequence feature vector and input into an upper layer of a Transformer layer of a language encoder, and the semantic embedding vector and the time sequence feature are interacted through a cross attention mechanism to obtain a semantic feature vector.
[0098] Specifically, the embodiment provides an implementation of generation of a semantic feature, cross-modal correlation such as "salt mist concentration increase + vibration frequency spectrum harmonic increase prompts structure corrosion risk" is analyzed through environment-structure coupling, and dynamic feature enhancement is realized through cross attention weight distribution such as focusing on high-frequency vibration features in a strong wind environment.
[0099] In some possible implementation manners of the present application, the generation of the risk warning signal comprises:
[0100] The cross-modal fusion feature is mapped to a risk level label, and the risk level label comprises three levels of normal, low risk and high risk.
[0101] Specifically, the embodiment provides an implementation of generation of a risk warning signal, and risk levels are mapped based on cross-modal fusion features, so as to provide a basis for precise operation and maintenance decisions, such as optional planned maintenance for low risk and immediate shutdown for high risk. High risk can also be triggered only when multi-modal feature cooperation is abnormal, such as salt mist + harmonic combination, to reduce false positives.
[0102] In some possible implementation manners of the present application, the pre-training of the visual language large model comprises:
[0103] The general visual language large model is fine-tuned using a historical monitoring data set of offshore wind turbines, the historical monitoring data set comprising: labeled abnormal data segments and corresponding spectral image abnormal region masks; and associated labels of environmental parameters and risk levels.
[0104] In some possible implementation manners of the present application, the fine-tuning adopts a multi-task joint loss function:
[0105]
[0106] wherein, is a data recovery loss, is a spectral continuity loss, is a risk classification loss, and alpha, beta and gamma are weight coefficients.
[0107] Specifically, the embodiment provides another implementation of pre-training of a visual language large model, which uses a field-specific dataset for fine-tuning (including an abnormal area mask + risk label) and a multi-task loss function (restoration loss + spectral constraint loss + risk classification loss), fine-tuning adapts the general VLM to the fan monitoring scene (greatly improves the precision compared with directly using CLIP), and realizes field adaptability; the data restoration quality and risk sensitivity are adjusted through weight coefficients (a, b, g), and the optimization objectives are balanced.
[0108] As described in the above embodiment, the prior art has the following technical problems:
[0109] Problem one, insufficient multi-modal data fusion:
[0110] The prior art only relies on time domain sensor data for abnormal data restoration and risk assessment, and fails to effectively combine frequency domain information (such as FFT spectrum), resulting in the inability to comprehensively capture fault features (such as resonance frequency shift, harmonic anomaly, etc.).
[0111] Problem two, low accuracy of abnormal data restoration:
[0112] Traditional numerical interpolation or time series prediction models (such as LSTM, ARIMA) have large restoration errors in long-time data loss or high-noise environments, and cannot utilize spectral characteristics to optimize restoration results.
[0113] Problem three, risk warning and data restoration are separated:
[0114] Existing methods usually treat data restoration and risk assessment as independent processes, resulting in restored data that may not conform to real physical laws (such as spectral continuity constraints), and risk warnings rely on artificial rules, making it difficult to adaptively learn potential correlations from multi-modal data.
[0115] Problem four, poor model generalization ability in small sample scenarios:
[0116] Offshore wind turbine support structure damage samples are scarce, and traditional neural networks rely on large amounts of labeled data for training, making it difficult to adapt to the small sample learning needs in actual operation and maintenance.
[0117] Problem five, lack of dynamic risk prediction:
[0118] Existing methods fail to incorporate real-time risk feedback (such as spectral anomaly pattern matching) during data restoration, resulting in a disconnect between restoration and warning, and the inability to achieve closed-loop optimization.
[0119] Therefore, the application provides a multi-modal offshore wind turbine monitoring method based on a visual language large model (VLPM). Time domain sensor data (wind speed, deformation, vibration) and frequency domain spectrum analysis (FFT transformation) are combined to realize the collaborative optimization of abnormal data recovery and risk warning by using a multi-modal large model, thereby realizing high-precision data repair and intelligent risk warning.
[0120] The application will be described in detail below with reference to specific embodiments.
[0121] In some specific embodiments of the application, first, the multi-modal data needs to be preprocessed. Specifically, different sensor time domain data can be converted into natural language descriptions, such as "t time wind speed 12.5 m / s, X direction vibration acceleration 0.15 g, Y direction 0.12 g". The time domain data can be converted into a frequency spectrum diagram, which can be obtained by using a classic short-time Fourier transform, and the processed frequency spectrum diagram is reserved as the visual input of the large model, as shown in Figure 2 .
[0122] Further, for the training of the visual language large model, training data needs to be prepared. In this embodiment, three categories of data are obtained by introducing expert system labeling:
[0123] The first category is abnormal recovery task data preparation, such as:
[0124] {
[0125] {“img”: frequency spectrum diagram, “prompt”: “present a piece of time series multi-modal data {time series data}, please combine the corresponding frequency spectrum diagram, please locate the abnormal time period of the above data”}.
[0126] “label”: [start time] to [end time] has an abnormality, the data after repair conforming to the physical law is: {restored time series data}.} ... ...
[0129]
[0130] In possible embodiments, abnormal data can be obtained by adding noise or various disturbances to normal data.
[0131] The second category is risk prediction task data preparation, such as:
[0132] {
[0133] {“img”: frequency spectrum diagram, “prompt”: “present a piece of time series multi-modal data {time series data}, please combine the corresponding frequency spectrum diagram, please evaluate the current state risk level (L / M / H), and explain the judgment basis”}.
[0134] "label": There is a risk, and the corresponding basis is ... ...
[0137] }
[0138] Category 3, abnormal recovery + risk prediction task data preparation, such as:
[0139] {
[0140] {"img": spectrum graph, "prompt": "There is a piece of time series multi-modal data {time series data}, please analyze whether there is an abnormal period based on the corresponding spectrum graph, if there is, please give the repaired data, and evaluate the current state risk level based on the repaired data."}
[0141] "label": [start time] to [end time] has an abnormality, the repaired data that conforms to the physical law is: {repaired time series data}. The current data has a risk of **, and the corresponding basis is ... ...
[0144] }
[0145] It should be noted that even if the data anomaly is repaired, the corresponding spectrum graph is not updated. In actual use, the spectrum graph can be updated offline before using the data format in the second type of data for risk prediction.
[0146] Further, for the model structure of the visual language large model, as shown in Figure 3 , the model structure follows the common graph-text multi-modal large model structure, consisting of a visual encoding module (time domain signals of multiple sensors need to be converted into different spectrum graphs), a text encoding module, and a decoder. It can utilize the existing mass of pre-training knowledge in deepseek or qwenvl (inheriting the understanding ability of 300+ visual concepts and 100+ text relationships) to improve the ability in vertical scenes, thereby having better zero-shot transfer learning ability.
[0147] Finally, for model fine-tuning, LoRA can be used for efficient parameter fine-tuning, only fine-tuning the low-rank matrix of the attention layer.
[0148] Specifically, the training of the visual language large model can be performed in two stages:
[0149] First stage: training the abnormal data recovery task alone, i.e. the first type of data preparation;
[0150] Second stage: joint training of three tasks data, i.e. the first-3 type of data preparation.
[0151] The two-stage fine-tuning setting of the embodiment of the present application is considering the training difficulty imbalance caused by the task characteristic difference. Since the fan monitoring belongs to the vertical scene, the traditional large model does not have such knowledge reserve, and the vertical scene data is limited, in order to better realize the model fine-tuning, the present application proposes a two-stage fine-tuning strategy for the scene.
[0152] Specifically, for the abnormal data recovery task, it is a typical regression problem, and the model needs to learn accurate numerical mapping; and the risk prediction task is a hybrid task of classification and generation, which contains discrete decision information.
[0153] Therefore, the present application divides the model fine-tuning into two stages, the first stage focuses on the bottom layer feature: the model is forced to first establish a robust time-frequency feature representation, and learns through the data recovery task: physical constraints (such as the maximum amplitude of the vibration signal), frequency domain correlation (spectrum pattern of a specific fault), etc. The second stage introduces high-level reasoning: on the basis of good features, decision logic is constructed, so as to better complete the risk prediction task.
[0154] The present application combines time domain sensor data (wind speed, deformation, vibration) and frequency domain spectrum analysis (FFT transform), and uses a multi-modal large model to realize high-precision data repair and risk intelligent early warning. Specifically: the present application further expands the time domain signal of the multi-sensor to the frequency domain, and uses a picture-text multi-modal large model to solve the fusion of multi-modal data from different sources; the present application proposes a two-stage fine-tuning strategy for the two tasks of abnormal data repair and risk intelligent early warning to realize the learning of bottom layer features to high-level reasoning information in turn.
[0155] The offshore wind turbine abnormal data recovery and risk early warning system provided by the present application is described below, and the offshore wind turbine abnormal data recovery and risk early warning system described below can be correspondingly referred to the offshore wind turbine abnormal data recovery and risk early warning method described above.
[0156] In some specific embodiments of the present application, as shown in Figure 4 The present application provides an offshore wind turbine abnormal data recovery and risk early warning system, comprising:
[0157] The multi-modal acquisition unit 41: real-time acquisition of multi-modal monitoring data of the offshore wind turbine, the multi-modal monitoring data comprising time domain sensor data, frequency spectrum image and environmental parameters, the frequency spectrum image being obtained based on the time domain sensor data;
[0158] The visual language large model 42: comprising a visual encoder, a language encoder and a cross-modal fusion module, configured to execute any one of the offshore wind turbine abnormal data recovery and risk early warning methods described above;
[0159] The joint output unit 43 is configured to synchronously transmit the recovered complete monitoring data and the risk warning result to the wind turbine monitoring system.
[0160] The offshore wind turbine abnormal data recovery and risk warning system provided by the embodiments of the present application has similar implementation principles and beneficial effects to those of the offshore wind turbine abnormal data recovery and risk warning method described above, and reference can be made to the implementation principles and beneficial effects of the offshore wind turbine abnormal data recovery and risk warning method described above, which will not be repeated here.
[0161] Figure 5 An example of an entity structure diagram of an electronic device is shown in Figure 5 As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can invoke the logic instructions in the memory 530 to execute the offshore wind turbine abnormal data recovery and risk warning method, which includes: receiving real-time multi-modal monitoring data of an offshore wind turbine, the multi-modal monitoring data including time-domain sensor data, a frequency spectrum image, and environmental parameters, the frequency spectrum image being converted based on the time-domain sensor data; inputting the multi-modal monitoring data into a pre-trained visual language large model, wherein the frequency spectrum image is input into a visual encoder of the visual language large model to extract frequency domain features; the time-domain sensor data and the environmental parameters are input into a language encoder of the visual language large model to generate time sequence features and semantic features, the time sequence features being used to represent time-domain dynamic patterns, and the semantic features being used to represent environmental states and potential impacts of the environmental states on device risks; performing cross-modal fusion on the frequency domain features, the time sequence features, and the semantic features through a cross-modal fusion module of the visual language large model to obtain cross-modal fusion features, and synchronously performing abnormal data recovery and risk warning based on the cross-modal fusion features to obtain recovered complete monitoring data and risk warning results.
[0162] In addition, the logic instructions in the memory 530 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0163] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the offshore wind turbine abnormal data recovery and risk early warning method provided by the above-mentioned method, which comprises: receiving real-time multi-modal monitoring data of an offshore wind turbine, the multi-modal monitoring data comprising time domain sensor data, a frequency spectrum image and environmental parameters, the frequency spectrum image being obtained by converting the time domain sensor data; inputting the multi-modal monitoring data into a pre-trained visual language large model, wherein the frequency spectrum image is input into a visual encoder of the visual language large model to extract frequency domain features; the time domain sensor data and the environmental parameters are input into a language encoder of the visual language large model to generate time sequence features and semantic features, the time sequence features being used to represent time domain dynamic patterns, and the semantic features being used to represent environmental states and potential effects of the environmental states on device risks; performing cross-modal fusion on the frequency domain features, the time sequence features and the semantic features through a cross-modal fusion module of the visual language large model to obtain cross-modal fusion features, and simultaneously performing abnormal data recovery and risk early warning based on the cross-modal fusion features to obtain recovered complete monitoring data and risk early warning results.
[0164] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the offshore wind turbine abnormal data recovery and risk early warning method provided by each of the above methods, and the method comprises: receiving multi-modal monitoring data of an offshore wind turbine in real time, wherein the multi-modal monitoring data comprises time domain sensor data, a frequency spectrum image, and environmental parameters, and the frequency spectrum image is obtained by converting the time domain sensor data; inputting the multi-modal monitoring data into a pre-trained visual language large model, wherein the frequency spectrum image is input into a visual encoder of the visual language large model to extract frequency domain features; the time domain sensor data and the environmental parameters are input into a language encoder of the visual language large model to generate time sequence features and semantic features, the time sequence features are used to represent time domain dynamic modes, and the semantic features are used to represent environmental states and potential influences of the environmental states on device risks; performing cross-modal fusion on the frequency domain features, the time sequence features, and the semantic features through a cross-modal fusion module of the visual language large model to obtain cross-modal fusion features, and simultaneously performing abnormal data recovery and risk early warning based on the cross-modal fusion features to obtain recovered complete monitoring data and a risk early warning result.
[0165] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0166] From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and necessary general hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0167] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for recovering abnormal data and providing risk warning for offshore wind turbines, characterized in that, include: Real-time reception of multimodal monitoring data from offshore wind turbines, including time-domain sensor data, spectral images, and environmental parameters, wherein the spectral images are obtained based on the time-domain sensor data; The multimodal monitoring data is input into a pre-trained visual language model, wherein the spectral image is input into the visual encoder of the visual language model to extract frequency domain features; the temporal sensor data and environmental parameters are input into the language encoder of the visual language model to generate temporal features and semantic features, wherein the temporal features are used to characterize temporal dynamic patterns, and the semantic features are used to characterize environmental states and the potential impact of environmental states on equipment risks. The cross-modal fusion module of the visual language big model performs cross-modal fusion of the frequency domain features, temporal features and semantic features to obtain cross-modal fusion features. Based on the cross-modal fusion features, abnormal data recovery and risk warning are performed simultaneously to obtain the recovered complete monitoring data and risk warning results.
2. The method for recovering abnormal data and providing risk warning for offshore wind turbines according to claim 1, characterized in that, The synchronous execution of abnormal data recovery and risk warning specifically includes: Execute S1-S2 synchronously: S1. Recover anomalous data based on a cross-modal attention mechanism and embed physical constraints; S2. Generate risk warning signals based on multimodal correlation.
3. The method for recovering abnormal data and providing risk warning for offshore wind turbines according to claim 2, characterized in that, The embedded physical constraints include spectral continuity constraints, and the embedding method of the physical constraints is as follows: During the training phase of the large visual language model, a preset loss function is added to the spectral continuity loss term.
4. The method for recovering abnormal data and providing risk warning for offshore wind turbines according to claim 1, characterized in that, The extraction of the frequency domain features includes: The spectral image is divided into N×N pixel blocks; Each pixel block is linearly projected as a feature vector; The self-attention mechanism is used to identify salient regions in the spectral image, including regions of abrupt changes in frequency band energy, regions of resonant peak shifts, and regions of harmonic anomalies. The feature vectors of the salient regions are weighted to generate spectral feature vectors.
5. The method for recovering abnormal data and providing risk warning for offshore wind turbines according to claim 1, characterized in that, The generation of the temporal features includes: The time-domain sensor data is sorted according to time to generate raw time-series data; The original time-series data is divided into fixed-length segments, each data point is converted into an embedding vector, and position encoding is added to each embedding vector; The language encoder captures long-range dependencies through a multi-layer self-attention mechanism and outputs temporal feature vectors.
6. The method for recovering abnormal data and providing risk warning for offshore wind turbines according to claim 5, characterized in that, The generation of the semantic features includes: The environment parameters are mapped into semantic embedding vectors through the fully connected layer and word embedding layer of the language encoder; The semantic embedding vector and the temporal feature vector are concatenated and then input into the upper Transformer layer of the language encoder. Through the cross-attention mechanism, the semantic embedding vector and the temporal feature vector interact to obtain the semantic feature vector.
7. The method for recovering abnormal data and providing risk warning for offshore wind turbines according to claim 1, characterized in that, The generation of the risk warning signal includes: The cross-modal fusion features are mapped to risk level labels, which include three levels: normal, low risk, and high risk.
8. The method for recovering abnormal data and providing risk warning for offshore wind turbines according to claim 1, characterized in that, The pre-training of the large visual language model includes: The general visual language model was fine-tuned using a historical monitoring dataset of offshore wind turbines. The historical monitoring dataset includes: labeled abnormal data segments and corresponding spectral image abnormal region masks; and labels relating environmental parameters to risk levels.
9. The method for recovering abnormal data and providing risk warning for offshore wind turbines according to claim 1, characterized in that, The fine-tuning employs a multi-task joint loss function: in, To recover lost data, For the loss of spectral continuity, The risk classification loss is represented by α, β, and γ, which are weighting coefficients.
10. A system for recovering abnormal data and providing risk warning for offshore wind turbines, characterized in that, include: Multimodal acquisition unit: acquires multimodal monitoring data of offshore wind turbines in real time. The multimodal monitoring data includes time-domain sensor data, spectral images, and environmental parameters. The spectral images are obtained based on the time-domain sensor data. Visual language large model: including a visual encoder, a language encoder and a cross-modal fusion module, configured to execute the offshore wind turbine abnormal data recovery and risk warning method as described in any one of claims 1-9; Joint output unit: used to synchronously transmit complete monitoring data and risk warning results after recovery to the wind turbine monitoring system.
Citation Information
Cited By
Power grid power equipment anomaly prediction method, system and equipment and storage medium
CN121485293A
Intelligent operation and maintenance auxiliary system for water turbine of hydropower station
CN121810269A