New energy equipment fault diagnosis method based on multi-modal data and AI large model
By using AI large models for time alignment and quality labeling in the multimodal data diagnosis of new energy equipment, and generating credibility and risk indices, the problem of alignment reliability and risk assessment in multimodal diagnosis is solved, and efficient fault diagnosis decision-making and operation and maintenance optimization are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINTA ZHONGGUANG SOLAR POWER CO LTD
- Filing Date
- 2025-12-10
- Publication Date
- 2026-04-10
AI Technical Summary
Existing multimodal new energy equipment diagnostic technologies struggle to explicitly quantify cross-modal alignment reliability under asynchronous and incomplete data conditions. Diagnostic credibility and fault risk do not achieve dual-domain synergy, and threshold strategies are difficult to adaptively adjust, leading to false alarms and unnecessary downtime.
By performing time alignment and quality labeling on multimodal monitoring data within a preset diagnostic time window, a pre-trained AI large model diagnostic network is used to generate fault candidate labels and model uncertainty, construct a diagnostic credibility index and a fault risk index, form a credibility-risk dual-domain hierarchical matrix and dynamic threshold, and make hierarchical diagnostic decisions.
It achieves the goal of maintaining safety priorities while reducing false alarms and unnecessary downtime, improving the traceability and on-site adaptability of diagnostic conclusions, and optimizing the cost-benefit ratio of operation and maintenance.
Smart Images

Figure CN121834546A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of new energy equipment state monitoring and intelligent fault diagnosis, and particularly relates to a new energy equipment fault diagnosis method based on multi-modal data and an AI large model. BACKGROUND
[0002] With the large-scale grid connection and centralized deployment of wind power, photovoltaic, energy storage and their supporting electrical equipment, new energy equipment gradually presents the engineering characteristics of high power density, strong coupling, multiple working conditions and long life cycle operation and maintenance. The field monitoring system has expanded from single sensing quantity to the cooperative collection of multi-source heterogeneous data such as vibration, acoustics, thermotics, electrical parameters, working conditions and operation and maintenance records. The fault diagnosis method has evolved from rule threshold and expert knowledge base to intelligent diagnosis system driven by multi-modal deep learning and pre-trained model. Although the multi-modal model can improve the recognition coverage of complex faults, it is still prone to problems such as unreliable alignment leading to overconfidence in diagnosis, undifferentiated risk level and diagnostic reliability, and overly rigid shutdown decision under conditions such as sampling asynchrony, timestamp drift, modal missing and noise disturbance in actual stations, thereby causing false positives and unnecessary shutdowns, affecting equipment availability and operation and maintenance input-output structure.
[0003] CN121071583A discloses a wind turbine starting condition intelligent diagnosis method and system based on multi-modal data fusion. Time sequence original data is constructed through multi-modal sensing data, a sound-vibration alignment mechanism is introduced, and a double-flow lightweight network is used to extract time sequence features. Cross-attention is used to complete time-frequency correlation alignment, and the wind turbine fault confidence is outputted. Whether the wind turbine is started is determined based on a pre-set confidence threshold. This scheme has reference value for lightweight diagnosis of wind turbine starting scene, but its decision main line still takes single confidence threshold driven as the core, lacks a judgment mechanism for structured linkage of alignment reliability, model uncertainty and fault consequence level, and is difficult to cover the engineering high-frequency gray area of high risk but insufficient evidence. The hierarchical distinction of shutdown and review strategy is still not sufficient.
[0004] CN121051429A discloses a key electrical equipment fault diagnosis method and system for energy storage power station based on multi-modal deep learning. A distributed data acquisition network is constructed to obtain multi-source heterogeneous data. Different sampling frequency data is mapped to a unified time axis through timestamp matching, interpolation and missing compensation. Further multi-modal feature extraction, fusion and model training are performed, and lightweight deployment is combined to adapt to edge side application. This scheme is helpful to improve the engineering usability of multi-modal data of energy storage power station, but its quantitative expression of alignment quality, diagnosis evidence packaging, chain structure of confidence and risk dual-domain decision is relatively limited, especially lacking a dynamic threshold updating idea coupled with false alarm cost and shutdown cost, which is difficult to form a traceable hierarchical decision-making closed loop between safety and economy.
[0005] In summary, the existing multi-modal new energy equipment diagnosis technology generally lacks explicit quantification of cross-modal alignment reliability, diagnosis reliability and fault risk are not formed in the dual domain coordination, and the threshold strategy is difficult to adapt to the change of the cost constraint and the working condition of the station. Therefore, the problem to be solved by the present application is how to construct a diagnosis reliability index and a fault risk index based on alignment confidence, model uncertainty and severity parameters under the engineering conditions of asynchronous and incomplete multi-source heterogeneous data, and form a reliability risk dual-domain grading matrix and a dynamic threshold accordingly, so as to give hierarchical decision of shutdown maintenance, supplementary sampling review, degradation observation and routine inspection. SUMMARY
[0006] This section is intended to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of the specification to avoid obscuring the purpose of this section, the abstract and the title. Such simplifications or omissions cannot be used to limit the scope of the present application.
[0007] In view of the above-mentioned existing problems, the present application is proposed.
[0008] To solve the above technical problems, the present application provides the following technical solutions: As a preferred scheme of the new energy equipment fault diagnosis method based on multi-modal data and AI large model, wherein: obtaining multi-modal monitoring data of the target new energy equipment within a preset diagnosis time window, and performing time alignment and quality labeling according to a unified time index to obtain multi-modal feature samples and alignment confidence corresponding to each modality; inputting the multi-modal feature samples into a pre-trained AI large model diagnosis network to obtain fault candidate labels and model uncertainty corresponding to the multi-modal feature samples, and forming a candidate fault evidence set in association with the alignment confidence; calculating a diagnosis reliability index based on the candidate fault evidence set, and calculating a fault risk index in combination with the severity parameter corresponding to the fault candidate label and the multi-modal monitoring data, constructing a reliability risk dual-domain grading matrix and generating a dynamic threshold; determining the reliability risk dual-domain grading matrix according to the dynamic threshold, and outputting a grading diagnosis decision: when the fault risk index is in the high risk area and the diagnosis reliability index is in the high reliability area, triggering a shutdown maintenance alarm and generating an operation instruction corresponding to the fault candidate label; when the fault risk index is in the high risk area and the diagnosis reliability index is in the low reliability area, triggering a supplementary sampling time window of the key modality and updating the multi-modal feature samples to obtain a review diagnosis result; trigger a degradation observation alarm and record the trend change of the multi-modal feature sample when the failure risk index is in a low risk zone and the diagnosis reliability index is in a high reliability zone; trigger a regular inspection and maintain the collection configuration of the next diagnosis time window when the failure risk index is in a low risk zone and the diagnosis reliability index is in a low reliability zone.
[0009] The present application has the following beneficial effects: the present application converts sampling asynchrony, missing and noise and other field disturbances into quantifiable alignment confidence by establishing a unified time index for multi-modal monitoring data within a preset diagnosis time window and performing time alignment and quality labeling; then inputs the multi-modal feature sample into a pre-trained AI large model diagnosis network to jointly output a failure candidate label and a model uncertainty, and encapsulates the two into a candidate failure evidence set together with the alignment confidence, so that the diagnosis result has a composite information basis of class evidence strength, model stability and alignment reliability; further generates a diagnosis reliability index based on the candidate failure evidence set, and couples a severity parameter corresponding to the failure candidate label with a multi-modal abnormality strength to generate a failure risk index, and brings the safety consequence and evidence sufficiency into the same judgment boundary through a reliability risk two-domain grading matrix and a dynamic threshold; finally, according to the dynamic threshold, executes a grading strategy, so that a high risk and high reliability scenario enters shutdown maintenance, a high risk and low reliability scenario enters key mode resampling and rechecking, a low risk and high reliability scenario enters degradation observation, and a low risk and low reliability scenario enters regular inspection, thereby solving the problems of false alarm diffusion caused by a fixed threshold in the prior art, over-disposal of insufficient evidence scenarios, and misjudgment of alignment errors as failure features, and achieving the technical effects of maintaining safety disposal priority while reducing unnecessary shutdown and false alarms, improving the traceability and field adaptability of diagnosis conclusions, and optimizing the cost-benefit ratio of operation and maintenance. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them: Figure 1 The flowchart of the new energy equipment failure diagnosis method based on multi-modal data and AI large model shown in the present application. DETAILED DESCRIPTION
[0011] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments.
[0012] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort should fall within the scope of protection of this invention.
[0013] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0014] According to an embodiment of the present invention, in combination Figure 1 The flowchart shown illustrates a fault diagnosis method for new energy equipment based on multimodal data and a large AI model, which specifically includes the following steps: S1. Obtain multimodal monitoring data of the target new energy equipment within a preset diagnostic time window, and perform time alignment and quality labeling according to a unified time index to obtain multimodal feature samples and alignment confidence scores corresponding to each modality. Note that the following should be noted in this step: In this embodiment, the target new energy equipment may be a wind turbine generator drive train, key electrical equipment of an energy storage power station, or other new energy equipment with multi-source sensing and operation and maintenance records.
[0015] In a preferred embodiment, the pre-training data of the pre-trained AI large-scale model diagnostic network includes at least vibration data, acoustic data, temperature and thermal imaging data, electrical parameter data, operating condition data, SCADA / control log data, image / video data, and operation and maintenance text data. The pre-training data is formed in the form of a unified time index sequence + multi-scale preset diagnostic time window. The basic time window length can be selected as 5 min, 15 min, and 60 min, and the sliding step size is set from 1 to 5 min according to the equipment category to cover multiple operating condition switching and multiple fault evolution stages. Time mapping, missing labeling, and initial quality labeling are performed on the multimodal data within the same time window, and then the data is used as pre-training sample units to be input into the pre-trained AI large-scale model diagnostic network.
[0016] In a preferred embodiment, the structure of the pre-trained AI large model diagnostic network is as follows: Modality coding layer: Modality adaptation substructure and modality coding substructure are set for different modes. For vibration / acoustic / electrical parameter modes, temporal convolution or temporal Transformer coding substructure is preferred. For temperature / thermal imaging modes, lightweight temporal attention coding substructure is preferred. For image / video modes, visual coding substructure is preferred. For operation and maintenance text modes, domain vocabulary embedding and sentence vector coding substructure is preferred. Shared semantic projection layer: Projects the adaptation features of each modality onto a shared semantic space of a unified dimension; Alignment confidence gated fusion layer: generate gating weights according to alignment confidence and weight fusion of modal representation vectors; Fault discrimination layer: output candidate log score vector corresponding to preset fault label set, and obtain candidate score through temperature calibration and normalization.
[0017] S1.1, determine the preset diagnosis time window corresponding to the target new energy equipment, and set a unified time index rule based on the preset diagnosis time window to obtain a unified time index sequence; As an example, the preset diagnosis time window adopts a fixed length + sliding update manner; when the wind turbine transmission chain is taken as the object, the preset diagnosis time window takes the latest 60 min monitoring segment and is updated with a sliding step of 5 min; when the energy storage converter is taken as the object, the preset diagnosis time window takes the latest 30 min monitoring segment and is updated with a sliding step of 2 min.
[0018] Illustratively, the unified time index rule is set based on the SCADA system main clock source, and the fixed sampling granularity Generate an equidistant index sequence; preferably , obtain a unified time index sequence , wherein is the starting time of the diagnosis time window, is the ending time of the diagnosis time window.
[0019] S1.2, collect multi-modal monitoring data of the target new energy equipment within the preset diagnosis time window, and record original time stamp and modal identifier for the multi-modal monitoring data to obtain multi-modal monitoring data with original time stamp; Wherein, the multi-modal monitoring data includes vibration data (such as three-axis acceleration sequence at the input end of the gearbox and the high-speed shaft), acoustic data (such as sound pressure level sequence of key frequency bands in the cabin), temperature and thermal imaging data (such as bearing seat temperature sequence and temperature rise characteristics of key regions of thermal image), electrical parameter data (such as voltage, current, harmonic content, power factor sequence), operating condition data (such as wind speed, speed, pitch angle, torque, load sequence), SCADA / control log data (such as control state switching events and alarm events), image / video data (such as blade surface inspection image frames), and operation and maintenance text data (such as work order description, maintenance notes, and fault phenomenon text). Write original time stamp and modal identifier for each record when collecting.
[0020] As an example, the vibration modal is sampled at 5 kHz; the acoustic modal is sampled at 16 kHz; the temperature modal is sampled at 1 Hz; the electrical parameter modal is sampled at 50-200 Hz; the image modal is sampled at 1-5 fps or is discretely collected according to the inspection task; the operation and maintenance text is associated with the time when the work order is generated and enters the diagnosis time window data pool.
[0021] S1.3. Perform time mapping and alignment on the multimodal monitoring data with original timestamps according to the unified time index sequence, and generate aligned data fragments according to the preset completion rules at the mapping conflict or missing indexes to obtain aligned multimodal data. It should be noted that the multimodal monitoring data with original timestamps are time-mapped and aligned with the unified time index sequence using the method of nearest neighbor index mapping + conflict resolution + missing data filling. For any modality, the original timestamp is mapped to the index time with the smallest distance in the unified time index sequence. If there are multiple records at the same index, the main record is retained in the order of priority of signal amplitude stability and data quality prediction, and the remaining records are merged into a conflict record set and associated with the corresponding index.
[0022] As an example, the preset completion rules are as follows: local linear interpolation is used for continuous numerical modes (vibration, acoustics, electrical parameters, temperature), and interval limits are set for the interpolation results; event / log modes are handled by event placeholders and the most recent state inheritance method; image / video modes are bound to the most recent frame index; and operation and maintenance text modes are bound to the window center index according to the text aggregation summary vector within the window; and thus, aligned multimodal data is obtained.
[0023] S1.4 Calculate the data quality score of each modality based on the missing proportion, noise amplitude and working condition consistency of the aligned multimodal data, and encapsulate the aligned multimodal data and data quality score into a multimodal sample unit; It should be noted that the missing percentage in this step is the index missing rate, which can be obtained from step S1.6. In this step, for any modality m, the data quality score is defined as: in, Score the data quality for the m-th mode; Weights for missing items; Weights for the noise term; Weights for consistency items; Let be the index missing rate of the m-th modality; This is the statistical measure of the noise amplitude of the m-th mode within the preset diagnostic time window; As a reference for noise amplitude normalization; This is the consistency index between the m-th mode and the operating condition data.
[0024] In this embodiment, the noise amplitude statistics can be obtained by weighting the high-frequency energy ratio, the spectral peak drift amplitude, or the time-domain peak coefficient; the operating condition consistency index is obtained based on the state clustering consistency of the operating condition variables in the same window; and then the aligned multimodal data and data quality scores are associated and encapsulated to obtain multimodal sample units.
[0025] S1.5. Generate multimodal feature vectors for multimodal sample units according to unified feature extraction rules, and combine the multimodal feature vectors with the data quality score to form multimodal feature samples; Furthermore, the unified feature extraction rules are set in the manner of intra-modal time-frequency / statistical features + cross-modal semantic alignment features.
[0026] For example, RMS, kurtosis, envelope spectrum peak value, and key meshing frequency band energy are extracted for vibration modes; Mel band energy and anomalous frequency band proportion are extracted for acoustic modes; temperature rise slope and hot spot area proportion are extracted for temperature / thermal imaging modes; harmonic distortion rate and load fluctuation amplitude are extracted for electrical parameter modes; texture and edge statistical vectors of defect candidate regions are extracted for image / video modes; and text embedding vectors aligned with domain vocabulary are extracted for operation and maintenance text modes. These are used to generate multimodal feature vectors, which, together with data quality scores, form multimodal feature samples.
[0027] S1.6. Based on the unified time index sequence, statistically align the number of valid indexes and missing indexes of each modality in the multimodal data, and calculate the index missing rate corresponding to each modality; For example, let the number of effective indices of the m-th mode on the unified time index sequence be... The number of missing indexes is The total number of indexes is ,but: in, Let $\frac{m}{m}$ be the index missing rate for the $m$-th modality. This represents the number of indices for which this modality is not mapped to a valid data segment within the current diagnostic time window; To unify the number of indexes in the time index sequence within this diagnostic time window.
[0028] S1.7 Perform deviation statistics on the mapping results between the original timestamps and the unified time index sequence of each mode, and calculate the mapping deviation value corresponding to each mode; As an example, let the original timestamp of the i-th record in the m-th modality be... The mapped unified time index corresponds to the time of the time. The number of records is ,but: in, This is the mapping deviation value; This represents the number of records for the m-th modality that participated in the mapping statistics within the diagnostic time window.
[0029] S1.8 Extract cross-modal synchronization reference features based on aligned multimodal data, and calculate the synchronization consistency index between each mode and the synchronization reference features to obtain the cross-modal synchronization consistency index corresponding to each mode. It should be noted that the synchronization reference feature is a combination of operating condition transition events and energy mutation features. For example, power command step, speed mutation, and SOC slope mutation can be selected as alignment anchor points. Synchronization response features of vibration, acoustic, electrical parameters and thermal modes are extracted in the neighborhood of the anchor points to obtain cross-modal synchronization consistency index.
[0030] For example, let the synchronization response vector of the m-th mode on the anchor point set A be... The synchronous reference feature vector is ,but: in, It serves as a cross-modal synchronization consistency index; It is a norm 2; To prevent constants with a denominator of zero.
[0031] S1.9. Based on the preset weighted fusion rules, the index missing rate, mapping deviation value and cross-modal synchronization consistency index are jointly scored, and the alignment confidence corresponding to each modality is generated according to the preset normalization rules.
[0032] In a preferred embodiment, the index missing rate is determined according to a preset weighted fusion rule. Mapping deviation value Cross-modal synchronization consistency index Perform joint scoring and generate alignment confidence scores corresponding to each modality according to preset normalization rules. ; For example, first construct an intermediate quantity of alignment quality. : in, , , To merge weights and satisfy ; This is the mapping deviation attenuation coefficient; It is an exponential mapping function; Then generate alignment confidence scores according to the unified normalization rules: in, Let m be the alignment confidence of the m-th modality; and For each modality within the current diagnostic time window The minimum and maximum values.
[0033] For example, when the electrical parameter modes of a certain energy storage power station have high consistency with the operating condition anchor point response, low missing rate, and small mapping deviation, then... Approaching 1; when the video mode of a certain wind turbine blade appearance experiences intermittent frame loss during periods of communication congestion, then... Enlarge The decrease provides a suppression signal for subsequent gating fusion.
[0034] S2. Input the multimodal feature samples into the pre-trained AI large-scale model diagnostic network to obtain the fault candidate labels and model uncertainties corresponding to the multimodal feature samples, and associate and align the confidence scores to form a candidate fault evidence set. Note that the following should be noted in this step: In a preferred embodiment, the pre-training data of the pre-trained AI large-scale model diagnostic network consists of historical monitoring data of new energy equipment from multiple sites, seasons, and operating conditions, including at least eight modalities: vibration, acoustics, temperature / thermal imaging, electrical parameters, operating conditions, SCADA / control logs, images / videos, and maintenance text. The pre-training data is composed of samples according to the diagnostic time window and unified time indexing rules consistent with S1, and adopts a hybrid annotation method of time window-level labels and event-level labels. Time window level labels correspond to typical fault families or health states; Event-level labels correspond to control and protection actions, alarm codes, and key component replacement events.
[0035] This implementation allows large models to simultaneously acquire cross-modal semantic alignment capabilities and temporal structure priors of fault evolution during the pre-training phase.
[0036] S2.1 Read the pre-trained AI large model diagnostic network corresponding to the target new energy equipment, and construct the multimodal feature samples into a multimodal input tensor X according to the network input format; It should be noted that the pre-trained AI large model diagnostic network is a hierarchical structure consisting of a multimodal encoder, a shared semantic projection, an aligned confidence gating fusion, and a fault discrimination layer. Its multimodal encoder includes: a temporal convolutional or temporal Transformer encoding substructure for vibration / acoustic / electrical parameters; a temporal lightweight attention encoding substructure for temperature / thermal imaging; a visual encoding substructure for images / videos; and a domain vocabulary embedding and sentence vector encoding substructure for operational text.
[0037] S2.2 Input the multimodal input tensor X into each modality encoding layer of the pre-trained AI large-scale model diagnostic network to obtain the modality representation vector corresponding to each modality. ; It should be noted that, based on the modality identifier of the multimodal input tensor, modality input sub-tensors corresponding to each modality are obtained by decomposing the tensor. ;Will Input the modality adaptation layer corresponding to this modality. Obtain adaptation features ;Will Input shared semantic projection layer Initial modal representation vectors are generated according to a unified dimension mapping rule. ;right Normalization and temporal aggregation are performed to obtain modal representation vectors. .
[0038] For example, its mathematical expression is as follows: in, For the mmm modal adaptation layer; For shared semantic projection layer; For normalization operators; It is a temporal aggregation operator, which can take either attention pooling or multi-scale statistical pooling.
[0039] S2.3. Based on the alignment confidence, the modal representation vectors are subjected to confidence-gated weighted fusion to obtain the fused diagnostic representation vector; Specifically, read the alignment confidence corresponding to each modality. The alignment confidence is then converted into the initial values of the modal gating weights according to a preset interval mapping rule. In this embodiment, a piecewise linear mapping is used: when hour, Select the low-weight interval; when hour, Take the middle weighted interval; when hour, Select the high-weighted interval; Then on Normalization is performed to obtain normalized gating weights. and modal characterization vector Modal weighting is used to form a fusion diagnostic representation vector. .
[0040] For example, its mathematical expression is as follows: Where M is the number of modes; is the normalized gating weight; H is the fusion diagnostic representation vector.
[0041] S2.4 Input the fused diagnostic representation vector into the fault discrimination layer of the pre-trained AI large model diagnostic network, and output the fault candidate labels and their candidate scores corresponding to the multimodal feature samples. In a preferred embodiment, a preset fault label set L is read, and the fused diagnostic representation vector H is input into the fault discrimination layer. Output the candidate logarithmic score vector s that corresponds one-to-one with L; perform temperature calibration and normalization on s to obtain the candidate score p, and then select the top K labels as fault candidate labels according to the sorting results.
[0042] For example, its mathematical expression is as follows: in, For temperature parameters; The candidate log score for label i; The candidate score for label i; when At that time, parallel candidate labels such as early gearbox wear, abnormal bearing lubrication, and loose sensor can be obtained, along with their corresponding candidate scores.
[0043] S2.5 Calculate the model uncertainty based on the distribution dispersion of candidate scores, and encapsulate the fault candidate labels, model uncertainty and alignment confidence into a candidate fault evidence set.
[0044] For example, the normalized entropy of the candidate scores is used as the source of uncertainty. Specifically, each evidence record in the candidate fault evidence set includes at least: candidate label identifier, candidate score, corresponding modal gating weight, aligned confidence statistic and uncertainty value, so as to perform multi-factor confidence synthesis in S3.
[0045] S3. Calculate the diagnostic confidence index based on the candidate fault evidence set, and calculate the fault risk index by combining the severity parameters corresponding to the fault candidate labels and multimodal monitoring data. Construct a confidence-risk dual-domain hierarchical matrix and generate dynamic thresholds. Note that the following should be noted in this step: In a preferred embodiment, the sources of severity parameters include at least: equipment manufacturing specifications and protection setting recommendations; historical failure consequence classification (such as downtime, probability of secondary damage, and grid connection impact level); operation and maintenance cost model (such as spare parts cost, labor and window period cost); and a one-to-one mapping is established between severity parameters and preset fault label sets to form a severity mapping table for risk synthesis.
[0046] S3.1 Aggregate the candidate scores corresponding to each candidate fault label in the candidate fault evidence set to obtain the aggregated candidate score value; Specifically, the candidate scores corresponding to each candidate fault label in the candidate fault evidence set are read, and the candidate scores of the same candidate fault label in multiple consecutive preset diagnostic time windows are aggregated in chronological order to obtain the aggregated candidate score value. The aggregation weight can be jointly determined by the time mean of alignment confidence and the time mean of data quality score, so that the time window with reliable alignment and higher ontology quality contributes more to the aggregated value. Finally, the aggregated candidate score value set corresponding one-to-one with each candidate fault label is output.
[0047] For example, in the wind turbine generator drivetrain scenario, the candidate scores are aggregated using a weighted moving average of the most recent three to five diagnostic time windows. When the wind speed jumps from the medium wind range to the high wind range, the aggregation mechanism suppresses the instantaneous fluctuation of the candidate scores, making the aggregated candidate score value corresponding to the main shaft bearing abnormality more consistent with the real trend of continuous temperature rise and load fluctuation.
[0048] S3.2 Read the model uncertainty and alignment confidence corresponding to the candidate fault evidence set and normalize them to obtain the uncertainty normalization value and alignment confidence normalization value. Specifically, the model uncertainty corresponding to the candidate fault evidence set is read and normalized to obtain the uncertainty normalization value. The alignment confidence corresponding to each mode is aggregated at the evidence layer to obtain the alignment confidence aggregate value. The aggregation method is preferably consistent with the normalization mechanism of the S2 gating weights to ensure that the gating fusion contribution and the evidence layer alignment reliability have a consistent statistical vector. Finally, the uncertainty normalization value and the alignment confidence aggregate value are output to S3.3 as input for confidence fusion.
[0049] For example, in the wind turbine generator drivetrain scenario, when the confidence level of vibration and acoustic alignment decreases while the confidence level of temperature, electrical parameters and operating conditions alignment remains stable, the aggregated value of alignment confidence is not excessively dragged down by a single low-alignment mode; in the energy storage converter scenario, when the confidence level of thermal imaging mode and temperature mode alignment remains high, the aggregated value of alignment confidence and the aggregated value of candidate scores show a positive reinforcing relationship, improving the stability of subsequent candidate confidence values.
[0050] S3.3 According to the preset multi-factor joint credibility fusion rule, the candidate score aggregate value, uncertainty normalized value and alignment confidence normalized value are weighted and synthesized to obtain the candidate credibility value. Specifically, the candidate score aggregation value, uncertainty normalization value, and alignment confidence aggregation value are read, and the three are weighted and synthesized according to preset weights to obtain the candidate confidence value. The candidate score aggregation value and alignment confidence aggregation value are used as positive factors, and the uncertainty normalization value is used as a negative factor. Weights and constraints are set, and the candidate confidence value is then output to S3.4 to generate the diagnostic confidence index.
[0051] For example, in the wind turbine generator drivetrain scenario, when the candidate score aggregation value increases with the continuous temperature rise and the energy increment of the characteristic frequency band, and the alignment confidence aggregation value remains in the medium-high range and the uncertainty normalization value is low, the candidate confidence value enters the high range; in the energy storage converter scenario, when thermal imaging anomalies and harmonic distortion jointly push up the candidate score aggregation value, and the gated fusion shows that the thermal imaging mode weight ratio is stable, the candidate confidence value forms a higher evidence support strength for the thermal anomalies of the power module.
[0052] S3.4. Perform interval mapping on the candidate confidence values according to the preset confidence mapping rules to generate a diagnostic confidence index corresponding to the candidate fault evidence set. As an example, the candidate confidence value is obtained using the following formula: Then, a diagnostic confidence index is generated according to the preset confidence mapping rules: in, This represents the candidate confidence value corresponding to the k-th fault candidate label; Aggregate the candidate scores; To align the aggregated confidence values; To normalize the uncertainty; As the score weight; To align weights; Uncertainty weight; The initial boundary is set as a high-confidence threshold; The initial boundary is set at a low confidence threshold; This is the discrete mapping result for the diagnostic credibility index.
[0053] In this embodiment, the preset confidence mapping rule example is a three-segment interval mapping, so as to match the discrete units of the subsequent dual-domain hierarchical matrix.
[0054] S3.5 Read the severity parameters corresponding to the fault candidate labels and establish a one-to-one severity mapping table between the severity parameters and the fault candidate labels; Specifically, the severity parameter in this embodiment is determined by the safety level of the consequences of failure of critical components in the equipment manufacturing specifications, the statistics of historical downtime losses, and the operation and maintenance resource consumption model.
[0055] For example, a higher severity range can be set for tags such as DC bus abnormality, power module thermal imbalance, and insulation degradation in the energy storage converter system; and a medium to low severity range can be set for tags such as sensor loosening and short-term communication abnormality.
[0056] S3.6. Extract risk feature quantities corresponding to fault candidate labels based on multimodal monitoring data, and normalize them to obtain normalized risk feature values. Specifically, based on multimodal monitoring data and fault candidate labels, a subset of modes with high correlation to the candidate labels is selected from vibration, acoustics, temperature / thermal imaging, electrical parameters, and operating conditions. Within a preset diagnostic time window, risk feature quantities corresponding to the candidate labels are extracted. The risk feature quantities include a joint vector of at least two types of indicators, such as frequency band energy increment, temperature rise slope, hot spot area ratio, harmonic distortion rate change amplitude, and load fluctuation amplitude. The risk feature quantities are normalized according to the historical stable operating condition baseline of the same equipment and the statistical baseline of the same equipment group to obtain the normalized risk feature value, and the normalized risk feature value is output to S3.7.
[0057] S3.7. According to the preset severity feature joint scoring rules, the severity parameter and the normalized value of the risk feature are weighted and synthesized to obtain the risk value corresponding to the fault candidate label. Specifically, the severity parameter and risk feature normalization value corresponding to the fault candidate label are read, and the two are weighted and synthesized according to the preset weight to obtain the risk value corresponding to the fault candidate label. The severity parameter has a higher weight than the risk feature weight so that high-consequence faults can also enter the higher risk concern range in the medium abnormal intensity stage. The risk value is output to S3.8 to generate the fault risk index.
[0058] For example, in the wind turbine generator drivetrain scenario, if the severity parameter corresponding to the main shaft bearing abnormality is in the high range and the normalized value of the risk feature shows a continuous upward trend, then the risk value enters the high range; in the energy storage converter scenario, if the severity parameter corresponding to the power module thermal abnormality is in the high range and the hot spot area growth rate continuously increases across the window, then the risk value quickly enters the medium-high range, providing a stable risk input for subsequent dual-domain matrix determination.
[0059] S3.8. Map the risk values to intervals according to the preset risk interval mapping rules to generate a fault risk index corresponding to the multimodal monitoring data.
[0060] As an example, the risk value can be obtained using the following formula: And generate a fault risk index according to the preset risk range mapping rules: in, This is the risk value corresponding to the k-th fault candidate label; This is the severity parameter corresponding to the k-th fault candidate label; This is the normalized value of the risk feature corresponding to the k-th fault candidate label; Severity weighting; For feature weights; This serves as the initial boundary for the high-risk threshold. The initial boundary is set as the low-risk threshold; This is the failure risk index.
[0061] Specifically, a dual-domain hierarchical matrix of reliability and risk is constructed based on the diagnostic reliability index and the failure risk index, and dynamic thresholds are generated, including: Confidence grading rules and risk grading rules are set for the diagnostic confidence index and the failure risk index, respectively, to obtain the initial values of the confidence level boundary and the initial values of the risk level boundary. A credibility-risk dual-domain hierarchical matrix is constructed based on the initial values of the credibility level boundary and the risk level boundary, and the diagnostic credibility index and the failure risk index are mapped to the matrix units of the credibility-risk dual-domain hierarchical matrix. Within the rolling statistical time window, summarize the diagnostic records corresponding to the dual-domain grading matrix of credibility risk, and calculate the false alarm rate statistics and downtime statistics. Based on the preset dual-objective threshold correction rules, the initial values of the credibility level boundary and the risk level boundary are corrected by combining the false alarm rate statistics and the downtime statistics, thereby generating a dynamic threshold corresponding to the credibility-risk dual-domain classification matrix.
[0062] As an example, the diagnostic confidence index is divided into low-confidence, medium-confidence, and high-confidence zones, and the fault risk index is divided into low-risk, medium-risk, and high-risk zones, thus constructing... Credibility risk dual-domain hierarchical matrix : Where C is the set of credibility levels; R is the set of risk levels; These are matrix units corresponding to the level combinations; Within the rolling statistical time window, summarize the diagnostic records corresponding to the matrix cells and calculate the false alarm rate and downtime rate statistics: in, This is the false alarm rate statistic; Number of false alarms; Total number of alarms; This is a statistical value for downtime rate; To calculate the cumulative downtime within the window; To calculate the cumulative running time within the statistics window; It is a very small positive number.
[0063] For example, the dynamic threshold update can be obtained by the following formula: in, Let be the confidence level boundary parameter for the t-th iteration; Let be the risk level boundary parameter for the t-th iteration; This is the false alarm rate statistic; This is a statistical value for downtime rate; This represents the distribution drift. , , This is the target reference value; , , , The step size coefficient is used to correct the threshold.
[0064] In a preferred embodiment, the dynamic threshold correction period can be one of 24 hours, 7 days, or 30 days, and can be configured according to the device type and maintenance plan.
[0065] In a preferred embodiment, upper and lower limits are set for the dynamic threshold boundary: the update range of the confidence boundary parameter is limited to a preset safety range to avoid short-term noise causing drastic drift of the level boundary; the risk boundary parameter sets a smaller downward adjustment space for high-consequence labels to retain a safety margin.
[0066] S4. Determine the credibility risk dual-domain classification matrix based on the dynamic threshold, and output the classification diagnosis decision: When the fault risk index is in the high-risk zone and the diagnostic confidence index is in the high-confidence zone, a shutdown maintenance alarm is triggered and an operation command corresponding to the fault candidate label is generated. When the fault risk index is in the high-risk zone and the diagnostic confidence index is in the low-confidence zone, the key modality resampling time window is triggered and the multimodal feature samples are updated to obtain the verification diagnostic results. When the fault risk index is in the low-risk zone and the diagnostic confidence index is in the high-confidence zone, a downgrade observation alarm is triggered and the trend changes of multimodal feature samples are recorded. When the fault risk index is in the low-risk zone and the diagnostic confidence index is in the low-confidence zone, a routine inspection is triggered and the data collection configuration for the next diagnostic time window is maintained.
[0067] For example, for critical electrical equipment in an energy storage power station, when the candidate label is power module thermal imbalance and Located in a high-risk area When in the high-confidence zone, the operation instructions include a set of instructions such as reducing the power command ramp limit, switching redundant branches, starting forced cooling in the cabin, and arranging a window period to replace the target module. The execution receipt index corresponding to the set of instructions is written to the control log.
[0068] For example, when the confidence level of electrical parameters aligned with vibration modes is significantly higher than that of acoustic modes in the diagnostics of wind turbine drivetrain, and the overall... When the confidence level is too low, the 1-2 key modes with the highest alignment confidence can be specified as the priority targets for resampling. The resampling time window is set to 10 minutes, and the resampling data is recalculated according to S1. Update the gating weights and review the candidate label sorting.
[0069] For example, in the early stage of blade thermal anomalies, if the thermal imaging is highly consistent with the operating conditions but the anomaly intensity is still low, this situation can be mapped to a low-risk-high-confidence combination, generating a downgraded observation instruction, recording the time series of hotspot location, area and temperature rise rate, and incorporating this trend series into the threshold drift calculation for the next statistical period.
[0070] For example, when multiple modalities are temporarily out of sync and the candidate scores have high dispersion, only the candidate evidence and the alignment confidence profile are recorded, without triggering additional sampling burden, thus avoiding the occupation of operation and maintenance resources due to false triggering of low confidence.
[0071] In a preferred embodiment, the risk interval is divided using a dynamic threshold. Based on the corresponding level boundaries, the confidence interval is divided using dynamic thresholds. The corresponding level boundary shall be used as the standard; when the false alarm rate F increases within the rolling statistical time window, the high confidence zone boundary may be appropriately adjusted upward; when the downtime rate D increases abnormally and is not accompanied by a concentrated appearance of high severity labels, the high risk zone boundary may be appropriately adjusted upward in order to balance security and availability.
[0072] In a preferred embodiment, key equipment related to the PCS / DC combiner and battery compartment thermal management of the energy storage power station is selected as the target new energy equipment. Traditional solutions based on a single temperature threshold or a single current ripple threshold are prone to frequent switching of alarm boundaries when facing high-temperature seasons, different SOC ranges, and different power scheduling strategies, resulting in prominent false alarms and conservative shutdowns. When using the method of this invention, thermal imaging, electrical parameters, operating conditions, and log modes are formed into aligned multimodal data under a unified time index. Alignment confidence drives the gating weights, suppressing weak quality modes caused by communication jitter during fusion. The candidate fault evidence set further explicitly writes the model uncertainty into the confidence synthesis. The severity parameter is combined with the operation and maintenance cost model to participate in risk synthesis.
[0073] For example, when the power module thermal imbalance candidate label maintains a high candidate score over two consecutive diagnostic time windows and the thermal imaging hotspot expansion rate increases, Migration to high-risk areas; if electrical parameters and thermal modes All are in the high range. If it is also in the high-confidence zone, the hierarchical matrix outputs a shutdown and maintenance command; if If the device is in the low confidence zone, it will first trigger a resampling and verification of thermal imaging and electrical parameters to avoid excessive downtime caused by single noise or alignment loss.
[0074] In a preferred embodiment, wind turbine gearbox / main shaft bearing and blade thermal anomalies are selected as example target scenarios. Traditional single-mode vibration diagnosis is prone to spectral drift and misjudgment during low wind speeds, frequent pitch changes, or control strategy switching phases. When blade thermal anomalies rely solely on infrared thresholds, it is difficult to distinguish between ambient temperature differences and actual defect hotspots. When using the method of this invention, vibration, acoustics, thermal imaging, operating conditions, and SCADA modes jointly participate in the aligned confidence score, and the gating fusion stage is based on... The modal characterization is weighted; when drastic changes in wind conditions lead to an increase in acoustic noise, the mode is weighted. and The sensitivity of the fusion diagnostic characterization to acoustic disturbances is reduced by downsampling.
[0075] For example, when the candidate label is early gearbox wear but there is no synchronous anomaly in thermal / electrical parameters, the normalized value of the risk characteristic is too low, making... It should remain in the low to medium range; if cross-modal synchronous responses such as torque jump, sudden increase in vibration meshing frequency band energy, and increase in oil temperature rise rate occur within the same time window, then rise, rise, The tiered matrix outputs a downgraded observation or window period maintenance instead of immediate downtime, thereby enhancing the controllability of the operation and maintenance strategy.
[0076] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for fault diagnosis of new energy equipment based on multimodal data and AI large-scale models, characterized in that, include: Acquire multimodal monitoring data of the target new energy equipment within a preset diagnostic time window, and perform time alignment and quality labeling according to a unified time index to obtain multimodal feature samples and alignment confidence corresponding to each mode; The multimodal feature samples are input into a pre-trained AI large model diagnostic network to obtain the fault candidate labels and model uncertainties corresponding to the multimodal feature samples, and the alignment confidence is associated to form a set of candidate fault evidence. The diagnostic confidence index is calculated based on the candidate fault evidence set, and the fault risk index is calculated by combining the severity parameter corresponding to the fault candidate label with the multimodal monitoring data. A confidence risk dual-domain hierarchical matrix is constructed and a dynamic threshold is generated. The credibility risk dual-domain classification matrix is judged based on the dynamic threshold, and a classification diagnosis decision is output: When the fault risk index is in the high-risk zone and the diagnostic confidence index is in the high-confidence zone, a shutdown maintenance alarm is triggered and an operation command corresponding to the fault candidate label is generated. When the fault risk index is in the high-risk zone and the diagnostic confidence index is in the low-confidence zone, the key modality's resampling time window is triggered and the multimodal feature samples are updated to obtain the verification diagnostic results. When the fault risk index is in the low-risk zone and the diagnostic confidence index is in the high-confidence zone, a downgrade observation alarm is triggered and the trend changes of the multimodal feature samples are recorded. When the fault risk index is in the low-risk zone and the diagnostic confidence index is in the low-confidence zone, a routine inspection is triggered and the data collection configuration for the next diagnostic time window is maintained.
2. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 1, characterized in that, Obtaining the multimodal feature samples includes: A preset diagnostic time window corresponding to the target new energy equipment is determined, and a unified time index rule is set based on the preset diagnostic time window to obtain a unified time index sequence; Multimodal monitoring data of the target new energy equipment is collected within the preset diagnostic time window, and the original timestamp and modality identifier are recorded for the multimodal monitoring data to obtain multimodal monitoring data with original timestamp; The unified time index sequence is used to perform time mapping and alignment on the multimodal monitoring data with original timestamps, and aligned data fragments are generated at mapping conflicts or missing indices according to preset completion rules to obtain aligned multimodal data. The quality score of each modality data is calculated based on the missing proportion, noise amplitude and working condition consistency of the aligned multimodal data, and the aligned multimodal data and the data quality score are associated and encapsulated into a multimodal sample unit; A multimodal feature vector is generated for the multimodal sample unit according to a unified feature extraction rule, and the multimodal feature vector and the data quality score are combined to form the multimodal feature sample; The multimodal monitoring data includes vibration data, acoustic data, temperature and thermal imaging data, electrical parameter data, operating condition data, SCADA / control log data, image / video data, and maintenance text data.
3. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 2, characterized in that, The alignment confidence corresponding to each modality is obtained, including: Based on the unified time index sequence, the number of valid indexes and missing indexes for each modality in the aligned multimodal data are statistically analyzed, and the index missing rate corresponding to each modality is calculated. The mapping results between the original timestamps of each modality and the unified time index sequence are statistically analyzed to calculate the mapping deviation value corresponding to each modality. Based on the aligned multimodal data, cross-modal synchronization reference features are extracted, and the synchronization consistency index between each modality and the synchronization reference features is calculated to obtain the cross-modal synchronization consistency index corresponding to each modality. The index missing rate, the mapping deviation value and the cross-modal synchronization consistency index are jointly scored according to the preset weight fusion rules, and the alignment confidence scores corresponding to each modality are generated according to the preset normalization rules.
4. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 3, characterized in that, The set of candidate fault evidence includes: Read the pre-trained AI large model diagnostic network corresponding to the target new energy equipment, and construct the multimodal feature samples into a multimodal input tensor according to the network input format; The multimodal input tensor is input into each modality encoding layer of the pre-trained AI large model diagnostic network to obtain the modality representation vector corresponding to each modality. The modal representation vector is subjected to confidence-gated weighted fusion based on the alignment confidence to obtain a fused diagnostic representation vector; The fused diagnostic representation vector is input into the fault discrimination layer of the pre-trained AI large model diagnostic network, and the fault candidate labels and their candidate scores corresponding to the multimodal feature samples are output. The model uncertainty is calculated based on the distribution dispersion of the candidate scores, and the fault candidate labels, the model uncertainty, and the alignment confidence are associated and encapsulated into a candidate fault evidence set.
5. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 4, characterized in that, The process of obtaining the modality representation vector corresponding to each modality includes: Based on the modality identifier of the multimodal input tensor, modality input sub-tensors corresponding to each modality are obtained; The modal input subtensors are respectively input into the modal adaptation layer in the pre-trained AI large model diagnostic network to obtain the adaptation features corresponding to each modality; The adaptation features are input into the shared semantic projection layer in the pre-trained AI large model diagnostic network, and initial modality representation vectors corresponding to each modality are generated according to a unified dimension mapping rule. The initial modal representation vector is normalized and time-series aggregated to obtain the modal representation vector corresponding to each mode.
6. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 5, characterized in that, The fusion diagnostic representation vector is obtained by: Read the alignment confidence corresponding to each modality, and convert the alignment confidence into the initial value of the gating weight of each modality according to the preset interval mapping rule; The initial values of the gating weights for each mode are normalized to obtain the normalized gating weights corresponding to each mode. The modal representation vector is weighted modally according to the normalized gating weights to obtain each modal weighted representation vector; The weighted representation vectors of each modality are fused at the vector level to obtain the fused diagnostic representation vector.
7. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 6, characterized in that, Output the candidate fault labels and their candidate scores, including: Read the preset fault label set corresponding to the target new energy equipment, and perform dimensional matching between the fused diagnostic representation vector and the preset fault label set; The fused diagnostic representation vector is input into the fault discrimination layer of the pre-trained AI large model diagnostic network, and a candidate log score vector corresponding one-to-one with the preset fault label set is output. The candidate logarithmic score vector is subjected to temperature calibration and normalization to obtain the candidate score corresponding to the preset fault label set; The top K labels are selected as the fault candidate labels based on the ranking of the candidate scores, and the candidate scores corresponding to the fault candidate labels are recorded.
8. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 4, characterized in that, Calculating the diagnostic confidence index includes: The candidate scores corresponding to each candidate fault label in the candidate fault evidence set are aggregated to obtain the aggregated candidate score value. Read the model uncertainty and the alignment confidence corresponding to the candidate fault evidence set and normalize them to obtain the uncertainty normalization value and the alignment confidence normalization value; According to the preset multi-factor joint credibility fusion rule, the candidate score aggregate value, the uncertainty normalized value and the alignment confidence normalized value are weighted and synthesized to obtain the candidate credibility value; The candidate confidence values are mapped to intervals according to a preset confidence mapping rule to generate a diagnostic confidence index corresponding to the candidate fault evidence set.
9. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 8, characterized in that, The calculation of the failure risk index includes: Read the severity parameter corresponding to the fault candidate label, and establish a one-to-one severity mapping table between the severity parameter and the fault candidate label; Based on the multimodal monitoring data, risk feature quantities corresponding to the fault candidate labels are extracted and normalized to obtain risk feature normalized values. According to the preset severity feature joint scoring rules, the severity parameter and the normalized value of the risk feature are weighted and synthesized to obtain the risk value corresponding to the fault candidate label; The risk values are mapped according to a preset risk interval mapping rule to generate a fault risk index corresponding to the multimodal monitoring data.
10. The method for fault diagnosis of new energy equipment based on multimodal data and AI large model according to claim 9, characterized in that, Based on the diagnostic reliability index and the failure risk index, a reliability-risk dual-domain hierarchical matrix is constructed and a dynamic threshold is generated, including: Confidence grading rules and risk grading rules are set for the diagnostic confidence index and the fault risk index, respectively, to obtain the initial values of the confidence level boundary and the initial values of the risk level boundary. A credibility-risk dual-domain hierarchical matrix is constructed based on the initial values of the credibility level boundary and the initial values of the risk level boundary, and the diagnostic credibility index and the failure risk index are mapped to the matrix units of the credibility-risk dual-domain hierarchical matrix. Within the rolling statistical time window, summarize the diagnostic records corresponding to the credibility risk dual-domain classification matrix, and calculate the false alarm rate statistics and downtime statistics. Based on the preset dual-objective threshold correction rule, the initial values of the confidence level boundary and the initial values of the risk level boundary are corrected by combining the false alarm rate statistics and the downtime statistics, thereby generating a dynamic threshold corresponding to the confidence risk dual-domain classification matrix.
Citation Information
Patent Citations
Energy storage power station key electrical equipment fault diagnosis method and system based on multi-mode deep learning
CN121051429A
Fan starting condition intelligent diagnosis method and system based on multi-modal data fusion
CN121071583A