Wind Farm Multi-Source Heterogeneous Data Processing Method and Device

By integrating GAN-LSTM and RNN-LSTM neural networks to process multi-source heterogeneous data in wind farms, the problem of low intelligent intelligence in the existing technology is solved, and efficient fault diagnosis and prediction is achieved.

CN115391523BActive Publication Date: 2025-07-22STATE GRID HUBEI ELECTRIC POWER RES INST +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210934927.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-05
Publication Date
2025-07-22
Estimated Expiration
2042-08-05

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process multi-source heterogeneous data in wind farms, especially unstructured data and images, resulting in low intelligence in fault diagnosis and poor model adaptability, requiring a lot of manual intervention and retraining.

Method used

The fusion method of GAN-LSTM and RNN-LSTM neural network is adopted to extract and label multi-source heterogeneous data, generate training samples, train and fuse network models, and realize intelligent diagnosis of the main electrical equipment of the wind farm.

Benefits of technology

It improves the accuracy and intelligence of fault diagnosis of main electrical equipment of wind farms, reduces manual intervention, and enhances the adaptability of the model and data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391523B_ABST
    Figure CN115391523B_ABST
Patent Text Reader

Abstract

A method and device for processing multi-source heterogeneous data in a wind farm. The method includes: obtaining multi-source heterogeneous data containing the operation information of the main electrical equipment in the wind farm; extracting a set of operation information of the main electrical equipment in the wind farm from the multi-source heterogeneous data; annotating the set of operation information of the main electrical equipment in the wind farm to generate a training sample set; training a GAN-LSTM network and an RNN-LSTM network through the training sample set and fusing the GAN-LSTM network and the RNN-LSTM network; inputting the collected real-time operation information of the main electrical equipment in the wind farm into the GAN-LSTM network, the RNN-LSTM network and the fused network to obtain a diagnosis result; and determining the diagnosis of the main electrical equipment in the wind farm according to the diagnosis result. The present invention can accurately extract a set of operation information of the main electrical equipment in the wind farm from multi-source heterogeneous data, and improve the accuracy and intelligence of diagnosis through a unique neural network model and training method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart grid, and specifically to a method and device for processing multi-source heterogeneous data in a wind farm. Background Art

[0002] Processing multi-source heterogeneous data in a wind farm and extracting features are of great significance for effectively diagnosing faults of main electrical equipment in the wind farm and accurately determining the operating state of the equipment. With the development of information technology, intelligent technologies have gradually been applied to the field of fault diagnosis. The most commonly used technologies are supervised learning methods such as wavelet networks, support vector machines, fuzzy clustering, grey clustering, rough sets, and Bayesian network classifiers. Supervised learning can make full use of existing knowledge and improve classification accuracy by repeatedly selecting and measuring samples, but it is affected by human subjectivity and is not suitable for unknown categories.

[0003] With the continuous popularization and deepening of smart grids, the amount of relevant data of main electrical equipment in wind farms has increased explosively, the data types have gradually become diversified, and the data timeliness has been continuously improved. Different platforms are not unified, including structured data and unstructured data such as images and videos. For structured data, there are many sources, involving hundreds of attributes, including physical, business, and operation aspects. With the continuous growth of monitoring equipment and information platforms, the data sources will continue to expand. In addition, these multi-source heterogeneous data are usually very noisy and may have missing values.

[0004] Currently, the method for processing wind farm data is to use supervised learning methods to analyze structured data, and the selection and evaluation of training samples will require more people and time. For the fault diagnosis of unstructured data and images, it stays at the stage of obtaining results through manual analysis and has a low intelligence level. In addition, when adding data sources, the existing models are no longer applicable, and professional data analysts need to redesign the models and retrain the parameters. Summary of the Invention

[0005] To solve the above problems existing in the prior art, the present invention provides a method and device for processing multi-source heterogeneous data in a wind farm that focuses on unstructured data and images.

[0006] A method for processing multi-source heterogeneous data in a wind farm includes the following steps:

[0007] Obtain multi-source heterogeneous data containing the operating information of main electrical equipment in the wind farm;

[0008] Extract a set of operating information of main electrical equipment in the wind farm from the multi-source heterogeneous data;

[0009] Annotate the set of operating information of main electrical equipment in the wind farm to generate a set of training samples;

[0010] Train the GAN-LSTM network and the RNN-LSTM network respectively using the training sample set, and fuse the GAN-LSTM network and the RNN-LSTM network to obtain a fused network;

[0011] Input the collected real-time operation information of the main electrical equipment of the wind farm into the GAN-LSTM network, the RNN-LSTM network and the fused network respectively to obtain a diagnosis result;

[0012] Determine the diagnosis of the main electrical equipment of the wind farm according to the diagnosis result.

[0013] Further, the method for extracting the operation information set of the main electrical equipment of the wind farm from the multi-source heterogeneous data adopts a multi-source heterogeneous data extraction method, which specifically includes:

[0014] Step 2.1: Define the data structure of the operation information of the main electrical equipment of the wind farm. The data structure consists of information elements and specific element attributes of the information elements. The information elements include location information elements, type information elements and time information elements. The location information element represents the location of the main electrical equipment of the wind farm, the type information element represents the events that occur to the main electrical equipment of the wind farm, and the time information element represents the start and end times of the event;

[0015] Step 2.1: Use the words that play a key role in describing the operation information of the main electrical equipment of the wind farm as feature words. According to the grammatical roles of these words in the multi-source heterogeneous data, define the feature word types for filling the element attributes of the operation information of the main electrical equipment of the wind farm, and construct a professional word library according to the feature word types;

[0016] Step 2.3: Based on the data structure of the operation information of the main electrical equipment of the wind farm defined in Step 2.1 and the feature word types defined in Step 2.2, combined with the syntactic structure features and syntactic structure features of the events described by the multi-source heterogeneous data for the main electrical equipment of the wind farm, formulate a basic extraction pattern, and expand the basic extraction pattern through rules to obtain an extraction pattern library;

[0017] Step 2.4: Use the collected multi-source heterogeneous data as the input text, and preprocess the input text to obtain the vocabulary sequence of the input text;

[0018] Step 2.5: Use the professional word library in Step 2.2 to identify the feature words that appear in the vocabulary sequence obtained in Step 2.4, and record the types of the feature words in the order of their appearance in the input text to generate the feature word type sequence of the input text, and filter the input text by judging whether the feature word types required for the element attributes of the operation information of the main electrical equipment of the wind farm are complete;

[0019] Step 2.6: Segment the input text into sentences. Based on the set of sentences obtained from the segmentation, split the sequence of feature word types of the input text obtained in Step 2.5 into a set of sequences of feature word types corresponding to the set of sentences. Use the dynamic time warping (DTW) distance to measure the similarity between each sequence of feature word types in the set of sequences of feature word types and the sequence of feature word types of each extraction pattern in the extraction pattern library. Select the extraction pattern with the highest similarity and less than the given threshold as the matching extraction pattern for this sentence.

[0020] Step 2.7: Traverse the set of sentences in the input text. If a sentence in the set of sentences obtains a matching extraction pattern in Step 2.6, fill the feature words in this sentence into the corresponding element attributes of the wind farm main electrical equipment operation information according to the element attribute sequence of this matching extraction pattern, generate the wind farm main electrical equipment operation information corresponding to this sentence, and obtain the set of wind farm main electrical equipment operation information with the extracted location information elements and type information elements of the input text.

[0021] Step 2.8: According to the different expression forms of time in the multi-source heterogeneous data, formulate a set of regular expressions for extracting the numerical values of time elements such as year, month, day, hour, minute, and second. Combine the judgment rules and use this set of regular expressions to extract the numerical values of time elements from the input text, and combine these numerical values of time elements into the element attributes of the event start time and the element attributes of the event end time to obtain the time information elements of the wind farm main electrical equipment operation information.

[0022] Step 2.9: Fill the time information elements extracted in Step 2.8 into the set of wind farm main electrical equipment operation information obtained in Step 2.7 to obtain a complete set of wind farm main electrical equipment operation information with the wind farm main electrical equipment operation information elements.

[0023] Furthermore, the extraction pattern includes two parts: a sequence of feature word types and an element attribute sequence. The sequence of feature word types is the order of the types of feature words used to describe events in the multi-source heterogeneous data. The function of the sequence of feature word types in the extraction pattern is to determine whether the multi-source heterogeneous data can match this extraction pattern. The element attribute sequence has the same length as the sequence of feature word types. The sequence items in the element attribute sequence are the element attributes corresponding to the sequence items in the same position in the sequence of feature word types in the wind farm main electrical equipment operation information. The function of the element attribute sequence is to map the feature words that appear in the multi-source heterogeneous data to the corresponding element attributes in the wind farm main electrical equipment operation information.

[0024] Furthermore, the preprocessing in Step 2.4 includes deleting duplicate information in the input text and performing Chinese word segmentation on the input text.

[0025] Further, after the traversal in step 2.7 is completed, it is judged whether the attributes of the positioning information elements and the attributes of the type information elements of the obtained operation information of the main electrical equipment in the wind farm are complete. If they are not complete, the supplement rules are used to fill in the attributes of the positioning information elements or the attributes of the type information elements missing in the operation information of the main electrical equipment in the wind farm.

[0026] Further, the collected real-time operation information of the main electrical equipment in the wind farm is respectively input into the GAN-LSTM network, the RNN-LSTM network and the fused network to obtain the first diagnosis result, the second diagnosis result and the third diagnosis result; the diagnosis of the main electrical equipment in the wind farm is determined according to the diagnosis results, specifically including:

[0027] 1) If the first diagnosis result, the second diagnosis result and the third diagnosis result are exactly the same, the faulty equipment and its location are determined according to any one of the diagnosis results;

[0028] 2) If the first diagnosis result, the second diagnosis result and the third diagnosis result are not exactly the same, the faulty equipment and its location are determined according to the relationship of the locations of the faulty equipment in various diagnosis results;

[0029] 3) If the first diagnosis result, the second diagnosis result and the third diagnosis result are completely different, the step of obtaining multi-source heterogeneous data containing the operation information of the main electrical equipment in the wind farm is returned for execution.

[0030] A wind farm multi-source heterogeneous data processing device includes:

[0031] A multi-source heterogeneous data acquisition module, configured to acquire multi-source heterogeneous data containing the operation information of the main electrical equipment in the wind farm;

[0032] An information extraction module, configured to extract a set of operation information of the main electrical equipment in the wind farm from the multi-source heterogeneous data;

[0033] A training sample generation module, configured to label the set of operation information of the main electrical equipment in the wind farm to generate a set of training samples;

[0034] A network training and fusion module, configured to train the GAN-LSTM network and the RNN-LSTM network respectively through the set of training samples, and fuse the GAN-LSTM network and the RNN-LSTM network to obtain a fused network;

[0035] A diagnosis module, configured to respectively input the collected real-time operation information of the main electrical equipment in the wind farm into the GAN-LSTM network, the RNN-LSTM network and the fused network to obtain a diagnosis result respectively, and determine the diagnosis of the main electrical equipment in the wind farm according to the diagnosis result.

[0036] Further, the information extraction module extracts the operation information set of the main electrical equipment of the wind farm from the multi-source heterogeneous data, specifically including:

[0037] Step 2.1: Define the data structure of the operation information of the main electrical equipment of the wind farm. The data structure consists of information elements and specific element attributes of the information elements. The information elements include location information elements, type information elements, and time information elements. The location information element represents the location of the main electrical equipment of the wind farm. The type information element represents the events that occur to the main electrical equipment of the wind farm. The time information element represents the start and end times of the event.

[0038] Step 2.2: Use the words that play a key role in describing the operation information of the main electrical equipment of the wind farm as feature words. According to the grammatical roles of these words in the multi-source heterogeneous data, define the feature word types for filling the element attributes of the operation information of the main electrical equipment of the wind farm, and construct a professional word library according to the feature word types.

[0039] Step 2.3: Based on the data structure of the operation information of the main electrical equipment of the wind farm defined in Step 2.1 and the feature word types defined in Step 2.2, combined with the syntactic structure features and syntactic structure features of the events that occur to the main electrical equipment of the wind farm described in the multi-source heterogeneous data, formulate a basic extraction pattern, and expand the basic extraction pattern through rules to obtain an extraction pattern library.

[0040] Step 2.4: Take the collected multi-source heterogeneous data as the input text, and preprocess the input text to obtain the vocabulary sequence of the input text.

[0041] Step 2.5: Use the professional word library in Step 2.2 to identify the feature words that appear in the vocabulary sequence obtained in Step 2.4, record the types of the feature words in the order of their appearance in the input text, generate the feature word type sequence of the input text, and filter the input text by judging whether the types of the feature words required for the element attributes of the operation information of the main electrical equipment of the wind farm are complete.

[0042] Step 2.6: Segment the input text into sentences. According to the sentence set obtained by segmentation, divide the feature word type sequence of the input text obtained in Step 2.5 into a set of feature word type sequences corresponding to the sentence set. Use the dynamic time warping (DTW) distance to measure the similarity between each feature word type sequence in the set of feature word type sequences and the feature word type sequences of each extraction pattern in the extraction pattern library, and select the extraction pattern with the highest similarity and less than the given threshold as the matching extraction pattern for the sentence.

[0043] Step 2.7: Traverse the sentence set of the input text. If a sentence in the sentence set obtains a matching extraction pattern in Step 2.6, fill the feature words in this sentence into the corresponding wind farm main electrical equipment operation information element attributes according to the element attribute sequence of the matching extraction pattern, generate the wind farm main electrical equipment operation information corresponding to this sentence, and obtain the wind farm main electrical equipment operation information set with the extracted positioning information elements and type information elements of the input text;

[0044] Step 2.8: According to the different expression forms of time in the multi-source heterogeneous data, formulate a set of regular expressions for extracting the numerical values of time elements such as year, month, day, hour, minute, and second. Combine the judgment rules and use this set of regular expressions to extract the time element numerical values from the input text, and combine these time element numerical values into the event start time element attribute and the event end time element attribute to obtain the time information element of the wind farm main electrical equipment operation information;

[0045] Step 2.9: Fill the time information element extracted in Step 2.8 into the wind farm main electrical equipment operation information set obtained in Step 2.7 to obtain the wind farm main electrical equipment operation information set with complete wind farm main electrical equipment operation information elements.

[0046] Further, after the traversal in Step 2.7 is completed, judge whether the attributes of the positioning information elements and the attributes of the type information elements of the obtained wind farm main electrical equipment operation information are complete. If they are not complete, use the supplementary rules to fill the attributes of the positioning information elements or the attributes of the type information elements missing in the wind farm main electrical equipment operation information.

[0047] Further, the diagnosis module inputs the collected real-time wind farm main electrical equipment operation information into the GAN-LSTM network, the RNN-LSTM network, and the fused network respectively, and obtains the first diagnosis result, the second diagnosis result, and the third diagnosis result respectively. Determining the diagnosis of the wind farm main electrical equipment according to the diagnosis results specifically includes:

[0048] 1) If the first diagnosis result, the second diagnosis result, and the third diagnosis result are exactly the same, determine the faulty equipment and its location according to any one of the diagnosis results;

[0049] 2) If the first diagnosis result, the second diagnosis result, and the third diagnosis result are not exactly the same, determine the faulty equipment and its location according to the relationship of the locations of the faulty equipment in various diagnosis results;

[0050] 3) If the first diagnosis result, the second diagnosis result, and the third diagnosis result are completely different, then return to execute the step of the multi-source heterogeneous data acquisition module to acquire the multi-source heterogeneous data containing the wind farm main electrical equipment operation information. A wind farm multi-source heterogeneous data processing system includes: a computer-readable storage medium and a processor;

[0051] The computer-readable storage medium is used to store executable instructions;

[0052] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the multi-source heterogeneous data processing method for a wind farm.

[0053] A non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-source heterogeneous data processing method for a wind farm is implemented.

[0054] The beneficial effects of the present invention are as follows:

[0055] In view of the multi-source heterogeneous data including unstructured data and images generated during the operation of the main electrical equipment in a wind farm, the present invention extracts the operation information of the main electrical equipment in the wind farm from the multi-source heterogeneous data, annotates the operation information and trains a neural network, and then conducts fault diagnosis and prediction on the main electrical equipment in the wind farm. Through the multi-source heterogeneous data extraction method, the present invention can accurately extract the operation information set of the main electrical equipment in the wind farm from the multi-source heterogeneous data; through the unique neural network model and training method, the accuracy and intelligence of diagnosis can be improved. Description of the Drawings

[0056] Figure 1 is the accuracy rate trend chart of fault recognition under different numbers of LSTM units;

[0057] Figure 2 is the ROC curve comparison and the percentages of FN / FP / TN / TP under different numbers of LSTM units;

[0058] Figure 3 is the accuracy rate trend chart of fault recognition under different activation units;

[0059] Figure 4 is the ROC curve comparison chart under different activation units;

[0060] Figure 5 is the percentages of FN / FP / TN / TP under different activation units

[0061] Figure 6 is the accuracy rate trend chart of fault recognition under different Batch sizes;

[0062] Figure 7 is the ROC curve comparison and the percentages of FN / FP / TN / TP under different Batch sizes. Detailed Embodiments

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] The present invention is directed to multi-source heterogeneous data containing unstructured data and images generated during the operation of main electrical equipment in a wind farm. The operation information of the main electrical equipment in the wind farm is extracted from the multi-source heterogeneous data, the operation information is labeled, and a neural network is trained, and then fault diagnosis and prediction are performed on the main electrical equipment in the wind farm.

[0065] An embodiment of the present invention provides a method for processing multi-source heterogeneous data in a wind farm, including the following steps:

[0066] Step 1: Obtain multi-source heterogeneous data containing the operation information of the main electrical equipment in the wind farm;

[0067] Step 2: Adopt a multi-source heterogeneous data extraction method to extract the operation information set of the main electrical equipment in the wind farm from the multi-source heterogeneous data.

[0068] The multi-source heterogeneous data extraction method mainly targets the unstructured data in the multi-source heterogeneous data, and specifically includes the following steps:

[0069] Step 2.1: Define the data structure of the operation information of the main electrical equipment in the wind farm to facilitate organizing and managing the operation information of the main electrical equipment in the wind farm in the form of a two-dimensional table. The data structure consists of information elements and specific element attributes of the information elements. The information elements include location information elements, type information elements, and time information elements. The location information element represents the location of the main electrical equipment in the wind farm, the type information element represents the events that occur to the main electrical equipment in the wind farm, and the time information element represents the start and end times of the event.

[0070] Step 2.2: Use the words that play a key role in describing the operation information of the main electrical equipment in the wind farm as feature words. According to the grammatical roles played by these words in the multi-source heterogeneous data, define the feature word types for filling the element attributes of the operation information of the main electrical equipment in the wind farm, and construct a professional word library according to the feature word types.

[0071] Step 2.3: Based on the data structure of the operating information of the main electrical equipment in the wind farm defined in Step 2.1 and the types of feature words defined in Step 2.2, combined with the syntactic structure features and morphological structure features of the events occurring in the main electrical equipment in the wind farm described in the multi-source heterogeneous data, formulate a basic extraction pattern, and expand the basic extraction pattern through rules to obtain an extraction pattern library. The extraction pattern includes two parts: a sequence of feature word types and a sequence of element attributes; the sequence of feature word types is the order of the types of feature words used when describing events in the multi-source heterogeneous data. The function of the sequence of feature word types in the extraction pattern is to determine whether the multi-source heterogeneous data can match this extraction pattern; the sequence of element attributes has the same length as the sequence of feature word types. The sequence items in the sequence of element attributes are the element attributes corresponding to the sequence items in the same position in the sequence of feature word types in the operating information of the main electrical equipment in the wind farm. The function of the sequence of element attributes is to map the feature words that appear in the multi-source heterogeneous data to the corresponding element attributes in the operating information of the main electrical equipment in the wind farm.

[0072] Step 2.4: Use the collected multi-source heterogeneous data as the input text, and preprocess the input text. This preprocessing includes deleting duplicate information in the input text and performing Chinese word segmentation on the input text to obtain the vocabulary sequence of the input text.

[0073] Step 2.5: Use the professional vocabulary in Step 2.2 to identify the feature words that appear in the vocabulary sequence obtained in Step 2.4, and record the types of feature words in the order of their appearance in the input text to generate the sequence of feature word types of the input text. Filter the input text by judging whether the types of feature words required for the element attributes of the operating information of the main electrical equipment in the wind farm are complete.

[0074] Step 2.6: Segment the input text into sentences. According to the set of sentences obtained from the segmentation, divide the sequence of feature word types of the input text obtained in Step 2.5 into a set of sequences of feature word types corresponding to the set of sentences. Use the dynamic time warping (DTW) distance to measure the similarity between each sequence of feature word types in the set of sequences of feature word types and the sequence of feature word types of each extraction pattern in the extraction pattern library, and select the extraction pattern with the highest similarity and less than the given threshold as the matching extraction pattern for this sentence.

[0075] Step 2.7: Traverse the sentence set of the input text. If a sentence in the sentence set obtains a matching extraction pattern in Step 2.6, fill the feature words in this sentence into the corresponding wind farm main electrical equipment operation information element attributes according to the element attribute sequence of the matching extraction pattern, and generate the wind farm main electrical equipment operation information corresponding to this sentence. After the traversal, judge whether the attributes of the positioning information element and the attributes of the type information element of the obtained wind farm main electrical equipment operation information are complete. If they are not complete, use the supplementation rules to fill the attributes of the positioning information element or the attributes of the type information element missing in the wind farm main electrical equipment operation information. Finally, obtain the wind farm main electrical equipment operation information set with the positioning information element and the type information element extracted from the input text.

[0076] Step 2.8: According to the different expression forms of time in the multi-source heterogeneous data, formulate a set of regular expressions for extracting the numerical values of time elements such as year, month, day, hour, minute, and second. Combine the judgment rules and use this set of regular expressions to extract the time element numerical values from the input text, and combine these time element numerical values into the event start time element attribute and the event end time element attribute to obtain the time information element of the wind farm main electrical equipment operation information.

[0077] Step 2.9: Fill the time information element extracted in Step 2.8 into the wind farm main electrical equipment operation information set obtained in Step 2.7 to obtain the wind farm main electrical equipment operation information set with complete wind farm main electrical equipment operation information elements.

[0078] Through the above 9 steps, the wind farm main electrical equipment operation information set can be extracted from the multi-source heterogeneous data.

[0079] Step 3: Label the wind farm main electrical equipment operation information set to generate a training sample set.

[0080] Step 4: Train the GAN-LSTM network and the RNN-LSTM network respectively through the training sample set, and fuse the GAN-LSTM network and the RNN-LSTM network to obtain the fused network.

[0081] In the past, speech text processing was usually a combination of neural networks and hidden Markov models. Using algorithms and computer hardware, acoustic models established through deep forward propagation networks have made quite remarkable progress in recent years. Considering sound, text processing is an internal dynamic process, and generative adversarial networks can be used as one of its candidate models. Dynamic means that the current processed text vector is associated with the context content. It cannot be an independent analysis of the current sample, but should set a comprehensive analysis of semantic information before and after the storage unit of the text information. This method applies a larger data state space and richer model dynamic performance.

[0082] The Generative Adversarial Network (GAN) is an important generative model in the field of deep learning. That is, two networks (the generator and the discriminator) are trained simultaneously and compete in the minimax algorithm. This adversarial approach avoids some difficulties of traditional generative models in practical applications and cleverly approximates some unsolvable loss functions through adversarial learning. It has a wide range of applications in the generation of data such as images, videos, natural language, and music.

[0083] Although the Recurrent Neural Network (RNN) performs the transformation from sentences to vectors in a principled way, it is usually difficult to learn long-term dependencies within a sequence due to the vanishing gradient problem. The recurrent neural network has two limitations: First, text analysis is actually context-dependent, while the RNN only has access to previous text, not subsequent text; Second, compared with the time step, the RNN has more difficulties in learning temporal correlations. The Bidirectional Long Short-Term Memory (BLSTM) network can be used for the first problem, while the Long Short-Term Memory model is for the second problem. The RNN repeating module only contains one neuron.

[0084] The LSTM model is an improvement of the traditional RNN model. Based on the RNN model, it adds a unit control mechanism to solve the long-term dependence problem of the RNN and the gradient explosion problem caused by long sequences. This model enables the RNN model to remember long-term information by designing special structural units. And by designing three "gate" structures: the forget gate layer, the input gate layer, and the output gate layer. When the control information passes through the unit, information can be selectively added and removed through the unit structure.

[0085] LSTM controls the transmission of information through gates, which are usually represented by the Sigmoid function. The key to LSTM is the cell state, whose horizontal line runs across the top of the figure. The cell state persists throughout the neural chain to transmit information, with only some small linear interactions. LSTM has the ability to remove or add information to the cell state, which is regulated by structures called gates. Gates are an optional way to let information pass through, and they consist of a Sigmoid neural network layer and a dot multiplication operation.

[0086] The role of the forget gate layer is to determine whether the input information from the upper layer is discarded or not, and it is used to control the hidden layer nodes at the last moment of the historical information stored. The forget gate calculates a value between 0 and 1 based on the state of the hidden layer at the previous time and the input of the current time node, and acts on the state of the cell at the previous moment to determine what information needs to be retained and discarded. "1" means "fully retain", while "0" means "completely delete the information". The output of the hidden layer unit (historical information) can be selectively processed through the handling of the forget gate.

[0087] The input gate layer is used to control the input of the unit state of the hidden gate layer. It can input information through multiple operations to determine whether to update the input information to the current state, so as to determine the information that needs to be updated and the information that needs to be stored. First, the input gate layer is established, and the Sigmoid function is used to determine which information should be updated. The output of the input gate layer is a value between 0 and 1 of the Sigmoid output, and then it acts on the input information to determine whether to update the corresponding value of the unit state. Among them, 1 means allowing information to pass through, and the corresponding value needs to be updated, and 0 means not allowing the corresponding value not to be updated. It can be seen that the input gate layer can remove some unnecessary information, and then a layer can be established by adding the candidate state of the neuron phase phasor, and the two are jointly calculated to update the value. The main purpose of updating the state of the neuron is to update the neuron state C t-1 at the previous moment to the state C t at the next moment. Multiply the state at the previous moment by f t and sum it with it×Ct, and remove the information that was considered negligible before to obtain C t . C t is the new candidate value, which depends on the number of times each state value is updated. In the case of a language model, this is actually deleting the information at the previous moment and adding a new information state, such as the decision in the previous step.

[0088] The output gate layer is used to control the output of the current hidden layer node and determine whether to output to the next hidden layer or the output layer. By controlling the output, it can be determined which information needs to be output. The value of its state is "0" or "1". "1" means that output is required, and "0" means that output is not required. After the final output value, the output control information about the current unit state can be found.

[0089] In step 4, the GAN-LSTM network is trained through the training sample set, and the specific steps are as follows:

[0090] For the operation information set of the main electrical equipment of the wind farm obtained by the multi-source heterogeneous data extraction processing method, each operation information of the main electrical equipment of the wind farm is labeled, and the operation information and its labeled information are used as training samples to train the GAN-LSTM neural network.

[0091] Specifically, the operation information with faults and its labels are used as training samples to train LSTM and GAN. Specifically: input the operation information and labeled information in the first training sample into LSTM to train LSTM, and obtain the final LSTM deep network. Based on the final LSTM, predict the future trend of the first training sample to obtain a prediction result, which is used as the second training sample set (including operation information and its labeled information). Input the second training sample set into the GAN network to train the generator and discriminator in the GAN network to obtain the final generator.

[0092] The process of training the LSTM specifically includes: dividing the first training sample into a training set and a validation set; inputting the training set into the LSTM to train the LSTM; inputting the validation set into the LSTM and calculating the relative error of the trained LSTM; if the relative error meets the preset condition, obtaining the final LSTM; if the relative error does not meet the preset condition, re-dividing the first training sample into a training set and a validation set and continuing the training.

[0093] The specific process of training the GAN includes: inputting random noise into the generator to obtain a noise data set, where the noise data set is of the same type as the second training sample; the step of training the discriminator: inputting the data set output by the generator and the second training sample into the discriminator to train the discriminator so that the resolution of the discriminator for the second training sample reaches the first threshold; the step of training the generator: inputting the second training sample into the generator to train the generator so that the similarity between the training data set generated by the generator and the second training sample is greater than the second threshold; repeatedly executing the step of training the discriminator and the step of training the generator until the resolution of the discriminator for the second training sample reaches the third threshold (greater than the first threshold), and taking the generator obtained by training at this time as the final generator.

[0094] In the past, speech text processing was usually a combination of neural networks and hidden Markov models. In recent years, acoustic models established through deep forward propagation networks have made considerable progress using algorithms and computer hardware. Considering sound, text processing is an internal dynamic process, and recurrent neural networks can be used as one of its candidate models. Dynamic means that the currently processed text vector is associated with the context content. It cannot be an independent analysis of the current sample, but should set a comprehensive analysis of semantic information before and after the storage unit of the text information. This method applies a larger data state space and richer model dynamic performance.

[0095] In a neural network, each neuron is a processing unit that takes the outputs of the nodes connected to it as inputs. Before emitting an output, each neuron first applies a non-linear activation function. It is because of this activation function that the neural network has the ability to model non-linear relationships. However, general neural models cannot explicitly simulate temporal relationships. The assumption that all data points are composed of vectors of a fixed length results in a significant reduction in the processing effect of the model when there is a strong correlation in the input phasors. Therefore, recurrent neural networks (RNNs) are introduced to endow the neural network with the ability to explicitly model time by adding self-connected hidden layers that span time points; the feedback of the hidden layer not only enters the output end but also enters the hidden layer of the next time step.

[0096] Traditional neural networks do not have a recurrent process in the middle layer. When specifying inputs \(x_0, x_1, x_2,\cdots, x\) t at that time, there will be some corresponding outputs \(h_0, h_1, h_2,\cdots, h\) t after the process of the neurons. In each training, there is no need for information transmission between neurons. The difference between a recurrent neural network and a traditional neural network is that for each training of an RNN, information needs to be transmitted between neurons. In this training, neurons need to use the state information after the action of the last neuron, similar to a recursive function.

[0097] Among them, the training method for the RNN-LSTM network can adopt the existing training methods, which will not be elaborated here.

[0098] In step 4, the GAN-LSTM network and the RNN-LSTM network are fused to obtain a fused network, specifically as follows:

[0099] Assume that the output channels of the GAN-LSTM network are \(X_1, X_2,\cdots, X\) c , and the output channels of the RNN-LSTM network are \(Y_1, Y_2,\cdots, Y\) c . After fusing the GAN-LSTM network and the RNN-LSTM network, the output channels are where \(K\) is the fusion coefficient.

[0100] The fused network model includes an input layer, a non-linear transformation layer, a linear fusion layer, and an output layer. The input layer includes two first branches with the same network structure. Each first branch includes a convolutional layer and a rectified linear unit. The non-linear transformation layer includes two second branches with the same network structure and respectively connected to the corresponding first branches. Each second branch includes 5 layers of networks, and each layer of network includes a convolutional layer, a batch normalization, and a ReLU activation function. The linear fusion layer fuses the results of the two second branches of the non-linear transformation layer to obtain an output result. The output layer includes a global average pooling layer, a randomly dropped neuron connection, and a fully connected layer. The output result of the linear fusion layer is output to the global average pooling layer.

[0101] In step 5, the collected real-time operation information of the main electrical equipment in the wind farm is respectively input into the GAN-LSTM network, the RNN-LSTM network, and the fused network to obtain a first diagnostic result, a second diagnostic result, and a third diagnostic result respectively. The diagnostic results include the status of each main electrical equipment, the event occurrence time, and the number.

[0102] In step 6, compare the first diagnostic result, the second diagnostic result, and the third diagnostic result to finally determine the diagnosis of the main electrical equipment in the wind farm.

[0103] 1) If the first diagnostic result, the second diagnostic result, and the third diagnostic result are exactly the same, determine the faulty device and its location according to any one of the diagnostic results.

[0104] 2) If the first diagnostic result, the second diagnostic result, and the third diagnostic result are not exactly the same, determine the faulty device and its location according to the relationship of the locations of the faulty devices in various diagnostic results.

[0105] Example 1: The first diagnostic result includes faulty device 1 (location 1) and faulty device 2 (location 2), the second diagnostic result includes faulty device 3 (location 3) and faulty device 4 (location 4), and the third diagnostic result includes faulty device 5 (location 5), faulty device 6 (location 6), and faulty device 7 (location 7). It can be seen that the third diagnostic result is different from both the first diagnostic result and the second diagnostic result. If faulty device 1 is the same as faulty device 3, location 1 is the same as location 3, faulty device 2 is the same as faulty device 4, and location 2 is the same as location 4, then determine the faulty device and its location according to the first diagnostic result.

[0106] Example 2: The first diagnostic result includes faulty device 1 (location 1) and faulty device 2 (location 2), the second diagnostic result includes faulty device 3 (location 3) and faulty device 4 (location 4), and the third diagnostic result includes faulty device 5 (location 5) and faulty device 6 (location 6). If faulty device 1 is the same as faulty device 3 and faulty device 5, and location 1 is the same as location 3 and location 5, then first determine that there is a fault in faulty device 1 on the branch where location 1 is located. If location 2, location 4, and location 6 are the same, but faulty device 2, faulty device 4, and faulty device 6 are different, then determine that there is a fault on the branch where location 1, location 3, and location 5 are located, and it is impossible to determine which specific device has a fault.

[0107] After that, obtain the historical operation parameters of faulty device 2, faulty device 4, and faulty device 6, and compare the current operation parameters with the historical operation parameters to determine the device most likely to have a fault. Among them, the historical operation parameters can be data under the same conditions, such as similar or the same time, similar or the same weather, etc.

[0108] 3) If the first diagnostic result, the second diagnostic result, and the third diagnostic result are completely different, re - execute the steps of this method, that is, return to execute the step of obtaining multi - source heterogeneous data containing the operation information of the main electrical equipment in the wind farm.

[0109] The embodiment of the present invention also provides a wind farm multi - source heterogeneous data processing device, including:

[0110] A multi - source heterogeneous data acquisition module, configured to acquire multi - source heterogeneous data containing the operation information of the main electrical equipment in the wind farm;

[0111] An information extraction module for extracting a set of operating information of the main electrical equipment of the wind farm from the multi-source heterogeneous data;

[0112] A training sample generation module for annotating the set of operating information of the main electrical equipment of the wind farm to generate a set of training samples;

[0113] A network training and fusion module for training the GAN-LSTM network and the RNN-LSTM network respectively through the set of training samples, and fusing the GAN-LSTM network and the RNN-LSTM network to obtain a fused network;

[0114] A diagnosis module for inputting the collected real-time operating information of the main electrical equipment of the wind farm into the GAN-LSTM network, the RNN-LSTM network and the fused network respectively to obtain a diagnosis result, and determining the diagnosis of the main electrical equipment of the wind farm according to the diagnosis result.

[0115] Among them, the information extraction module extracts a set of operating information of the main electrical equipment of the wind farm from the multi-source heterogeneous data, specifically including:

[0116] Step 2.1: Define the data structure of the operating information of the main electrical equipment of the wind farm. The data structure consists of information elements and specific element attributes of the information elements. The information elements include location information elements, type information elements and time information elements. The location information element represents the location of the main electrical equipment of the wind farm, the type information element represents the events occurring to the main electrical equipment of the wind farm, and the time information element represents the start and end times of the event;

[0117] Step 2.1: Use the words that play a key role in the process of describing the operating information of the main electrical equipment of the wind farm as feature words. According to the grammatical roles of these words in the multi-source heterogeneous data, define the types of feature words for filling the element attributes of the operating information of the main electrical equipment of the wind farm, and construct a professional word library according to the types of feature words;

[0118] Step 2.3: Based on the data structure of the operating information of the main electrical equipment of the wind farm defined in Step 2.1 and the types of feature words defined in Step 2.2, combined with the syntactic structure features and syntactic structure features of the events occurring to the main electrical equipment of the wind farm described in the multi-source heterogeneous data, formulate a basic extraction pattern, and expand the basic extraction pattern through rules to obtain an extraction pattern library;

[0119] Step 2.4: Use the collected multi-source heterogeneous data as the input text, and preprocess the input text to obtain the vocabulary sequence of the input text;

[0120] Step 2.5: Identify the feature words that appear in the vocabulary sequence obtained in Step 2.4 using the professional vocabulary in Step 2.2, record the types of feature words in the order of their appearance in the input text, generate the feature word type sequence of the input text, and filter the input text by determining whether the types of feature words required for the element attributes of the main electrical equipment operation information in the wind farm are complete;

[0121] Step 2.6: Segment the input text, and according to the set of sentences obtained from the segmentation, divide the feature word type sequence of the input text obtained in Step 2.5 into a set of feature word type sequences corresponding to the set of sentences. Use the dynamic time warping (DTW) distance to measure the similarity between each feature word type sequence in this set of feature word type sequences and the feature word type sequences of each extraction pattern in the extraction pattern library, and select the extraction pattern with the highest similarity and less than the given threshold as the matching extraction pattern for this sentence;

[0122] Step 2.7: Traverse the set of sentences in the input text. If the sentences in the set of sentences obtain a matching extraction pattern in Step 2.6, fill the feature words in this sentence into the corresponding main electrical equipment operation information element attributes of the wind farm according to the element attribute sequence of this matching extraction pattern, generate the main electrical equipment operation information of the wind farm corresponding to this sentence, and obtain the set of main electrical equipment operation information of the wind farm with the location information elements and type information elements already extracted;

[0123] Step 2.8: According to the different expression forms of time in the multi-source heterogeneous data, formulate a set of regular expressions for extracting the numerical values of time elements such as year, month, day, hour, minute, and second. Combine the judgment rules to extract the time element numerical values from the input text using this set of regular expressions, and combine these time element numerical values into the event start time element attribute and the event end time element attribute to obtain the time information element of the main electrical equipment operation information in the wind farm;

[0124] Step 2.9: Fill the time information element extracted in Step 2.8 into the set of main electrical equipment operation information of the wind farm obtained in Step 2.7 to obtain the set of main electrical equipment operation information of the wind farm with complete main electrical equipment operation information elements.

[0125] Based on the simulation and analysis of the power grid fault inspection report:

[0126] The complete neural network training method for all model parameters has been described in detail in the previous text. Below, a fault detection report of a certain wind farm will be used as the analysis object. Through the processing of the above network model, machine learning can be used to classify and analyze unstructured data in different situations. Based on the training of the network model with a large number of single fault samples, a test set is imported for the accuracy test of fault types. In the embodiment of the present invention, three variables are selected and the fault recognition rates are compared in the fault report. When the other two variables are fixed, different moving times are used to verify the test samples. These three variables are: the number of LSTM units, the type of activation unit, and Batch size. Batch size is the size of the data processed in each batch and is the depth of the dedicated training method for learning. It can not only reduce the number of weight adjustments, prevent overfitting, but also speed up the training.

[0127] 1. Analysis of multi-source heterogeneous data sets

[0128] (1) Fault type analysis

[0129] The description of the corpus is as follows. The fault inspection report records the income of grid personnel during daily maintenance by inspecting grid equipment, lines, and protection devices. The accumulation of statements one by one constitutes the main body of the report. Among them, the information in the fault inspection report is mainly composed of 6 main information bodies such as "DeviceInfo", "TripInfo", "Faultinfo", "DigitalStatus", "DigitalEvent", "SettingValue" and several common information. The TripInfo information body can contain multiple optional FaultInfo information. The FaultInfo information body indicates the action current and voltage and can clearly reflect and display the fault conditions and operation processes through the report. The content source of the DeviceInfo information can be a fixed value or a configuration file. The information of Faultinfo, DigitalStatus, DigitalEvent, and SettingValue can vary according to the protection type or manufacturer. Faultinfo can be used as auxiliary information for a single action message or fault parameters for the entire action group. The content of each information body is as follows:

[0130] 1) DeviceInfo: Describes the information part of the recording device.

[0131] 2) TripInfo: Records some protection action events during the fault process.

[0132] 3) FaultInfo: Records information such as fault current, fault voltage, fault phase, and fault distance during the fault recording process.

[0133] 4) DigialStatus: Record the signal in front of the device into the self-check signal status.

[0134] 5) DigitalEvent: Record the changes of events such as self-check signals during the fault protection process; all switches are sorted according to the action time, and the action time and return time are recorded simultaneously.

[0135] 6) SettingValue: Record the actual value set by the device during the fault.

[0136] According to the dynamic fault recording of the power system, in the present invention, all faults are divided into the following five categories, and corresponding labels are given after each record: mechanical faults, electrical faults, secondary equipment faults, faults caused by the external environment, and faults caused by human factors.

[0137] In the embodiment of the present invention, the fault inspection reports in the recent 10 years are selected as the data set. In the used data set, the specific types of fault causes and fault reasons, and their statistical percentages are shown in Table 1. The size of a single sample data varies from 21 kb to 523 kb; the truncation size: 10 (every 10 bytes are truncated into a phrase); the training samples and test samples are randomly selected each time to ensure the universality of model testing.

[0138] Table 1 Statistics of different fault types in the data set

[0139]

[0140] (2) Semantic relationship analysis of the data set

[0141] In semantic analysis, the present invention also analyzes the used data set. In the present invention, nine categories are selected to cover the semantic relationships between most entity pairs, and they do not overlap. However, there are some very similar relationships, which may cause difficulties in the recognition task, such as Entity-Origin (EO), Entity-Destination (ED), and Content-Container (CC) often appear in a sample at the same time. Similarly, there are Component-Whole (CW) and Member-Collection (MC). The brief introduction and examples of the nine relationships are as follows:

[0142] (1) Causal relationship: These cancers are caused by radiation exposure.

[0143] (2) Relationship between person and institution: Telephone operator

[0144] (3) Relationship between product and manufacturer: A factory produces suits.

[0145] (4) Content - Container Relationship: Weigh a bottle of honey.

[0146] (5) Entity - Origin Relationship: A letter from a foreign country.

[0147] (6) Entity - Destination Relationship: The boy went to bed.

[0148] (7) Component - Whole Relationship: My apartment has a big kitchen.

[0149] (8) Member - Set Relationship: There are many trees in the forest.

[0150] (9) Message - Topic Relationship: The lecture is about semantics.

[0151] The specific distribution of the number of samples in each category is shown in Table 2:

[0152] Table 2 Statistical Distribution of Relationship Categories in Samples

[0153]

[0154] 2. Simulation and Analysis Based on Different Numbers of LSTM Units

[0155] In this experiment, the type of activation unit and the Batch size are kept unchanged. At the same time, the number of LSTM units is gradually increased, and the number of traversals is increased under the condition of the same number of LSTM units. The number of LSTM training samples is 10,000, and the number of test samples is 3,000; the activation unit uses Sigmoid; Batch size: 20. The relationship between the accuracy rate and the number of LSTM units is shown in Table 3, and its trend is as Figure 1 shown.

[0156] Table 3 Accuracy Rates of Fault Identification under Different Numbers of LSTM Units

[0157]

[0158] From Table 3 and Figure 1 it can be seen that: when the number of units in the LSTM remains constant, as the number of traversals increases, the fault identification accuracy rate is higher. When the number of LSTM units is the same, the more the number of LSTM units, the better the performance. However, when the number of LSTM units is kept at 512, there is a significant decrease in the accuracy rate. The reason for the decrease is that as the required data volume increases, if more than 512 LSTM units are needed, the corresponding optimization parameters need to be adjusted.

[0159] To further analyze the data, the Receiver Operating Characteristic curve (ROC curve) system was added to the results. Due to the different performances of different numbers of LSTM units, the present invention repeated the experiments under different numbers of traversal times and selected three relatively representative numbers of LSTM units for analysis: 64, 128, and 256. The Area under Curve (AUC) reflects the ability of the recognition algorithm to correctly distinguish between two types of targets. The larger the AUC, the better the performance of the algorithm. False Negative (FN), False Positive (FP), True Negative (TN), and True Positive (TP) are important parameters in the ROC curve. Specificity is defined as the True Negative Rate (TNR), and Sensitivity is defined as the True Positive Rate (TPR). In the following experiments, the threshold was set to 0.5. If the accuracy of fault recognition under different activation units was higher than the threshold, the test result was determined to be positive. As can be seen from Table 4 and Figure 2 it can be seen that in the proposed algorithm, as the number of LSTM units increases, the performance of the algorithm tends to be better within a certain interval.

[0160] Table 4 Analysis of ROC Curve and AUC under Different Numbers of LSTM Units

[0161]

[0162] 3. Simulation and Analysis Based on Different Activation Units

[0163] The number of LSTM units and Batch size in this simulation remained unchanged. At the same time, four different activation units were selected, and the number of traversal times was increased under the condition of the same activation unit. The types of activation functions were: Sofamax, Relu, tanh, and sigmoid. The number of LSTM training samples was 10,000, and the number of test samples was 3,000; the number of LSTM units was 128; the Batch size was 20. The relationship between the accuracy rate and different activation units is shown in Table 5, and its trend is as Figure 3 shown.

[0164] Table 5 Accuracy Rate of Fault Recognition under Different Activation Units

[0165]

[0166] As can be seen from Table 5 and Figure 3It can be seen that under the condition of the same activation unit, as the number of traversals increases, the fault recognition accuracy is higher. Among the same number of traversals, better accuracy can be obtained by using Softmax and Sigmoid activation units, followed by Relu. It can be seen that the more the number of traversals, the closer the performance of Relu and Sigmoid, but the change of the result obtained by using tanh is not obvious. Therefore, when choosing an activation function, Softmax and Sigmoid are more suitable for text processing under the conditions selected in the text.

[0167] The present invention also selects the above four activation functions for ROC analysis through repeated experiments under different numbers of traversals. In the following simulation, the threshold is set to 0.5. If the accuracy of fault recognition under different activation units is higher than the threshold, the test result is considered positive. From Table 6 and Figure 4 、 Figure 5 it can be seen that the results obtained by Softmax and Sigmoid activation functions are the best, and at the same time, the above conclusion is verified.

[0168] Table 6 AUC results of ROC curves under different activation units

[0169]

[0170] 4. Simulation and analysis based on different Batch sizes

[0171] In this simulation, the number of LSTM units and activation units is kept unchanged. At the same time, the size of a single batch of data Batchsize is gradually increased, and the number of traversals under the same Batch size condition is increased. The number of LSTM training samples is 10,000, and the number of test samples is 3,000; the number of LSTM units: 128; the activation unit: sigmoid. The relationship between the accuracy rate and different Batch sizes is shown in Table 7, and its trend is as Figure 6 shown.

[0172] Table 7 Accuracy rate of fault recognition under different Batch sizes

[0173]

[0174] From Table 7 and Figure 6It can be seen that under the condition of the same Batch size, as the number of traversals increases, the fault recognition accuracy is higher. Among the same number of traversals, when the Batch size value is 20, the accuracy is higher than the other two cases. When the Batch size value is 10, the accuracy rate increases with the increase of the number of traversals, but the accuracy rate lacks continuous improvement and is in an underfitting state. When the Batch size value is 50, compared with the previous two cases, the accuracy rate decreases significantly because too much data in each batch leads to overfitting.

[0175] The present invention still selects the above three different Batch sizes for ROC analysis through repeated experiments under different numbers of traversals. In the following simulation, the threshold is set to 0.48. If the accuracy of fault recognition under different Batch sizes is higher than the threshold, the test result is determined to be positive. From Table 8 and Figure 7 It can be seen that when the Batch size value is 20, it shows the best performance. However, when the Batch size value is 50, the overall ROC curve is more smoothed. It should be noted that for different data sets showing different characteristics, there should be no fixed selection range for Batch size. Since the inspection report requires a certain word length to represent the corresponding characteristics, the optimal value of Batch size is 20.

[0176] Table 8 AUC statistics of ROC curves under different Batch sizes

[0177]

[0178] On the other hand, the present invention provides a wind farm multi-source heterogeneous data processing system, including: a computer-readable storage medium and a processor;

[0179] The computer-readable storage medium is used to store executable instructions;

[0180] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the wind farm multi-source heterogeneous data processing method described in the first aspect.

[0181] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the wind farm multi-source heterogeneous data processing method described in the first aspect.

[0182] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0183] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0184] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

[0187] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for processing multi-source heterogeneous data in a wind farm, characterized in that: The method includes the following steps: Obtain multi-source heterogeneous data including the operation information of the main electrical equipment in the wind farm; Extract the operation information set of the main electrical equipment in the wind farm from the multi-source heterogeneous data; Annotate the operation information set of the main electrical equipment in the wind farm to generate a training sample set; Train the GAN-LSTM network and the RNN-LSTM network respectively through the training sample set, and fuse the GAN-LSTM network and the RNN-LSTM network to obtain a fused network; Input the collected real-time operation information of the main electrical equipment in the wind farm into the GAN-LSTM network, the RNN-LSTM network and the fused network respectively to obtain a diagnosis result; Determine the diagnosis of the main electrical equipment in the wind farm according to the diagnosis result; The method for extracting the operation information set of the main electrical equipment in the wind farm from the multi-source heterogeneous data includes the following steps: Step 2.1: Define the data structure of the operation information of the main electrical equipment in the wind farm. The data structure consists of information elements and specific element attributes of the information elements. The information elements include location information elements, type information elements and time information elements. The location information element represents the location of the main electrical equipment in the wind farm, the type information element represents the events occurring to the main electrical equipment in the wind farm, and the time information element represents the start and end times of the events; Step 2.2: Use the words that play a key role in describing the operation information of the main electrical equipment in the wind farm as feature words. According to the syntactic roles of these words in the multi-source heterogeneous data, define the types of feature words used to fill the element attributes of the operation information of the main electrical equipment in the wind farm, and construct a professional word library according to the types of feature words; Step 2.3: Based on the data structure of the operation information of the main electrical equipment in the wind farm defined in Step 2.1 and the types of feature words defined in Step 2.2, combined with the syntactic structure features and syntactic structure features of the events occurring to the main electrical equipment in the multi-source heterogeneous data, formulate a basic extraction pattern, and expand the basic extraction pattern through rules to obtain an extraction pattern library; Step 2.4: Use the collected multi-source heterogeneous data as the input text, and preprocess the input text to obtain the vocabulary sequence of the input text; Step 2.5: Use the professional word library in Step 2.2 to identify the feature words appearing in the vocabulary sequence obtained in Step 2.4, record the types of feature words in the order of appearance of the feature words in the input text, generate the feature word type sequence of the input text, and filter the input text by judging whether the types of feature words required for the element attributes of the operation information of the main electrical equipment in the wind farm are complete; Step 2.6: Segment the input text into sentences. According to the sentence set obtained by the segmentation, divide the feature word type sequence of the input text obtained in Step 2.5 into a set of feature word type sequences corresponding to the sentence set. Use the dynamic time warping (DTW) distance to measure the similarity between each feature word type sequence in the set of feature word type sequences and the feature word type sequences of each extraction pattern in the extraction pattern library, and select the extraction pattern with the highest similarity and less than a given threshold as the matching extraction pattern for the sentence; Step 2.7: Traverse the sentence set of the input text. If the sentences in the sentence set obtain a matching extraction pattern in Step 2.6, fill the feature words in the sentence into the corresponding element attributes of the main electrical equipment operation information of the wind farm according to the element attribute sequence of the matching extraction pattern, generate the main electrical equipment operation information of the wind farm corresponding to the sentence, and obtain the set of main electrical equipment operation information of the wind farm with the extracted positioning information elements and type information elements; Step 2.8: According to the different expression forms of time in the multi-source heterogeneous data, formulate a set of regular expressions for extracting the numerical values of time elements such as year, month, day, hour, minute, and second. Combine the judgment rules and use this set of regular expressions to extract the numerical values of time elements from the input text, and combine these numerical values of time elements into the element attributes of the event start time and the element attributes of the event end time to obtain the time information elements of the main electrical equipment operation information of the wind farm; Step 2.9: Fill the time information elements extracted in Step 2.8 into the set of main electrical equipment operation information of the wind farm obtained in Step 2.7 to obtain a complete set of main electrical equipment operation information of the wind farm with the main electrical equipment operation information elements; 2. The multi-source heterogeneous data processing method for a wind farm according to claim 1, wherein: The extraction pattern includes two parts: a feature word type sequence and an element attribute sequence; the feature word type sequence is the sequence of the types of feature words used to describe events in the multi-source heterogeneous data arranged in the order of precedence. The function of the feature word type sequence in the extraction pattern is to determine whether the multi-source heterogeneous data can match this extraction pattern; the element attribute sequence has the same length as the feature word type sequence. The sequence items in the element attribute sequence are the element attributes corresponding to the sequence items in the same position in the feature word type sequence in the main electrical equipment operation information of the wind farm. The function of the element attribute sequence is to map the feature words that appear in the multi-source heterogeneous data to the corresponding element attributes in the main electrical equipment operation information of the wind farm; 3. The multi-source heterogeneous data processing method for a wind farm according to claim 1, characterized in that: The preprocessing in Step 2.4 includes deleting the duplicate information in the input text and performing Chinese word segmentation on the input text; 4. The method for processing multi-source heterogeneous data of a wind farm according to claim 1, characterized in that: After the traversal in Step 2.7 is completed, judge whether the attributes of the positioning information elements and the attributes of the type information elements of the obtained main electrical equipment operation information of the wind farm are complete. If they are not complete, use the supplementary rules to fill in the missing attributes of the positioning information elements or the attributes of the type information elements of the main electrical equipment operation information of the wind farm; 5. The method for processing multi-source heterogeneous data in a wind farm according to claim 1, characterized in that: Input the collected real-time main electrical equipment operation information of the wind farm into the GAN-LSTM network, the RNN-LSTM network, and the fused network respectively to obtain the first diagnosis result, the second diagnosis result, and the third diagnosis result; determine the diagnosis of the main electrical equipment of the wind farm according to the diagnosis results, specifically including: 1) If the first diagnosis result, the second diagnosis result, and the third diagnosis result are exactly the same, determine the faulty equipment and its location according to any one of the diagnosis results; 2) If the first diagnosis result, the second diagnosis result, and the third diagnosis result are not exactly the same, determine the faulty equipment and its location according to the relationship of the locations of the faulty equipment in various diagnosis results; 3) If the first diagnostic result, the second diagnostic result, and the third diagnostic result are completely different, return to the step of obtaining multi-source heterogeneous data including the operating information of the main electrical equipment of the wind farm.

6. A multi-source heterogeneous data processing device for a wind farm, characterized in that, Including: A multi-source heterogeneous data acquisition module, used to acquire multi-source heterogeneous data including the operating information of the main electrical equipment of the wind farm; An information extraction module, used to extract the operating information set of the main electrical equipment of the wind farm from the multi-source heterogeneous data; A training sample generation module, used to label the operating information set of the main electrical equipment of the wind farm to generate a training sample set; A network training and fusion module, used to train the GAN-LSTM network and the RNN-LSTM network respectively through the training sample set, and fuse the GAN-LSTM network and the RNN-LSTM network to obtain a fused network; A diagnosis module, used to input the collected real-time operating information of the main electrical equipment of the wind farm into the GAN-LSTM network, the RNN-LSTM network, and the fused network respectively to obtain a diagnostic result, and determine the diagnosis of the main electrical equipment of the wind farm according to the diagnostic result; The information extraction module extracts the operating information set of the main electrical equipment of the wind farm from the multi-source heterogeneous data, specifically including: Step 2.1: Define the data structure of the operating information of the main electrical equipment of the wind farm. The data structure consists of information elements and specific element attributes of the information elements. The information elements include location information elements, type information elements, and time information elements. The location information element represents the location of the main electrical equipment of the wind farm, the type information element represents the events that occur to the main electrical equipment of the wind farm, and the time information element represents the start and end times of the event; Step 2.1: Use the words that play a key role in describing the operating information of the main electrical equipment of the wind farm as feature words. According to the syntactic roles played by these words in the multi-source heterogeneous data, define the type of feature words used to fill the element attributes of the operating information of the main electrical equipment of the wind farm, and construct a professional vocabulary according to the type of feature words; Step 2.3: Based on the data structure of the operating information of the main electrical equipment of the wind farm defined in Step 2.1 and the type of feature words defined in Step 2.2, combined with the syntactic structure features and syntactic structure features of the events that occur to the main electrical equipment of the wind farm described in the multi-source heterogeneous data, formulate a basic extraction pattern, and expand the basic extraction pattern through rules to obtain an extraction pattern library; Step 2.4: Use the collected multi-source heterogeneous data as the input text, preprocess the input text to obtain the vocabulary sequence of the input text; Step 2.5: Use the professional vocabulary in Step 2.2 to identify the feature words that appear in the vocabulary sequence obtained in Step 2.4, record the type of feature words in the order of appearance of the feature words in the input text, generate the feature word type sequence of the input text, and filter the input text by judging whether the type of feature words required for the element attributes of the operating information of the main electrical equipment of the wind farm is complete; Step 2.6: Segment the input text into sentences. According to the set of sentences obtained from the segmentation, segment the sequence of feature word types of the input text obtained in Step 2.5 into a set of sequences of feature word types corresponding to the set of sentences. Use the dynamic time warping (DTW) distance to measure the similarity between each sequence of feature word types in the set of sequences of feature word types and the sequences of feature word types of each extraction pattern in the extraction pattern library. Select the extraction pattern with the highest similarity and less than the given threshold as the matching extraction pattern for this sentence; Step 2.7: Traverse the set of sentences in the input text. If a sentence in the set of sentences obtains a matching extraction pattern in Step 2.6, fill the feature words in this sentence into the corresponding element attributes of the operating information of the main electrical equipment in the wind farm according to the element attribute sequence of this matching extraction pattern, generate the operating information of the main electrical equipment in the wind farm corresponding to this sentence, and obtain the set of operating information of the main electrical equipment in the wind farm with the location information elements and type information elements already extracted from the input text; Step 2.8: According to the different expression forms of time in the multi-source heterogeneous data, formulate a set of regular expressions for extracting the numerical values of time elements such as year, month, day, hour, minute, and second. Combine the judgment rules and use this set of regular expressions to extract the numerical values of time elements from the input text, and combine these numerical values of time elements into the element attributes of the event start time and the element attributes of the event end time to obtain the time information elements of the operating information of the main electrical equipment in the wind farm; Step 2.9: Fill the time information elements extracted in Step 2.8 into the set of operating information of the main electrical equipment in the wind farm obtained in Step 2.7 to obtain a complete set of operating information of the main electrical equipment in the wind farm with the operating information elements; 7. The multi-source heterogeneous data processing device for a wind farm according to claim 6, characterized in that: After the traversal in Step 2.7 is completed, judge whether the attributes of the location information elements and the attributes of the type information elements of the obtained operating information of the main electrical equipment in the wind farm are complete. If they are not complete, use the supplementary rules to fill in the missing attributes of the location information elements or the attributes of the type information elements of the operating information of the main electrical equipment in the wind farm; 8. The multi-source heterogeneous data processing device for a wind farm according to claim 6, characterized in that: The diagnostic module inputs the collected real-time operating information of the main electrical equipment in the wind farm into the GAN-LSTM network, the RNN-LSTM network, and the fused network respectively, and obtains the first diagnostic result, the second diagnostic result, and the third diagnostic result respectively. Determine the diagnosis of the main electrical equipment in the wind farm according to the diagnostic results, specifically including: 1) If the first diagnostic result, the second diagnostic result, and the third diagnostic result are exactly the same, determine the faulty equipment and its location according to any one of the diagnostic results; 2) If the first diagnostic result, the second diagnostic result, and the third diagnostic result are not completely the same, determine the faulty equipment and its location according to the relationship of the locations of the faulty equipment in various diagnostic results; 3) If the first diagnostic result, the second diagnostic result, and the third diagnostic result are completely different, return to execute the step of the multi-source heterogeneous data acquisition module to obtain the multi-source heterogeneous data containing the operating information of the main electrical equipment in the wind farm; 9. A multi-source heterogeneous data processing system for a wind farm, characterized in that, Including: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the wind farm multi-source heterogeneous data processing method according to any one of claims 1-5.

10. A non-transitory computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, the wind farm multi-source heterogeneous data processing method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Train control onboard device failure diagnosis method with LSTM (Long Short Term Memory Network) and neural network combined

    CN108536123A

  • Fan surge operation fault identification method and system

    CN112052551A