A method and system for collecting and transmitting data of energy storage power station

By standardizing the data of energy storage power stations and quantifying segmented texts, combining clustering analysis and sentence embedding model, the problem of redundant data in data collection and transmission of energy storage power stations is solved, and data transmission efficiency and intelligence level are improved.

CN119226446BActive Publication Date: 2025-05-20SHANDONG SIJI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411754762.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-20
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Energy storage power stations have redundant data problems in data collection, processing and analysis, resulting in an increase in storage and transmission load pressure, affecting intelligent management and economic benefits.

Method used

By standardizing the data of energy storage power stations and quantizing segmented texts, clustering analysis is performed to identify redundant and anomaly data, and data transmission is carried out using a keyword-based sentence embedding model.

Benefits of technology

It effectively reduces redundant data, improves data transmission efficiency, reduces storage and transmission loads, and improves the intelligence level and economic benefits of energy storage power stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119226446B_ABST
    Figure CN119226446B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of energy storage data transmission, and in particular to a data collection and transmission method and system for an energy storage power station. The present invention obtains a digital vector sequence and a text vector sequence by standardizing and quantizing the original logs of each electric energy storage unit in a segmented manner; performs a first clustering analysis on the original logs based on the text vector sequence to obtain the feature category of each electric energy storage unit; analyzes the distribution characteristics of the original logs in the feature category to obtain the redundant weight of each electric energy storage unit; analyzes the correlation and difference between the digital vector sequences, and obtains the abnormal expression of the original logs in combination with the redundant weights; performs a second clustering analysis on the original logs according to the abnormal expression to obtain a normal log set; uses the normal log set to train a keyword-based sentence embedding model, and transmits the original logs based on the sentence embedding model, which can effectively reduce the load pressure of storage and transmission and improve transmission efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of energy storage data transmission, and particularly to a method and system for data collection and transmission of an energy storage power station. Background Art

[0002] As an important part of modern power systems, energy storage power stations aim to solve the problems of intermittency and instability of renewable energy generation. In the operation of energy storage power stations, log data plays a crucial role. These log data record information such as the real-time operation status of the power station, energy storage and release conditions, and equipment performance, providing an important basis for operation and maintenance decisions. By analyzing the log data, the scheduling strategy of the energy storage system can be optimized, the charge and discharge efficiency can be improved, and the operation cost can be reduced. In addition, the log data also provides data support for fault diagnosis, predictive maintenance, and performance evaluation, facilitating intelligent management.

[0003] With the progress of technology and the growth of market demand, the application scenarios of energy storage power stations are constantly expanding, including participation in the power market, demand response, and power grid regulation. However, to fully unleash the potential of energy storage power stations, more investment is still needed in data collection, processing, and analysis capabilities to improve their intelligent level and economic benefits.

[0004] Since an energy storage power station is composed of multiple groups of electrical energy storage units, that is, each electrical energy storage unit has multiple batteries, the condition of the electrical energy storage units is mainly described in the form of monitoring logs during the monitoring process of the energy storage power station. However, in the actual monitoring process, there are a large number of identical texts in the monitoring logs, that is, the similarity of the log texts is extremely high. Then, a large amount of redundant data will be generated during storage, which will lead to the load pressure on the storage and transmission of the collected monitoring logs. Summary of the Invention

[0005] In order to solve the above technical problems, the purpose of the present invention is to provide a method and system for data collection and transmission of an energy storage power station.

[0006] According to the first aspect of the embodiments of the present invention, a method for data collection and transmission of an energy storage power station is provided, and the specific technical solution adopted is as follows:

[0007] Monitor the data of each electrical energy storage unit of the energy storage power station to obtain the original logs, and the original logs include digital data and text data;

[0008] Perform standardization processing on the digital data to obtain a digital vector sequence, and perform segmented text quantization on the text data to obtain a text vector sequence;

[0009] Based on the text vector sequence, perform the first clustering analysis on the original logs to obtain the characteristic categories of each electrical energy storage unit;

[0010] Analyze the distribution characteristics of the original logs within the feature category to obtain the redundancy weights of each electrical energy storage unit;

[0011] Analyze the correlation and difference between the digital vector sequences, and combine the redundancy weights to obtain the abnormal manifestation of the original logs;

[0012] According to the abnormal manifestation, perform a second clustering analysis on the original logs to obtain a normal log set and a suspected abnormal log set;

[0013] Use the normal log set to train a keyword-based sentence embedding model, and perform monitoring data transmission based on the sentence embedding model.

[0014] In some embodiments of the present invention, performing segmented text quantization on the text data to obtain a text vector sequence includes:

[0015] According to the stop word list, remove the stop words in the text data to obtain effective text;

[0016] Use the Jieba word segmentation algorithm to perform word segmentation on the effective text to obtain the keywords of the effective text;

[0017] Perform equal-proportion segmentation on all the original logs to obtain segmentation segments;

[0018] Based on the keywords, sequentially construct a dictionary for the text data within the segmentation segments corresponding to all the original logs to obtain a word vector dictionary;

[0019] Based on the word vector dictionary, perform text vectorization on the text data to obtain a text vector sequence.

[0020] In some embodiments of the present invention, based on the text vector sequence, performing a first clustering analysis on the original logs to obtain the feature categories of each electrical energy storage unit includes:

[0021] Perform modulus conversion on the word vectors in the text vector sequence to obtain the text sequence of the original logs and the modulus values of the word vectors;

[0022] Use the Mann-Kendall (MK) trend test algorithm to perform test analysis on the text sequence to obtain the test statistic corresponding to the text vector sequence;

[0023] Calculate the average value of the modulus values of all the word vectors in the text vector sequence to obtain the data distribution of the text vector sequence;

[0024] Based on the test statistic and the data distribution, perform the first clustering analysis on the original log using the ISODATA algorithm to obtain the characteristic categories of each electrical energy storage unit.

[0025] In some embodiments of the present invention, based on the test statistic and the data distribution, perform the first clustering analysis on the original log using the ISODATA algorithm to obtain the characteristic categories of each electrical energy storage unit, including:

[0026] Construct a two-dimensional analysis space with the data distribution as the horizontal axis and the test statistic as the vertical axis;

[0027] Based on the two-dimensional analysis space, perform the first clustering analysis on the original log using the ISODATA algorithm to obtain several categories;

[0028] Obtain the number of original logs of each electrical energy storage unit in each of the categories, and select the category with the largest number of original logs containing the electrical energy storage unit as the characteristic category of the electrical energy storage unit.

[0029] In some embodiments of the present invention, analyze the distribution characteristics of the original logs within the characteristic category to obtain the redundancy weights of each electrical energy storage unit, including:

[0030] Obtain the number of original logs of each electrical energy storage unit in its corresponding characteristic category, and obtain the total number of original logs contained in the characteristic category corresponding to each electrical energy storage unit;

[0031] Calculate the ratio of the number of original logs of each electrical energy storage unit in its corresponding characteristic category to the total number of original logs contained in the characteristic category corresponding to each electrical energy storage unit to obtain the redundancy weights of each electrical energy storage unit.

[0032] In some embodiments of the present invention, analyze the correlation and difference between the digital vector sequences, and combine the redundancy weights to obtain the abnormal performance of the original log, including:

[0033] Count the number of word vectors in all the digital vector sequences to obtain a standard digital vector sequence;

[0034] Perform DTW matching on each of the digital vector sequences with the standard digital vector sequence to obtain a regularized digital vector sequence corresponding to each of the digital vector sequences;

[0035] Calculate the Pearson correlation coefficient between each of the regularized digital vector sequences and the standard digital vector sequence to obtain the correlation corresponding to each of the regularized digital vector sequences;

[0036] Calculate the modulus difference between each of the regular digital vector sequences and the corresponding vectors in the standard digital vector sequence to obtain the difference corresponding to each of the regular digital vector sequences;

[0037] Based on the correlation and the difference, and in combination with the redundancy weight, obtain the abnormal performance of the original log.

[0038] In some embodiments of the present invention, according to the abnormal performance, perform a second clustering analysis on the original log to obtain a normal log set and a suspected abnormal log set, including:

[0039] Based on the abnormal performance of all the original logs, use the density clustering algorithm DBSCAN to perform a second clustering analysis, where the radius is 0.1 and the number of neighboring points within the radius is the number of energy storage units, to obtain several clusters;

[0040] Take the original logs that can form a cluster as normal logs, and take the original logs that cannot form a cluster as differential logs;

[0041] All the normal logs form a normal log set, and all the differential logs form a suspected abnormal log set.

[0042] In some embodiments of the present invention, based on the sentence embedding model, perform monitoring data transmission, including:

[0043] Transmit the suspected abnormal log set, and transmit the keywords of the original logs in the normal log set;

[0044] At the analysis end, use the sentence embedding model and the keywords of the original logs to restore the original logs in the normal log set.

[0045] According to the second aspect of the embodiments of the present invention, a data acquisition and transmission system for an energy storage power station is provided, including: a memory and a processor, where:

[0046] The memory is used to store program code;

[0047] The processor is used to read the program code stored in the memory and execute the method described in the first aspect of the embodiments of the present invention.

[0048] In some embodiments of the present invention, the processor includes:

[0049] A log data monitoring module, which is used to perform data monitoring on each energy storage unit of the energy storage power station to obtain original logs, and the original logs include digital data and text data;

[0050] A log data processing module, which is used to perform normalization processing on the digital data to obtain a digital vector sequence, and perform segmented text quantization on the text data to obtain a text vector sequence;

[0051] A redundant weight analysis module, which is used to perform a first clustering analysis on the original log based on the text vector sequence to obtain the characteristic categories of each energy storage unit; and analyze the distribution characteristics of the original log within the characteristic categories to obtain the redundant weights of each energy storage unit;

[0052] A log judgment module, which is used to analyze the correlation and difference between the digital vector sequences, and combine the redundant weights to obtain the abnormal performance of the original log; and perform a second clustering analysis on the original log according to the abnormal performance to obtain a normal log set and a suspected abnormal log set;

[0053] A sentence embedding model training module, which is used to train a keyword-based sentence embedding model using the normal log set and perform monitoring data transmission based on the sentence embedding model.

[0054] Compared with the prior art, a data acquisition and transmission method and system for an energy storage power station provided by the present invention has the following beneficial effects:

[0055] First, by performing segmented text quantization on the text data in the original log to obtain a text vector sequence, it is avoided that word vectors with differences will be arranged at the back of the dictionary, and the interval from normal word vectors is relatively far, that is, the vector difference is too large, resulting in the distortion of the meaning of the vector.

[0056] Second, by analyzing the difference in the text vector sequence caused by the internal resistance of the same energy storage unit, it is used as the redundant weight of a digital vector sequence with higher sensitivity when analyzing the similarity of the energy storage unit.

[0057] Third, by analyzing all the text vector sequences and digital vector sequences of the same energy storage unit respectively, and using the redundant weight obtained by analyzing the text vector sequence as the weight when analyzing the digital vector sequence, the accuracy of judging the abnormal performance of the original log is improved, and further the reliability of the normal log set and the suspected abnormal log set is improved.

[0058] Fourth, by training a keyword-based sentence embedding model using the normal log set, it can ensure the transmission quality while effectively reducing the storage and transmission load pressure and improving the transmission efficiency. Description of the Drawings

[0059] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0060] Figure 1 It is a schematic diagram of the basic process of a data acquisition and transmission method for an energy storage power station provided by an embodiment of the present invention;

[0061] Figure 2 It is a schematic diagram of the basic composition of a data acquisition and transmission system for an energy storage power station provided by an embodiment of the present invention. Detailed implementation manners

[0062] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features, and effects of a data acquisition and transmission method and system for an energy storage power station according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. Terms such as "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the element.

[0064] The following specifically describes the specific solution of a data acquisition and transmission method for an energy storage power station provided by the present invention in conjunction with the accompanying drawings.

[0065] Please refer to Figure 1 , which shows the basic process of a data acquisition and transmission method for an energy storage power station provided by an embodiment of the present invention.

[0066] As Figure 1 shown, a data acquisition and transmission method for an energy storage power station provided by an embodiment of the present invention specifically includes:

[0067] S100: Monitor the data of each electrical energy storage unit in the energy storage power station to obtain the original log, where the original log includes digital data and text data.

[0068] The monitoring log data of the electrical energy storage unit usually includes the operating status data of each electrical energy storage unit (battery, SOC status, voltage, current, temperature), performance indicators (charge and discharge efficiency, input and output of system power, records of faults and abnormal events), as well as environmental data and maintenance records.

[0069] Therefore, first install sensors at the corresponding positions where each electrical energy storage unit needs to be monitored, and use a data acquisition system (SCADA) to collect and summarize the data of each sensor. During the summarization, the data transmission between the sensors, devices, and the acquisition system is realized through an industrial standard protocol (MODBUS), so as to obtain the original log of each electrical energy storage unit in the data acquisition system at each time node. A log is generated every 15 minutes, that is, there is a unique original log for all electrical energy storage units at each time node. Each original log contains the recorded digital data and text data (text description log). And because the data acquisition system has the same rule for summarizing all data, the positions where the digital data and text data of the same description object exist in the log are the same.

[0070] S200: Perform standardization processing on the digital data to obtain a digital vector sequence, and perform segmented text quantization on the text data to obtain a text vector sequence.

[0071] Currently, the main steps for text keyword extraction are removing stop words, model training, keyword extraction, and obtaining a keyword set. However, when extracting keywords from the original log, the stop words belonging to digital data may be more important than those of text data. Using the same rule to extract keywords from digital data and text data will lead to extraction deviation. Therefore, in this embodiment, the digital data is respectively standardized to obtain a digital vector sequence, and the text data is segmented and text-quantized to obtain a text vector sequence. The specific operations are as follows:

[0072] First, for the digital data, obtain all characters (strings), identifiers, numbers, and stop words in each original log through a stop word table. Arrange all the characters (strings) (digital text units) and numbers in the order of the text content, and construct a word vector table based on these characters to obtain the digital vector sequence in each original log. Since the data acquisition system has the same rule for summarizing all data, the digital vector sequences of different original logs only differ in numbers, and the characters (strings) as digital units are the same.

[0073] Meanwhile, perform segmented text quantization on the text data to obtain a text vector sequence, specifically including:

[0074] First, according to the stop word list, remove the stop words in the text data to obtain valid text; use the Jieba word segmentation algorithm to perform word segmentation on the valid text to obtain the keywords of the valid text, that is, all the keywords in each original log can be obtained.

[0075] Then, perform equal-proportion segmentation on all the original logs to obtain segmented segments; based on the keywords, sequentially construct a dictionary for the text data within the segmented segments corresponding to all the original logs to obtain a word vector dictionary. In the traditional method of vectorizing text, it is necessary to complete the construction of the word vector dictionary for all the keywords of each piece of text before constructing the word vector dictionary for the next original log. Then, after vectorizing the original log based on the constructed word vector dictionary, there will be no regularity in the text. To ensure that after vectorizing the original log using the constructed word vector dictionary, the resulting text vector sequence still has the described regularity. Therefore, in this embodiment, all the original logs are segmented in equal proportion to obtain segmented segments, and the number of segmented segments is ten. Then, based on the keywords, sequentially construct a dictionary for the text data within the segmented segments corresponding to all the original logs to obtain a word vector dictionary. Further elaboration: construct a dictionary for all the text data corresponding to the same segment after segmentation. When the dictionary construction of the first segment of all the original logs is completed, construct the dictionary for the second segment, so that the texts with similar constructed word vector dictionaries have an approximate trend, that is, the information and content positions of similar logs are together in the text vector. Then, when there is a difference in the position between the text vector of an original log and the text vectors of other original logs, it means that there is a difference in the similarity between this text vector and the other text vectors belonging to the same group, which can more prominently show the differences between the original logs. If the traditional dictionary construction method is used, the vectors of the words with differences will be arranged at the back of the dictionary, and the interval from the normal word vectors is relatively far, that is, the vector difference is too large, resulting in the distortion of the meaning of the vectors. Based on the word vector dictionary, perform text vectorization on the text data to obtain a text vector sequence.

[0076] Finally, based on the word vector dictionary, perform text vectorization on the text data to obtain a text vector sequence, that is, the text vector sequence corresponding to each original log can be obtained.

[0077] S300: Based on the text vector sequence, perform the first clustering analysis on the original logs to obtain the characteristic categories of each electric energy storage unit.

[0078] After obtaining the digital vector sequence and text vector sequence of all the original logs, since the information described by the digital vector sequence and the text vector sequence is different, that is, the digital vector sequence describes the sensor detection information of the energy storage unit, while the text vector sequence describes the operating state of the energy storage unit, it is necessary to analyze all the text vector sequences and digital vector sequences of the same energy storage unit separately, so as to distinguish the normal original logs from the abnormal original logs.

[0079] For different energy storage units, the internal resistance of their battery packs (the internal resistance of a battery refers to the internal resistance of the battery, which is a scalar, usually expressed in ohms. The smaller the internal resistance, the better the stability of the battery output voltage and the greater the output power) is different, and this internal resistance will affect the input of current, the output of current, and the power situation. Therefore, it is first necessary to analyze the difference in the text vector sequence caused by the internal resistance influence coefficient of the same energy storage unit, which is used as the redundant weight of the more sensitive digital vector sequence when analyzing the similarity of the energy storage unit.

[0080] The information that the text vector can reflect is more ambiguous than the digital vector. Since the influence of the internal resistance is mainly on the overall energy storage unit, and the corresponding text vector is also a description of the state of the energy storage power supply, in this embodiment, all the text vectors of the original logs are analyzed based on the text vector sequence, and then it is analyzed whether the same energy storage unit meets the uniqueness, so as to obtain the redundant weight.

[0081] Based on the above analysis, in this embodiment, first, based on the text vector sequence, the first clustering analysis is performed on the original logs to obtain the characteristic categories of each energy storage unit. Further, it includes: modulus normalization of the word vectors in the text vector sequence to obtain the text sequence of the original logs and the modulus of the word vectors; using the Mann-Kendall (MK) trend test algorithm to perform test analysis on the text sequence to obtain the test statistic corresponding to the text vector sequence; calculating the average value of the moduli of all the word vectors in the text vector sequence to obtain the data distribution of the text vector sequence; based on the test statistic and the data distribution, using the ISODATA algorithm to perform the first clustering analysis on the original logs to obtain the characteristic categories of each energy storage unit. Furthermore, based on the test statistic and the data distribution, using the ISODATA algorithm to perform the first clustering analysis on the original logs to obtain the characteristic categories of each energy storage unit, including: using the test statistic as the horizontal axis and the data distribution as the vertical axis to construct a two-dimensional analysis space; based on the two-dimensional analysis space, using the ISODATA algorithm to perform the first clustering analysis on the original logs to obtain several categories; obtaining the number of original logs of each energy storage unit in each category, and selecting the category with the largest number of original logs of the energy storage unit as the characteristic category of the energy storage unit.

[0082] Taking the th original log as an example, the further elaboration is as follows:

[0083] For all word vectors in the text vector sequence of the th original log, perform modulus conversion. Each word vector corresponds to a modulus value, so the text vector sequence is converted into a sequence composed of multiple modulus values, called the text sequence. That is, the text sequence of the th original log and the modulus values of all word vectors can be obtained. Then use the Mann-Kendall (MK) trend analysis algorithm to perform a test analysis on the text sequence of the th original log, and obtain the test statistic of the text sequence corresponding to the th original log. This test statistic mainly describes the trend of the text vector sequence of the th original log; affected by the method of constructing the word vector dictionary of the text vector, the trends of similar text vector sequences are the same, while for a text vector sequence that is offset compared to the normal text vector sequence, due to the difference in the modulus values of its word vectors from those of other text vector sequences, the results of the test statistic will be different. When the test statistic is greater than 0, it indicates an upward trend; when it is less than 0, it indicates a downward trend; when it is equal to 0, it indicates a constant trend. Therefore, the test statistic can measure the overall situation described by the text data in different original logs.

[0084] Since the test statistic can only describe the trend relationship between different text vector sequences and cannot describe the similarity or difference relationship of vector values, in this embodiment, by calculating the average value of the modulus values of all word vectors in the text vector sequence, the data distribution of the text vector sequence is obtained as the th original log corresponding data distribution. Then, the numerical distributions of similar original logs and their trend directions under this numerical distribution are approximately the same.

[0085] Based on the above logic, a two-dimensional analysis space is constructed. The horizontal axis of this two-dimensional analysis space is the data distribution, and the vertical axis is the test statistic. Use the ISODATA algorithm to cluster the original logs to obtain several categories, where the initial input value of the ISODATA algorithm is the same as the number of electrical energy storage units.

[0086] Among the obtained several categories, if only the original logs of the same electrical energy storage unit are included in the same category, and do not include or include very few original logs of other electrical energy storage units, it indicates that this electrical energy storage unit is special compared to other electrical energy storage units. Then it indicates that it may be affected by the internal resistance of the battery, resulting in an offset of its data. Therefore, obtain the number of original logs of each electrical energy storage unit in each category, and select the category with the largest number of original logs containing the electrical energy storage unit as the characteristic category of this electrical energy storage unit.

[0087] S400: Analyze the distribution characteristics of the original logs within the feature category to obtain the redundancy weights of each electrical energy storage unit.

[0088] Analyze the distribution characteristics of the original logs within the feature category to obtain the redundancy weights of each electrical energy storage unit. Specifically, by obtaining the number of original logs of each electrical energy storage unit in its corresponding feature category, and obtaining the total number of original logs included in the feature category corresponding to each electrical energy storage unit; then calculating the ratio of the number of original logs of each electrical energy storage unit in its corresponding feature category to the total number of original logs included in the feature category corresponding to each electrical energy storage unit, to obtain the redundancy weights of each electrical energy storage unit. Taking the th electrical energy storage unit as an example, its redundancy weight is calculated as follows:

[0089]

[0090] In the formula: represents the redundancy weight of the th electrical energy storage unit; represents the number of original logs of the th electrical energy storage unit in its corresponding feature category; represents the total number of original logs included in the feature category corresponding to the th electrical energy storage unit.

[0091] represents the ratio of the number of original logs of the th electrical energy storage unit in its corresponding feature category to the total number of original logs included in the feature category corresponding to the th electrical energy storage unit. The larger this ratio, the greater the proportion of the th electrical energy storage unit in its feature category, which means that this category contains fewer original logs of other electrical energy storage units, indicating that this electrical energy storage unit is more special compared to other electrical energy storage units, indicating that this electrical energy storage unit is more likely to be affected by internal resistance and cause problems in the overall system, and the corresponding redundancy weight of this electrical energy storage unit is larger.

[0092] Similarly, obtain the redundancy weights of all electrical energy storage units.

[0093] S500: Analyze the correlation and difference between digital vector sequences, and combine with the redundancy weights to obtain the abnormal performance of the original logs.

[0094] The digital vector sequence mainly contains numbers and their units. Since the data acquisition system aggregates data in the same order, the units match each other. Therefore, before analyzing the correlation and difference between digital vector sequences, different original logs need to be regularized first. In this embodiment, the standard digital vector sequence is obtained by counting the number of word vectors in all digital vector sequences; each digital vector sequence is matched with the standard digital vector sequence by DTW to obtain the regularized digital vector sequence corresponding to each digital vector sequence. The specific operation is as follows:

[0095] First, count the number of word vectors in all digital vector sequences. The digital vector sequence with the most word vectors is recorded as the standard digital vector sequence; and count the number of times each word vector appears in the standard digital vector sequence, and then use Otsu threshold segmentation to obtain all word vectors in the larger part, which are recorded as unit word vectors (characters representing units of numbers, etc.).

[0096] Then, perform DTW matching on the digital vector sequence corresponding to each original log and the standard digital vector sequence. Since the unit word vectors are invariant, the standard digital vector sequence is segmented based on every two unit word vectors of the standard digital vector sequence; in the digital vector sequence, the word vector with the smallest DTW distance from each unit word vector of the standard digital vector sequence is used as the unit word vector in the digital vector sequence of each original log. It should be noted that the DTW here is not the generalized DTW. The principle of DTW is that two digital vector sequences are traversed and subtracted to obtain the difference, and then the element with the smallest difference in each sequence is selected as the matching element (it may be one-to-many or many-to-one), and then the sum of all differences is the DTW distance. In this embodiment, the word vector with the smallest DTW distance means selecting the word vector with the smallest difference in one-to-many or many-to-one; then, based on the DTW matching result, equal-length interpolation is performed to make the length of the digital vector sequence of each original log the same as that of the standard digital vector sequence. That is, since the DTW matching is random, there may be one-to-many. Here, it is equivalent to finding two anchor points on the left and right. To make the number of vectors in the middle the same, equal-length interpolation needs to be performed, and then the regularized digital vector sequence of each original log is obtained.

[0097] After obtaining the regular digital vector sequences corresponding to each digital vector sequence, calculate the Pearson correlation coefficient between each regular digital vector sequence and the standard digital vector sequence to obtain the correlation corresponding to each regular digital vector sequence, that is, the correlation corresponding to each original log can be obtained. At the same time, calculate the modulus difference between the corresponding numerical vectors in each regular digital vector sequence and the standard digital vector sequence to obtain the difference corresponding to each regular digital vector sequence, where the numerical vector represents the vector in the regular digital vector sequence and the standard digital vector sequence except the unit word vector. According to the correlation and difference, and combined with the redundancy weight, the abnormal performance of the original log is obtained.

[0098] Construct the th abnormal performance of the original log of the calculation formula is:

[0099]

[0100] In the formula: represents the th abnormal performance of the th original log of the th electric energy storage unit; represents the Pearson correlation coefficient between the regular digital vector sequence of the th original log of the th electric energy storage unit and the standard digital vector sequence, that is, the correlation corresponding to the regular digital vector sequence of the th original log, ; represents the modulus of the th word vector except the unit word vector in the regular digital vector sequence of the th original log of the

[0101] Redundancy weight The larger the value taken, the more likely it is that the abnormality of the electric energy storage unit is caused by the internal resistance, and the smaller the abnormal performance of the corresponding electric energy storage unit; the absolute value of the Pearson correlation coefficient The larger the value, the more likely it is that the electric energy storage unit is normal, and the smaller the abnormal manifestation of the corresponding electric energy storage unit; Indicates the th th regular digital vector sequence and standard digital vector sequence of the th word vector in the original log of the electric energy storage unit, excluding the unit word vector, that is, the th difference in the modulus value of the word vector corresponding to the regular digital vector sequence of the original log. The larger this value, the greater the difference between the regular digital vector sequence of the electric energy storage unit and the standard digital vector sequence, and the greater the abnormal manifestation of the corresponding electric energy storage unit. It should be noted that for the situation of th th regular digital vector sequence of the original log of the electric energy storage unit and the standard digital vector sequence are not correlated. At this time, it can be directly determined that the original log is a differential log and classified into the suspected abnormal log set.

[0102] Here, only the word vectors in the regular digital vector sequence of the original log, excluding the unit word vector, are used to analyze the abnormal manifestation, in order to avoid the problem that the value of the abnormal manifestation is weakened due to the same unit word vectors.

[0103] S600: According to the abnormal manifestation, perform a second clustering analysis on the original logs to obtain a normal log set and a suspected abnormal log set.

[0104] According to the abnormal manifestation, perform a second clustering analysis on the original logs to obtain a normal log set and a suspected abnormal log set. Specifically, based on the abnormal manifestations of all the original logs of all the electric energy storage units, use the density clustering algorithm DBSCAN to perform a second clustering analysis, where the radius is 0.1 and the number of neighboring points within the radius is the number of electric energy storage units, to obtain several clusters; the original logs that can form clusters are used as ordinary logs, and the original logs that cannot form clusters are used as differential logs; all the ordinary logs form the normal log set, and all the differential logs form the suspected abnormal log set.

[0105] S700: Train a keyword-based sentence embedding model using the normal log set and perform monitoring data transmission based on the sentence embedding model.

[0106] For the information in the normal log set, it has a very high similarity. In this embodiment, a keyword-based sentence embedding model is trained based on the normal log set, so that the keyword is used as the storage information of each ordinary log in the normal log set, and other words that are not keywords are discarded. Then, when the ordinary logs in the normal log set need to be called finally, only the keyword needs to be input into the keyword-based sentence embedding model to obtain the complete log information.

[0107] The training process of the keyword-based sentence embedding model is as follows:

[0108] The BERT model is a revolutionary pre-trained model proposed by Google in 2018. It uses the Transformer architecture and significantly improves the performance of natural language processing tasks through bidirectional context understanding. In the BERT-KPE project, BERT is trained as a keyword locator that can identify which words or phrases are crucial for the topic of the original text. Traditional methods usually rely on statistics and rules, while BERT-KPE introduces deep learning and adaptively masters the rules of keyword selection by learning a large amount of unlabeled text.

[0109] Therefore, in this embodiment, the normal log set is used as the training set and the validation set, and the ratio of the training set to the validation set is 3:7 for model training. Among them:

[0110] The training architecture is: the Transformer architecture;

[0111] The loss function is: the mean squared error loss function of the triplet;

[0112] The training steps are as follows:

[0113] Input layer: Define an input layer using the Input class. This layer accepts a string array with a shape of (1,) (i.e., a sentence) and is named "sentence";

[0114] Backbone network processes the encoded input: Pass the encoded input to the backbone network backbone;

[0115] Global average pooling: Use the GlobalAveragePooling1D layer to perform global average pooling on the output of the backbone network;

[0116] Create a model: Use the Model class to combine the input layer, preprocessor, backbone network, and global average pooling layer into a complete model roberta_encoder. The input of this model is a sentence, and the output is the embedding representation of the sentence.

[0117] Based on the above training steps, a keyword-based sentence embedding model is obtained.

[0118] It should be noted that the normal log set used for training the sentence embedding model uses historical data or data generated in the early stage, that is, the sentence embedding model is obtained first; the original logs of each electrical energy storage unit in the energy storage power station are generated in real time (a log is generated every 15 minutes), and the pre-trained sentence embedding model is used to restore the real-time original log data.

[0119] Monitor data transmission is performed based on a sentence embedding model. Specifically, when transmitting the original log data, the suspected abnormal log set and the normal log set in the original log are identified. The entire suspected abnormal log set is transmitted, the keywords of the original logs in the normal log set are transmitted, and the pre-trained sentence embedding model is used at the analysis end to restore the original logs in the normal log set.

[0120] Based on the same inventive concept as the above method, this embodiment also provides an energy storage power station data acquisition and transmission system.

[0121] Please refer to Figure 2 , which shows the basic composition of an energy storage power station data acquisition and transmission system provided by an embodiment of the present invention.

[0122] As Figure 2 shown, an energy storage power station data acquisition and transmission system includes: a memory 10 and a processor 20, where:

[0123] The memory 10 is used to store program code;

[0124] The processor 20 is used to read the program code stored in the memory 10 and execute data monitoring on each electric energy storage unit of the energy storage power station to obtain original logs, where the original logs include digital data and text data; perform standardization processing on the digital data to obtain a digital vector sequence, and perform segmented text quantization on the text data to obtain a text vector sequence; based on the text vector sequence, perform a first clustering analysis on the original logs to obtain the characteristic categories of each electric energy storage unit; analyze the distribution characteristics of the original logs within the characteristic categories to obtain the redundancy weights of each electric energy storage unit; analyze the correlation and difference between the digital vector sequences, and combine the redundancy weights to obtain the abnormal performance of the original logs; according to the abnormal performance, perform a second clustering analysis on the original logs to obtain a normal log set and a suspected abnormal log set; use the normal log set to train a keyword-based sentence embedding model, and perform monitor data transmission based on the sentence embedding model.

[0125] Further, the processor 20 includes a log data monitoring module 21, a log data processing module 22, a redundancy weight analysis module 23, a log judgment module 24, and a sentence embedding model training module 25. Wherein:

[0126] The log data monitoring module 21 is used to perform data monitoring on each electric energy storage unit of the energy storage power station to obtain original logs, where the original logs include digital data and text data;

[0127] The log data processing module 22 is used to perform standardization processing on the digital data to obtain a digital vector sequence, and perform segmented text quantization on the text data to obtain a text vector sequence;

[0128] The redundant weight analysis module 23 is configured to perform a first clustering analysis on the original logs based on the text vector sequence to obtain the characteristic categories of each energy storage unit; and analyze the distribution characteristics of the original logs within the characteristic categories to obtain the redundant weights of each energy storage unit.

[0129] The log judgment module 24 is configured to analyze the correlation and difference between digital vector sequences, and combine the redundant weights to obtain the abnormal performance of the original logs; and perform a second clustering analysis on the original logs according to the abnormal performance to obtain a normal log set and a suspected abnormal log set.

[0130] The sentence embedding model training module 25 is configured to train a keyword-based sentence embedding model using the normal log set and perform monitoring data transmission based on the sentence embedding model.

[0131] It should be noted that the above sequence of embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized.

Claims

1. A method for data collection and transmission of energy storage power station, characterized in that: The method comprises: Performing data monitoring on each electric energy storage unit of the energy storage power station to obtain an original log, wherein the original log includes digital data and text data; Performing standardization processing on the digital data to obtain a digital vector sequence, and performing segmented text quantization on the text data to obtain a text vector sequence; Based on the text vector sequence, performing a first cluster analysis on the original log to obtain a feature category of each electric energy storage unit; Analyzing the distribution characteristics of the original logs in the characteristic category to obtain the redundancy weight of each electric energy storage unit; Analyze the correlation and difference between the digital vector sequences, and combine the redundant weights to obtain the abnormal performance of the original log; According to the abnormal performance, a second cluster analysis is performed on the original log to obtain a normal log set and a suspected abnormal log set; Using the normal log set to train a keyword-based sentence embedding model, and monitoring data transmission based on the sentence embedding model; Analyze the distribution characteristics of the original logs in the characteristic category to obtain the redundancy weight of each electric energy storage unit, including: Obtain the number of original logs of each electric energy storage unit in its corresponding feature category, and obtain the total number of original logs contained in the feature category corresponding to each electric energy storage unit; The ratio of the number of original logs of each electric energy storage unit in its corresponding feature category to the total number of original logs in the feature category corresponding to each electric energy storage unit is calculated to obtain the redundancy weight of each electric energy storage unit.

2. The energy storage power station data collection and transmission method according to claim 1 is characterized in that: The text data is segmented and quantized to obtain a text vector sequence, including: According to the stop word list, the stop words in the text data are removed to obtain a valid text; Using the Jieba word segmentation algorithm, the valid text is segmented to obtain keywords of the valid text; All the original logs are divided into equal parts to obtain divided segments; Based on the keywords, a dictionary is constructed for the text data in the segmented segments corresponding to all the original logs in turn to obtain a word vector dictionary; Based on the word vector dictionary, the text data is vectorized to obtain a text vector sequence.

3. The energy storage power station data collection and transmission method according to claim 1, characterized in that: Based on the text vector sequence, a first cluster analysis is performed on the original log to obtain feature categories of each electric energy storage unit, including: Modulo-value the word vectors in the text vector sequence to obtain the text sequence of the original log and the modulus value of the word vector; Using the Mann-Kendall (MK) trend test algorithm, the text sequence is tested and analyzed to obtain a test statistic corresponding to the text vector sequence; Calculating the average value of the modulus values ​​of all word vectors in the text vector sequence to obtain the data distribution of the text vector sequence; Based on the test statistic and the data distribution, the ISODATA algorithm is used to perform a first cluster analysis on the original log to obtain the characteristic category of each electric energy storage unit.

4. The energy storage power station data collection and transmission method according to claim 3 is characterized in that: Based on the test statistic and the data distribution, the ISODATA algorithm is used to perform a first cluster analysis on the original log to obtain the characteristic categories of each electric energy storage unit, including: The data distribution is the horizontal axis, and the test statistic is the vertical axis, constructing a two-dimensional analysis space; Based on the two-dimensional analysis space, the ISODATA algorithm is used to perform a first cluster analysis on the original logs to obtain several categories; The number of original logs of each electric energy storage unit in each of the categories is obtained, and the category containing the largest number of original logs of the electric energy storage unit is selected as the characteristic category of the electric energy storage unit.

5. The energy storage power station data collection and transmission method according to claim 1, characterized in that: Analyze the correlation and difference between the digital vector sequences, and combine the redundant weights to obtain the abnormal performance of the original log, including: Counting the number of word vectors in all the digital vector sequences to obtain a standard digital vector sequence; Perform DTW matching on each of the digital vector sequences and the standard digital vector sequence to obtain a regular digital vector sequence corresponding to each of the digital vector sequences; Calculating the Pearson correlation coefficient between each of the regular digital vector sequences and the standard digital vector sequence to obtain the correlation corresponding to each of the regular digital vector sequences; Calculate the module value difference between each of the regular digital vector sequences and the corresponding vector in the standard digital vector sequence to obtain the difference corresponding to each of the regular digital vector sequences; According to the correlation and the difference, and in combination with the redundancy weight, the abnormal performance of the original log is obtained; Performing DTW matching on each of the digital vector sequences and the standard digital vector sequence to obtain a regular digital vector sequence corresponding to each of the digital vector sequences includes: Count the number of occurrences of each word vector in the standard digital vector sequence, and then use the Otsu threshold segmentation to obtain a larger portion of all word vectors as unit word vectors; Perform DTW matching on each of the digital vector sequences and the standard digital vector sequence, and in each of the digital vector sequences, the word vector with the smallest DTW distance with each unit word vector of the standard digital vector sequence is used as the unit word vector in each of the digital vector sequences; Based on the DTW matching result, equal-length interpolation is performed so that the length of each digital vector sequence is the same as that of the standard digital vector sequence, thereby obtaining a regular digital vector sequence corresponding to each digital vector sequence.

6. The energy storage power station data collection and transmission method according to claim 1 or 5, characterized in that: According to the abnormal performance, a second cluster analysis is performed on the original log to obtain a normal log set and a suspected abnormal log set, including: Based on the abnormal performance of all original logs, the density clustering algorithm DBSCAN was used for the second clustering analysis, where the radius was 0.1 and the number of neighboring points within the radius was the number of energy storage units, and several clusters were obtained; The original logs that can form a cluster are regarded as ordinary logs, and the original logs that cannot form a cluster are regarded as differential logs. All common logs constitute a normal log set, and all difference logs constitute a suspected abnormal log set.

7. The energy storage power station data collection and transmission method according to claim 1, characterized in that: Monitoring data transmission based on the sentence embedding model includes: The suspected abnormal log set is transmitted, and the keywords of the original logs in the normal log set are transmitted; At the analysis end, the original log in the normal log set is restored using the sentence embedding model and the keywords of the original log.

8. A data collection and transmission system for an energy storage power station, characterized in that: The system comprises: a memory and a processor, wherein: The memory is used to store program code; The processor is configured to read the program code stored in the memory and execute the method according to any one of claims 1 to 7.

9. The energy storage power station data acquisition and transmission system according to claim 8, characterized in that: The processor comprises: A log data monitoring module is used to monitor the data of each electric energy storage unit of the energy storage power station to obtain an original log, wherein the original log includes digital data and text data; A log data processing module, used for performing standardization processing on the digital data to obtain a digital vector sequence, and performing segmented text quantization on the text data to obtain a text vector sequence; A redundancy weight analysis module, configured to perform a first cluster analysis on the original log based on the text vector sequence to obtain a feature category of each electric energy storage unit; and to analyze the distribution characteristics of the original log within the feature category to obtain a redundancy weight of each electric energy storage unit; A log judgment module is used to analyze the correlation and difference between the digital vector sequences, and obtain the abnormal performance of the original log in combination with the redundant weight; and according to the abnormal performance, perform a second cluster analysis on the original log to obtain a normal log set and a suspected abnormal log set; A sentence embedding model training module, used to train a keyword-based sentence embedding model using the normal log set, and monitor data transmission based on the sentence embedding model; Analyze the distribution characteristics of the original logs in the characteristic category to obtain the redundancy weight of each electric energy storage unit, including: Obtain the number of original logs of each electric energy storage unit in its corresponding feature category, and obtain the total number of original logs contained in the feature category corresponding to each electric energy storage unit; The ratio of the number of original logs of each electric energy storage unit in its corresponding feature category to the total number of original logs in the feature category corresponding to each electric energy storage unit is calculated to obtain the redundancy weight of each electric energy storage unit.

Citation Information

Patent Citations

  • Log data processing method and device, equipment and storage medium

    CN114328106A

  • Abnormal log analysis method and device and electronic equipment

    CN114528845A