Method and system for analyzing and optimizing power data based on natural language processing
By employing natural language processing and multiple attention units with varying window lengths, the method addresses the challenge of dynamic consumption patterns in electric power data analysis, enhancing prediction accuracy and planning precision.
Patent Information
- Application Number
- CN202510470512.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
AI Technical Summary
Traditional power data analysis methods are difficult to capture the complex nonlinear relationships and long-distance dependencies in power consumption data, resulting in the inaccurate power prediction results output from the power consumption prediction model, which affects the power company's formulation of power consumption plans.
Using a natural language processing method, by setting up multiple attention units of different window lengths, combining sentiment analysis and feature extraction models, the attention units are dynamically adjusted to adapt to the dynamic changes of applied power data over time, and power consumption prediction is performed using the BI-LSTM layer and the fully connected layer.
It improves the accuracy of feature extraction and enhances the accuracy of electricity usage prediction, allowing power companies to formulate electricity usage plans more accurately and optimize power scheduling and energy utilization.
Smart Images

Figure CN120316477A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to digital data processing technology, and in particular to an analysis and optimization method and system for power data based on natural language processing. Background Art
[0002] In order to be able to adjust the electricity consumption plan of a certain industry more precisely, currently power companies often use the method of analyzing power data based on the electricity consumption data of a certain industry to predict the electricity consumption demand of that industry in a future period. Thus, when formulating an electricity consumption plan for that industry, based on the predicted electricity consumption demand of that industry, power dispatching is optimized and energy utilization efficiency is improved.
[0003] Traditional power data analysis methods include using statistical methods for power data analysis and using machine learning models for power data analysis. However, no matter which of the above analysis methods is used, it is difficult to capture the complex non-linear relationships and long-distance dependence relationships in the electricity consumption data. To solve the defects existing in the traditional power data analysis methods, currently a method of using the attention mechanism to analyze power data based on electricity consumption data has been proposed. In power data analysis, this method can help the electricity consumption prediction model capture important feature information in the electricity consumption data. Among them, the electricity consumption prediction model is used to predict the electricity consumption of a certain industry in a certain target period.
[0004] However, the above method of using the attention mechanism to analyze power data based on electricity consumption data also has certain defects. Specifically, since the electricity consumption data corresponding to each industry usually has the characteristic of changing dynamically with time, that is, the electricity consumption data fluctuates with time. For example, the electricity consumption of each industry will have seasonal changes, with more electricity consumption in winter and summer, and less electricity consumption in spring and autumn; there are differences in the electricity consumption of each industry on weekdays and on weekends, with more electricity consumption on weekdays and less electricity consumption on weekends; and whether using the global attention mechanism or the local attention mechanism to extract features from the electricity consumption data, the extraction range is constant and may be difficult to adapt to this dynamic change. Therefore, it may lead to inaccurate feature extraction.
[0005] Since the above method of using the attention mechanism to analyze power data based on electricity consumption data may lead to inaccurate feature extraction, it will ultimately lead to inaccurate power prediction results output by the electricity consumption prediction model, thereby affecting the power company's formulation of an electricity consumption plan for that industry. Summary of the Invention
[0006] The present invention sets a plurality of attention units with different window lengths in the feature extraction model connected to the electricity consumption prediction model, and obtains the information corresponding to the target industry in each historical period from each website. The specific technical solution is as follows:
[0007] In a first aspect, the present invention provides an analysis and optimization method for power data based on natural language processing, including the following processes:
[0008] S100: Sentiment analysis; including the following specific processes:
[0009] S110: Obtain the information Q1 corresponding to the target industry from various websites within the historical period T1; the data collection module can obtain the information Q2 corresponding to the target industry from various websites within the historical period T2; and so on; the data collection module can obtain the information Q n corresponding to the target industry from various websites within the historical period T n ; the information includes electricity news corresponding to the target industry, electricity policies and regulations corresponding to the target industry, and electricity dynamics and trends corresponding to the target industry;
[0010] S120: The sentiment analysis model performs sentiment analysis on the information Q1 - Q n corresponding to the target industry in each historical period, and obtains the sentiment coefficients I1 - I n corresponding to each historical period T1 - T n ;
[0011] S200: Feature extraction; including the following specific processes:
[0012] S210: Call from the data center the electricity consumption data corresponding to the electricity consumption enterprise m of the target industry in the target region within each historical period T n within the historical period T n,m ; count the electricity consumption data corresponding to m electricity consumption enterprises within the historical period T n and determine the electricity consumption coding information E n = {e n,1 , e n,2 ,..., e n,m}; the electricity consumption data corresponding to multiple electricity consumption enterprises in the target industry, and determine the electricity consumption coding information E1 - E n ;
[0013] S220: The pre-set feature extraction model extracts features from the electricity consumption coding information E1 - E n corresponding to each historical period T1 - T n to obtain the feature vectors F1 - F n corresponding to each historical period T1 - T n ;
[0014] S300: Electricity consumption prediction; specifically including the following processes: For each historical period T1 - T nPower consumption coding information E1~E n The corresponding eigenvectors F1~F n Input to the electricity consumption forecasting model; the output of the electricity consumption forecasting model is related to a target period T x The corresponding electricity consumption forecast data E x ; The electricity consumption prediction model includes a BI-LSTM layer and a fully connected layer; wherein the BI-LSTM layer is connected to the fully connected layer, and the BI-LSTM layer is used to simultaneously capture the forward and backward dependencies of the input data to better understand the context information.
[0015] Preferably, in S120, the sentiment analysis model includes a one-hot encoding module, a word embedding module, BERT and a DENSE layer; the one-hot encoding module is connected to the word embedding module, the word embedding module is connected to BERT, and BERT is connected to the DENSE layer; the one-hot encoding module is used to one-hot encode the information and generate a vector form suitable for the machine learning model; the word embedding module maps the vector corresponding to the information to a low-dimensional vector space, thereby reducing the dimension of the vector; BERT is used to perform sentiment analysis; and the DENSE layer is used to implement sentiment classification.
[0016] Preferably, in S220, the feature extraction model includes an attention adaptation unit, a plurality of attention units and a summary module;
[0017] The attention unit includes: a global attention unit, a first local attention unit, a second local attention unit, and a third local attention unit; the attention adaptation unit is used to adjust the emotion coefficients I1 to I2 based on the emotion coefficients I1 to I3 sent by the emotion analysis module. n , select the most appropriate attention unit; the global attention unit is used to consider the information of the entire input sequence and is suitable for capturing long-distance dependencies; the window length of the first local attention unit is 3, the window length of the second local attention unit is 5, and the window length of the third local attention unit is 7; the summary module is used to summarize and output the output results of the global attention unit, the first local attention unit, the second local attention unit, and the third local attention unit.
[0018] Preferably, the calculation formula of the variance of the sentiment coefficient corresponding to the global attention unit is as follows:
[0019]
[0020] in, Represents the emotion coefficients I1~I corresponding to the global attention unit n The variance of I i Indicates the relationship between each historical period T1~T n The corresponding sentiment coefficient is Indicates that all historical periods T1 to Tn The average value of the corresponding sentiment coefficients, and its calculation formula is as follows:
[0021]
[0022] The calculation formula for the variance of the sentiment coefficients I1 to I3 corresponding to the first local attention unit is as follows:
[0023]
[0024] Where, represents the variance of the sentiment coefficients I1 to I3 corresponding to the first local attention unit, I i represents the sentiment coefficients I1 to I3 corresponding to the historical periods T1 to T3, represents the average value of the sentiment coefficients I1 to I3 corresponding to the historical periods T1 to T3, and its calculation formula is as follows:
[0025]
[0026] The calculation formula for the variance of the sentiment coefficients I2 to I4 corresponding to the first local attention unit 1 is as follows:
[0027]
[0028] Where, represents the variance of the sentiment coefficients I2 to I4 corresponding to the first local attention unit, I i represents the sentiment coefficients I2 to I4 corresponding to the historical periods T2 to T4, represents the average value of the sentiment coefficients I2 to I4 corresponding to the historical periods T2 to T4, and its calculation formula is as follows:
[0029]
[0030] The calculation formula for the variance of the sentiment coefficients I3 to I5 corresponding to the first local attention unit is as follows:
[0031]
[0032] Where, represents the variance of the sentiment coefficients I3 to I5 corresponding to the first local attention unit, I i represents the sentiment coefficients I3 to I5 corresponding to the historical periods T3 to T5, represents the average value of the sentiment coefficients I3 to I5 corresponding to the historical periods T3 to T5, and its calculation formula is as follows:
[0033]
[0034] The formula for calculating the variance of the sentiment coefficients I1 to I5 corresponding to the second local attention unit is as follows:
[0035]
[0036] Among them, represents the variance of the sentiment coefficients I1 to I5 corresponding to the second local attention unit, and I i represents the sentiment coefficients I1 to I5 corresponding to historical periods T1 to T5, represents the average value of the sentiment coefficients I1 to I5 corresponding to historical periods T1 to T5, and its calculation formula is as follows:
[0037]
[0038] The formula for calculating the variance of the sentiment coefficients I2 to I6 corresponding to the second local attention unit is as follows:
[0039]
[0040] Among them, represents the variance of the sentiment coefficients I2 to I6 corresponding to the second local attention unit, and I i represents the sentiment coefficients I2 to I6 corresponding to historical periods T2 to T6, represents the average value of the sentiment coefficients I2 to I6 corresponding to historical periods T2 to T6, and its calculation formula is as follows:
[0041]
[0042] The formula for calculating the variance of the sentiment coefficients I3 to I7 corresponding to the second local attention unit is as follows:
[0043]
[0044] Among them, represents the variance of the sentiment coefficients I3 to I7 corresponding to the second local attention unit, and I i represents the sentiment coefficients I3 to I7 corresponding to historical periods T3 to T7, represents the average value of the sentiment coefficients I3 to I7 corresponding to historical periods T3 to T7, and its calculation formula is as follows:
[0045]
[0046] The formula for calculating the variance of the sentiment coefficients I1 to I7 corresponding to the third local attention unit is as follows:
[0047]
[0048] Among them, Represents the variance of the sentiment coefficients I1 to I7 corresponding to the third local attention unit, I i Represents the sentiment coefficients I1 to I7 corresponding to the historical periods T1 to T7, Represents the average value of the sentiment coefficients I1 to I7 corresponding to the historical periods T1 to T7, and its calculation formula is as follows:
[0049]
[0050] The calculation formula for the variance of the sentiment coefficients I2 to I8 corresponding to the third local attention unit is as follows:
[0051]
[0052] Among them, Represents the variance of the sentiment coefficients I2 to I8 corresponding to the third local attention unit, I i Represents the sentiment coefficients I2 to I8 corresponding to the historical periods T2 to T8, Represents the average value of the sentiment coefficients I2 to I8 corresponding to the historical periods T2 to T8, and its calculation formula is as follows:
[0053]
[0054] The calculation formula for the variance of the sentiment coefficients I3 to I9 corresponding to the third local attention unit is as follows:
[0055]
[0056] Among them, Represents the variance of the sentiment coefficients I3 to I9 corresponding to the third local attention unit, I i Represents the sentiment coefficients I3 to I9 corresponding to the historical periods T3 to T9, Represents the average value of the sentiment coefficients I3 to I9 corresponding to the historical periods T3 to T9, and its calculation formula is as follows:
[0057]
[0058] Thus, when the attention adaptation unit determines the variance corresponding to the global attention unit The variance corresponding to the first local attention unit The variance corresponding to the second local attention unit The variance corresponding to the third local attention unit In this case, it can be judged which of the above variances is the largest, and the attention unit corresponding to the largest variance is determined as the attention unit for feature extraction of the power consumption coding data E3 in the historical period T3.
[0059] Second aspect, the present invention provides an analysis and optimization system for power data based on natural language processing, including: a data collection module, a sentiment analysis module, and a power prediction module;
[0060] The power prediction module includes: a feature extraction unit and a power prediction unit;
[0061] The data collection module is connected to the feature extraction unit in the power prediction module, and is used to call the power consumption data corresponding to the electricity-consuming enterprises in each historical period in the corresponding industry from the data center;
[0062] The data collection module is also connected to the sentiment analysis module, and is used to obtain the information corresponding to each historical period in the corresponding industry from each website;
[0063] The sentiment analysis module is connected to the feature extraction unit in the power prediction module, and is used to perform sentiment analysis on the information corresponding to each historical period in the corresponding industry received by using a pre-set sentiment analysis model, and determine the sentiment coefficient of the corresponding industry in each historical period;
[0064] The feature extraction unit is connected to the power prediction unit, and is used to perform feature extraction on the power consumption data corresponding to the electricity-consuming enterprises in each historical period received by using the sentiment coefficient corresponding to each historical period and based on a feature extraction model, and generate a feature vector corresponding to the electricity-consuming enterprises in each historical period;
[0065] The power prediction unit is used to perform power prediction by using a power prediction model and based on the feature vectors corresponding to the electricity-consuming enterprises in each historical period sent by the feature extraction unit, and determine the power prediction results corresponding to one or more target periods.
[0066] The beneficial effects of the present invention are as follows:
[0067] (1) By setting attention units with multiple different window lengths and selecting the attention unit most suitable for each historical period, the present invention extracts features from the electricity consumption coding information corresponding to each historical period. Therefore, the attention unit used to extract the electricity consumption coding information corresponding to each historical period can adapt to the characteristics of the dynamic change of the power consumption data over time, and can improve the accuracy of feature extraction, ultimately making the power prediction result more accurate.
[0068] (2) The present invention takes into account the correlation between the sentiment coefficient corresponding to the target industry in each historical period and the characteristic that the electricity consumption data changes dynamically over time; and when selecting the attention unit corresponding to a specific historical period, with the sentiment coefficient of this historical period as the center and the window lengths of each attention unit as the range, multiple variance results are determined; thereby, the attention unit corresponding to the largest variance is determined as the target attention unit for processing the electricity consumption coding information corresponding to this historical period; wherein, the larger the variance of the determined sentiment coefficient, the greater the electricity consumption fluctuation, and the attention unit with the corresponding window length can better extract features from the electricity consumption coding information corresponding to this historical period. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a schematic diagram of the analysis and optimization system for power data according to an embodiment of the present invention.
[0070] Figure 2 is a schematic diagram of the structure of the sentiment analysis model according to an embodiment of the present invention.
[0071] Figure 3 is a schematic diagram of the structure of the feature extraction model according to an embodiment of the present invention.
[0072] Figure 4 is a schematic diagram of the structure of the electricity consumption prediction model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] Referring to the attached Figure 1 , the analysis and optimization system for power data based on natural language processing includes: a data collection module, a sentiment analysis module, a power prediction module, and a data output module; wherein the power prediction module includes a feature extraction unit and a power prediction unit.
[0074] Among them, the data collection module is connected to the feature extraction unit in the power prediction module, and is used to call the electricity consumption data corresponding to the electricity-consuming enterprises in the corresponding industry in each historical period from the data center. In this embodiment, with "week" as the unit, each historical period is each week. And in each historical period, the electricity consumption data corresponding to the electricity-consuming enterprises in the corresponding industry can be, for example, the total electricity consumption data corresponding to the electricity-consuming enterprises in the corresponding industry in each week.
[0075] The data collection module is also connected to the sentiment analysis module, and is used to obtain the information corresponding to the corresponding industry in each historical period from various websites. The information can be, for example, text information, and the information includes electricity news corresponding to the corresponding industry, electricity policies and regulations corresponding to the corresponding industry, and electricity dynamics and trends corresponding to the corresponding industry, etc.
[0076] The sentiment analysis module is connected to the feature extraction unit in the power prediction module, and is used to perform sentiment analysis on the received information corresponding to the corresponding industry in each historical period by using a preset sentiment analysis model, and determine the sentiment coefficient of the corresponding industry in each historical period. Among them, the sentiment coefficient ranges from 0 to 1, and the higher the value, the more optimistic the mood; the lower the value, the more negative the mood.
[0077] The feature extraction unit is connected to the power prediction unit, and is used to extract features from the received power consumption data corresponding to the electricity-consuming enterprises in each historical period by using the received sentiment coefficient corresponding to each historical period and based on a feature extraction model, and generate feature vectors corresponding to the electricity-consuming enterprises in each historical period.
[0078] The power prediction unit is connected to the data output module, and is used to perform power prediction by using a power prediction model and based on the feature vectors corresponding to the electricity-consuming enterprises in each historical period sent by the feature extraction unit, and determine the power prediction results corresponding to one or more target periods. That is, when the power company wants to predict the power consumption data corresponding to a certain industry in a specific week, the power prediction unit can output the power consumption data corresponding to the industry in the specific week. When the power company wants to predict the power consumption data corresponding to a certain industry in multiple weeks, the power prediction unit can output the power consumption data corresponding to the industry in multiple weeks.
[0079] The data output module is used to output the power prediction results corresponding to one or more target periods.
[0080] The analysis and optimization method for power data based on natural language processing disclosed in this embodiment includes: sentiment analysis, feature extraction, and power prediction. Among them, sentiment analysis is mainly used to perform sentiment analysis on the information corresponding to the corresponding industry in each historical period, and determine the sentiment coefficient corresponding to each historical period. Feature extraction is mainly used to select the corresponding range of attention units according to the sentiment coefficients corresponding to different historical periods to extract features from the power consumption data corresponding to the electricity-consuming enterprises in each historical period, and generate feature vectors corresponding to the power consumption data in each historical period. The power prediction stage is mainly used to determine the power prediction results corresponding to one or more target periods.
[0081] The overall solution will be described in sequence below:
[0082] S100: Sentiment analysis; includes the following specific processes:
[0083] S110: The data collection module obtains the information corresponding to the target industry in the target area in each historical period from various websites.
[0084] Specifically, the target industry of the above-mentioned target area can be any one of multiple industries such as the service industry, raw material production industry, handicraft production industry, or food production industry, etc., which will not be elaborated here. And in this application, taking "week" as the unit, each historical period is each week. The information corresponding to the target industry can be, for example, text information related to this industry, and the information includes, for example, electricity news corresponding to the target industry, electricity policies and regulations corresponding to the target industry, and electricity dynamics and trends corresponding to the target industry, etc.
[0085] Thus, the data collection module can obtain the information Q1 corresponding to the target industry from various websites within the historical period T1; the data collection module can obtain the information Q2 corresponding to the target industry from various websites within the historical period T2; and so on; the data collection module can obtain the information Q n corresponding to the target industry from various websites within the historical period T n .
[0086] Thus, through the above method, the data collection module can obtain the information Q1 - Q corresponding to the target industry within the target area from various websites. n .
[0087] S120: The data acquisition module sends the information Q1 - Q obtained from various websites as above n to the sentiment analysis module, and the sentiment analysis module uses a pre-set sentiment analysis model to perform sentiment analysis on the information Q1 - Q corresponding to the target industry in each historical period n , so as to obtain the sentiment coefficients I1 - I corresponding to each historical period T1 - T n . n .
[0088] Specifically, as shown in the appendix Figure 2 , the sentiment analysis model includes a one-hot encoding module, a word embedding module, BERT, and a DENSE layer; the DENSE layer is the fully connected layer. Among them, the one-hot encoding module is connected to the word embedding module, the word embedding module is connected to BERT, and BERT is connected to the DENSE layer. And the one-hot encoding module is used to perform one-hot encoding on the information and generate a vector form suitable for the machine learning model; the information is the text information. The word embedding module maps the vector corresponding to the information to a low-dimensional vector space, thereby reducing the dimension of the vector. BERT is mainly used for sentiment analysis. The DENSE layer is mainly used to implement sentiment classification.
[0089] When the sentiment analysis module receives the information Q1 - Q n , it performs sentiment analysis on the information Q1 - Q nPerform word segmentation and input the segmented words into the sentiment analysis model. Among them, word segmentation refers to the process of splitting continuous text into meaningful word or symbol units, and word segmentation mainly provides input data for subsequent tasks such as part-of-speech tagging, syntactic analysis, and semantic analysis.
[0090] Information Q1 - Q n After passing through the one-hot encoding module, word embedding module, BERT, and DENSE layer, the sentiment coefficients I1 - I corresponding to Information Q1 - Q n in different historical periods T1 - T n are output. n These sentiment coefficients I1 - I n For example, they can range from 0 to 1. And the higher the value, the more optimistic the mood; the lower the value, the more depressed the mood. For example, the sentiment coefficient I1 corresponding to Information Q1 in historical period T1, the sentiment coefficient I2 corresponding to Information Q2 in historical period T2,..., the sentiment coefficient I n corresponding to Information Q n in historical period T n .
[0091] Through the above method, the mood analysis module can determine the corresponding sentiment coefficients I1 - I of Information Q1 - Q n in different historical periods T1 - T n . n
[0092] S200: Feature extraction; includes the following specific processes:
[0093] S210: The data collection module calls the power consumption data corresponding to multiple power-consuming enterprises in the target industry of the target region from the data center in each historical period.
[0094] Specifically, when analyzing the power consumption data of a certain target industry in a certain region, m power-consuming enterprises corresponding to the target industry can be selected as the research objects, and the power consumption data can be, for example, the total power consumption data corresponding to each power-consuming enterprise in each historical period. Among them, the target industry can be any one of multiple industries such as the service industry, raw material production industry, handicraft production industry, or food production industry, which will not be elaborated here. And the target industry includes, for example, power-consuming enterprise 1 - power-consuming enterprise m.
[0095] In this embodiment, taking "week" as the unit, each historical period is each week. Thus, the data collection module can call the power consumption e 1,1 ~e 1,m of each power-consuming enterprise 1 - m in historical period T1 from the data center.For example, the data collection module calls the electricity consumption e of electricity-consuming enterprise 1 within the historical period T1 from the data center 1,1 , the data collection module calls the electricity consumption e of electricity-consuming enterprise 2 within the historical period T1 from the data center 1,2 ,..., the data collection module calls the electricity consumption e of electricity-consuming enterprise m within the historical period T1 from the data center 1,m . Thus, the data collection module can count the electricity consumption coding information E1 = {e 1,1 , e 1,2 ,..., e 1,m} corresponding to m electricity-consuming enterprises within the historical period T1
[0096] Similarly, the data collection module can call the electricity consumption e of each electricity-consuming enterprise 1 to m within the historical period T2 from the data center 2,1 ~e 2,m . For example, the data collection module calls the electricity consumption e of electricity-consuming enterprise 1 within the historical period T2 from the data center 2,1 , the data collection module calls the electricity consumption e of electricity-consuming enterprise 2 within the historical period T2 from the data center 2,2 ,..., the data collection module calls the electricity consumption e of electricity-consuming enterprise m within the historical period T2 from the data center 2,m . Thus, the data collection module can count the electricity consumption coding information E2 = {e 2,1 , e 2,2 ,..., e 2,m} corresponding to m electricity-consuming enterprises within the historical period T2
[0097] And so on
[0098] The data collection module can call the electricity consumption e of each electricity-consuming enterprise 1 to m within the historical period T n from the data center n,1 ~e n,m . For example, the data collection module calls the electricity consumption e of electricity-consuming enterprise 1 within the historical period T n from the data center n,1 , the data collection module calls the electricity consumption e of electricity-consuming enterprise 2 within the historical period T n from the data center n,2 ,..., the data collection module calls the electricity consumption e of electricity-consuming enterprise m within the historical period T n from the data center n,m . Thus, the data collection module can count the electricity consumption coding information E n = {e n , e n,1 ,..., e n,2 ,..., e n,m} corresponding to m electricity-consuming enterprises within the historical period T
[0099] Through the above method, the data collection module can call the power consumption data corresponding to multiple electricity-consuming enterprises in the target industry within the target area from the data center, and determine the electricity consumption coding information E1 to E n .
[0100] S220: The data acquisition module sends the determined electricity consumption coding information E1 to E corresponding to each historical period T1 to T n to the feature extraction unit, and the feature extraction unit uses a pre-set feature extraction model to extract features from the electricity consumption coding information E1 to E corresponding to each historical period T1 to T n and obtains the feature vectors F1 to F corresponding to each historical period T1 to T n . n n n . n .
[0101] Refer to the appendix Figure 3 , the feature extraction model includes an attention adaptation unit, multiple attention units, and a summary module. Among them, the multiple attention units include, for example, a global attention unit, a first local attention unit, a second local attention unit, and a third local attention unit.
[0102] Among them, the attention adaptation unit is used to select the most suitable attention unit based on the sentiment coefficients I1 to I sent by the sentiment analysis module n , for example, a global attention unit, a first local attention unit, a second local attention unit, or a third local attention unit.
[0103] The global attention unit is used to consider the information of the entire input sequence. In this application, it is used to indicate all the power consumption data information included in the electricity consumption coding information, and is suitable for capturing long-distance dependencies.
[0104] The first local attention unit, the second local attention unit, and the third local attention unit are used to consider the local window of the input sequence and are suitable for processing local patterns; in the present invention, they are used to indicate partial power consumption data information in the electricity consumption coding information. It should be noted that the window length of the first local attention unit is 3, the window length of the second local attention unit is 5, and the window length of the third local attention unit is 7.
[0105] The summary module is used to summarize and output the output results of the global attention unit, the first local attention unit, the second local attention unit, and the third local attention unit.
[0106] After receiving the electricity consumption coding information E1 to E corresponding to each historical period T1 to T n by the feature extraction unit nIn the case of, input the electricity consumption coding information E1 to E n into a pre-set feature extraction model. The attention adaptation unit in the feature extraction model uses the sentiment coefficients I1 to I n corresponding to each historical period T1 to T n to determine the attention unit that is most adaptable to the electricity consumption coding information E1 to E n of each historical period T1 to T n .
[0107] For example, the attention adaptation unit needs to determine the attention unit that is most adaptable to the electricity consumption coding information E3 of historical period T3. Then the attention adaptation unit first needs to determine the variances of the sentiment coefficients corresponding to each attention unit respectively; here each attention unit includes: a global attention unit, a first local attention unit, a second local attention unit or a third local attention unit.
[0108] The calculation formula for the variance of the sentiment coefficient corresponding to the global attention unit is as follows:
[0109]
[0110] Where, represents the variance of the sentiment coefficients I1 to I n corresponding to the global attention unit, I i represents the sentiment coefficients corresponding to each historical period T1 to T n , represents the average value of the sentiment coefficients corresponding to all historical periods T1 to T n , and the calculation formula is as follows:
[0111]
[0112] For the first, second, and third local attention units, since the first, second, and third local attention units can only consider the electricity consumption coding information within the window range, and for the first, second, and third local attention units, their window lengths are different, so the local features that the first, second, and third local attention units can finally capture are different. Therefore, in this embodiment, it is necessary to first determine the local attention units suitable for processing the electricity consumption coding information of different historical periods based on the sentiment coefficients corresponding to each historical period.
[0113] Since the window lengths of the first, second, and third local attention units are different, for the sentiment coefficient of any one historical period among multiple historical periods, taking the sentiment coefficient of this historical period as the center and the window length as the range, multiple variance results can be determined.
[0114] For example, the attention adaptation unit needs to determine the local attention unit that is most suitable for the historical period T3. Then the attention adaptation unit determines that the emotion coefficient corresponding to the historical period T3 is I3. Also, since the window length of the first local attention unit is 3, taking the emotion coefficient I3 as the center and 3 as the range, the variances corresponding to the emotion coefficients I1, I2, and I3, the variances corresponding to the emotion coefficients I2, I3, and I4, and the variances corresponding to the emotion coefficients I3, I4, and I5 can be determined.
[0115] The calculation formula for the variance of the emotion coefficients I1 to I3 corresponding to the first local attention unit is as follows:
[0116]
[0117] Among them, represents the variance of the emotion coefficients I1 to I3 corresponding to the first local attention unit, and I i represents the emotion coefficients I1 to I3 corresponding to the historical periods T1 to T3, represents the average value of the emotion coefficients I1 to I3 corresponding to the historical periods T1 to T3, and the calculation formula is as follows:
[0118]
[0119] The calculation formula for the variance of the emotion coefficients I2 to I4 corresponding to the first local attention unit is as follows:
[0120]
[0121] Among them, represents the variance of the emotion coefficients I2 to I4 corresponding to the first local attention unit, and I i represents the emotion coefficients I2 to I4 corresponding to the historical periods T2 to T4, represents the average value of the emotion coefficients I2 to I4 corresponding to the historical periods T2 to T4, and the calculation formula is as follows:
[0122]
[0123] The calculation formula for the variance of the emotion coefficients I3 to I5 corresponding to the first local attention unit is as follows:
[0124]
[0125] Among them, represents the variance of the emotion coefficients I3 to I5 corresponding to the first local attention unit, and I i represents the emotion coefficients I3 to I5 corresponding to the historical periods T3 to T5, Denote the average value of the sentiment coefficients I3 to I5 corresponding to the historical periods T3 to T5. The calculation formula is as follows:
[0126]
[0127] Similarly, since the window length of the second local attention unit is 5, taking the sentiment coefficient I3 as the center and 5 as the range, the variances corresponding to the sentiment coefficients I1, I2, I3, I4, and I5, the variances corresponding to the sentiment coefficients I2, I3, I4, I5, and I6, and the variances corresponding to the sentiment coefficients I3, I4, I5, I6, and I7 can be determined.
[0128] The calculation formula for the variance of the sentiment coefficients I1 to I5 corresponding to the second local attention unit is as follows:
[0129]
[0130] Among them, Denote the variance of the sentiment coefficients I1 to I5 corresponding to the second local attention unit. I i Denote the sentiment coefficients I1 to I5 corresponding to the historical periods T1 to T5. Denote the average value of the sentiment coefficients I1 to I5 corresponding to the historical periods T1 to T5. The calculation formula is as follows:
[0131]
[0132] The calculation formula for the variance of the sentiment coefficients I2 to I6 corresponding to the second local attention unit is as follows:
[0133]
[0134] Among them, Denote the variance of the sentiment coefficients I2 to I6 corresponding to the second local attention unit. I i Denote the sentiment coefficients I2 to I6 corresponding to the historical periods T2 to T6. Denote the average value of the sentiment coefficients I2 to I6 corresponding to the historical periods T2 to T6. The calculation formula is as follows:
[0135]
[0136] The calculation formula for the variance of the sentiment coefficients I3 to I7 corresponding to the second local attention unit is as follows:
[0137]
[0138] Among them, Denotes the variance of the sentiment coefficients I3 to I7 corresponding to the local first attention unit, I i Denotes the sentiment coefficients I3 to I7 corresponding to the historical periods T3 to T7, Denotes the average value of the sentiment coefficients I3 to I7 corresponding to the historical periods T3 to T7. The calculation formula is as follows:
[0139]
[0140] Similarly, since the window length of the third local attention unit is 7, taking the sentiment coefficient I3 as the center and 7 as the range, the variances corresponding to the sentiment coefficients I1, I2, I3, I4, I5, I6, and I7, the variances corresponding to the sentiment coefficients I2, I3, I4, I5, I6, I7, and I8, and the variances corresponding to the sentiment coefficients I3, I4, I5, I6, I7, I8, and I9 can be determined.
[0141] The calculation formula for the variance of the sentiment coefficients I1 to I7 corresponding to the third local attention unit is as follows:
[0142]
[0143] Among them, Denotes the variance of the sentiment coefficients I1 to I7 corresponding to the third local attention unit, I i Denotes the sentiment coefficients I1 to I7 corresponding to the historical periods T1 to T7, Denotes the average value of the sentiment coefficients I1 to I7 corresponding to the historical periods T1 to T7. The calculation formula is as follows:
[0144]
[0145] The calculation formula for the variance of the sentiment coefficients I2 to I8 corresponding to the third local attention unit is as follows:
[0146]
[0147] Among them, Denotes the variance of the sentiment coefficients I2 to I8 corresponding to the third local attention unit, I i Denotes the sentiment coefficients I2 to I8 corresponding to the historical periods T2 to T8, Denotes the average value of the sentiment coefficients I2 to I8 corresponding to the historical periods T2 to T8. The calculation formula is as follows:
[0148]
[0149] The calculation formula for the variance of the sentiment coefficients I3 to I9 corresponding to the third local attention unit is as follows:
[0150]
[0151] Among them, represents the variance of the sentiment coefficients I3 to I9 corresponding to the third local attention unit, and I i represents the sentiment coefficients I3 to I9 corresponding to the historical periods T3 to T0, represents the average value of the sentiment coefficients I3 to I9 corresponding to the historical periods T3 to T9, and the calculation formula is as follows:
[0152]
[0153] Thus, when the attention adaptation unit determines the variance corresponding to the global attention unit the variance corresponding to the first local attention unit the variance corresponding to the second local attention unit the variance corresponding to the third local attention unit it is possible to determine which of the above variances is the largest and determine the attention unit corresponding to the largest variance as the attention unit for feature extraction of the power consumption coding data E3 of the historical period T3; for example, it can be the global attention unit, the first local attention unit, the second local attention unit, or the third local attention unit. Among them, the larger the variance, the greater the power consumption fluctuation corresponding to the historical period; the smaller the variance, the smaller the power consumption fluctuation corresponding to the historical period.
[0154] The above is only an example using the power consumption coding data E3 corresponding to the historical period T3. The method for determining the attention unit corresponding to the power consumption coding data of other historical periods is as above and will not be elaborated here.
[0155] Thus, when the attention adaptation unit determines the attention units corresponding to the power consumption coding data E1 to E n corresponding to each historical period T1 to T n feature extraction is respectively performed on the power consumption coding data E1 to E n corresponding to each historical period T1 to T n using the corresponding attention units. For example, if the attention adaptation unit determines that the global attention unit is the most suitable for the power consumption coding data E3 of the historical period T3, the feature extraction unit uses the global attention unit to perform feature extraction on the power consumption coding data E3 of the historical period T3.
[0156] If the attention adaptation unit determines that the first local attention unit is the most suitable for the electricity consumption coding data E2 in the historical period T2, the feature extraction unit uses the first local attention unit to extract features from the electricity consumption coding data E2 in the historical period T2.
[0157] Further, the aggregation module aggregates the data output by the global attention unit, the first local attention unit, the second local attention unit, and the third local attention unit, and finally outputs the corresponding feature vectors F1 to F corresponding to the electricity consumption coding information E1 to E in each historical period T1 to T. n of n the corresponding n .
[0158] S300: Electricity consumption prediction; specifically, it includes the following process:
[0159] S310: The feature extraction unit sends the feature vectors F1 to F corresponding to the electricity consumption coding information E1 to E in each historical period T1 to T n of n to the electricity consumption prediction unit. The electricity consumption prediction unit uses a pre-set electricity consumption prediction model and, based on the feature vectors F1 to F, n predicts the electricity consumption situation in the target period and outputs the electricity consumption prediction result. Among them, the target period can be one period or multiple periods, which is set according to the needs of the power company and will not be elaborated here. n Specifically, when the feature extraction unit uses a pre-set feature extraction model to output the feature vectors F1 to F corresponding to the electricity consumption coding information E1 to E in each historical period T1 to T,
[0160] the electricity consumption prediction unit further inputs the feature vectors F1 to F corresponding to the electricity consumption coding information E1 to E in each historical period T1 to T n of n to the electricity consumption prediction model. n In n the case of n the electricity consumption prediction unit further inputs the feature vectors F1 to F corresponding to the electricity consumption coding information E1 to E in each historical period T1 to T n into the electricity consumption prediction model.
[0161] Refer to Appendix Figure 4 , the electricity consumption prediction model includes a BI-LSTM layer and a fully connected layer. Among them, the BI-LSTM layer is connected to the fully connected layer. And the BI-LSTM layer is used to capture the forward and backward dependencies of the input data simultaneously to better understand the context information.
[0162] Thus, when the electricity consumption prediction unit inputs the feature vectors F1 to F corresponding to the electricity consumption coding information E1 to E in each historical period T1 to T n of n the corresponding nWhen input into the electricity consumption prediction model, the fully connected layer of the electricity consumption prediction model can output electricity consumption prediction data corresponding to the target industry in the target period. Among them, the target period can be, for example, one period or multiple periods.
[0163] For example, if you want to determine the electricity consumption prediction data corresponding to the target industry within a target period T x the electricity consumption prediction model can output the electricity consumption prediction data E x corresponding to a target period T x .
[0164] If you want to determine the electricity consumption prediction data corresponding to the target industry within multiple target periods the electricity consumption prediction model can output the electricity consumption prediction data corresponding to multiple target periods
Claims
1. An analysis and optimization method for power data based on natural language processing, characterized in that It includes the following processes: S100: Sentiment analysis; It includes the following specific processes: S110: Obtain the information Q1 corresponding to the target industry within the historical period T1 from various websites; the data collection module can obtain the information Q2 corresponding to the target industry within the historical period T2 from various websites; and so on; the data collection module can obtain the information Q corresponding to the target industry within the historical period T n within, the information Q corresponding to the target industry n ; The information includes electricity news corresponding to the target industry, electricity policies and regulations corresponding to the target industry, and electricity dynamics and trends corresponding to the target industry; S120: The sentiment analysis model performs sentiment analysis on the information Q1-Q corresponding to the target industry in each historical period n to obtain the sentiment coefficients I1-I corresponding to each historical period T3-T n ; n ; S200: Feature extraction; It includes the following specific processes: S210: Invoke from the data center the electricity consumption data corresponding to the electricity consumption e of the electricity-consuming enterprise m in the target industry in the target region during each historical period T n within the historical period T n,m ; Count the electricity consumption data corresponding to m electricity-consuming enterprises within the historical period T n and determine the electricity consumption coding information E n = {e n,1 , e n,2 ,..., e n,m}; The electricity consumption data corresponding to multiple electricity-consuming enterprises in the target industry, and determine the electricity consumption coding information E1 to E n ; S220: A pre-set feature extraction model extracts features from the electricity consumption coding information E1 to E corresponding to each historical period T1 to T n to obtain the feature vectors F1 to F corresponding to each historical period T1 to T n ; n n ; S300: Power consumption prediction; specifically including the following process: Input the power consumption coding information E1 to E corresponding to each historical period T1 to T n into the corresponding feature vectors F1 to F n and input them into the power consumption prediction model; the power consumption prediction model outputs the power consumption prediction data E corresponding to a target period T n ; the power consumption prediction model includes a BI-LSTM layer and a fully connected layer; x x Among them, the BI-LSTM layer is connected to the fully connected layer. The BI-LSTM layer is used to capture the forward and backward dependencies of the input data simultaneously to better understand the context information.
2. The analysis and optimization method of power data based on natural language processing according to claim 1, characterized in that In the S120, the sentiment analysis model includes a one-hot encoding module, a word embedding module, BERT, and a DENSE layer; the one-hot encoding module is connected to the word embedding module, the word embedding module is connected to BERT, and BERT is connected to the DENSE layer; The one-hot encoding module is used to perform one-hot encoding on the information and generate a vector form suitable for the machine learning model; the word embedding module maps the vector corresponding to the information to a low-dimensional vector space to reduce the dimension of the vector; BERT is used for sentiment analysis; the DENSE layer is used to implement sentiment classification.
3. The analysis and optimization method of power data based on natural language processing according to claim 1, characterized in that In the S220, the feature extraction model includes an attention adaptation unit, multiple attention units, and a summary module; The attention unit includes: a global attention unit, a first local attention unit, a second local attention unit, and a third local attention unit; the attention adaptation unit is used to select the most suitable attention unit based on the sentiment coefficients I1 to I sent by the sentiment analysis module n , and select the most appropriate attention unit; the global attention unit is used to consider the information of the entire input sequence and is suitable for capturing long-range dependencies; the window length of the first local attention unit is 3, the window length of the second local attention unit is 5, and the window length of the third local attention unit is 7; the aggregation module is used to aggregate and output the output results of the global attention unit, the first local attention unit, the second local attention unit, and the third local attention unit.
4. The analysis and optimization method of power data based on natural language processing according to claim 3, characterized in that The formula for calculating the variance of the sentiment coefficient corresponding to the global attention unit is as follows: Among them, represents the variance of the sentiment coefficients I1 to I corresponding to the global attention unit n , I i represents the sentiment coefficients corresponding to each historical period T1 to T n , represents the average value of the sentiment coefficients corresponding to all historical periods T1 to T n , and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I1-I3 corresponding to the first local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I1 to I3 corresponding to the first partial attention unit, I i represents the sentiment coefficients I1 to I3 corresponding to the historical periods T1 to T3, represents the average value of the sentiment coefficients I1 to I3 corresponding to the historical periods T1 to T3, and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I2-I4 corresponding to the first local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I2 to I4 corresponding to the first partial attention unit, I i represents the sentiment coefficients I2 to I4 corresponding to the historical periods T2 to T4, represents the average value of the sentiment coefficients I2 to I4 corresponding to the historical periods T2 to T4, and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I3-I5 corresponding to the first local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I3 to I5 corresponding to the first local attention unit, I i represents the sentiment coefficients I3 to I5 corresponding to the historical periods T3 to T5, represents the average value of the sentiment coefficients I3 to I5 corresponding to the historical periods T3 to T5, and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I1-I5 corresponding to the second local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I1 to I5 corresponding to the second partial attention unit, I i represents the sentiment coefficients I1 to I5 corresponding to the historical periods T1 to T5, represents the average value of the sentiment coefficients I1 to I5 corresponding to the historical periods T1 to T5, and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I2-I6 corresponding to the second local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I2 to I6 corresponding to the second partial attention unit, I i represents the sentiment coefficients I2 to I6 corresponding to the historical periods T2 to T6, represents the average value of the sentiment coefficients I2 to I6 corresponding to the historical periods T2 to T6, and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I3-I7 corresponding to the second local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I3 to I7 corresponding to the second partial attention unit, I i represents the sentiment coefficients I3 to I7 corresponding to the historical periods T3 to T7, represents the average value of the sentiment coefficients I3 to I7 corresponding to the historical periods T3 to T7, and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I1-I7 corresponding to the third local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I1 to I7 corresponding to the third partial attention unit, I i represents the sentiment coefficients I1 to I7 corresponding to the historical periods T1 to T7, represents the average value of the sentiment coefficients I1 to I7 corresponding to the historical periods T1 to T7, and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I2-I8 corresponding to the third local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I2 to I8 corresponding to the third partial attention unit, I i represents the sentiment coefficients I2 to I8 corresponding to the historical periods T2 to T8, represents the average value of the sentiment coefficients I2 to I8 corresponding to the historical periods T2 to T8, and its calculation formula is as follows: The formula for calculating the variance of the sentiment coefficients I3-I9 corresponding to the third local attention unit is as follows: Among them, represents the variance of the sentiment coefficients I3 to I9 corresponding to the third partial attention unit, I i represents the sentiment coefficients I3 to I9 corresponding to the historical periods T3 to T9, represents the average value of the sentiment coefficients I3 to I9 corresponding to the historical periods T3 to T9, and its calculation formula is as follows: Thus, when the attention adaptation unit determines the variance corresponding to the global attention unit the variance corresponding to the first local attention unit the variance corresponding to the second local attention unit the variance corresponding to the third local attention unit it can determine which of the above variances is the largest, and determine the attention unit corresponding to the largest variance as the attention unit used for feature extraction of the power consumption coding data E3 in this historical period T3.
5. An analysis and optimization system for power data based on natural language processing, characterized in that, It includes: A data collection module, a sentiment analysis module, and a power prediction module; The power prediction module includes: a feature extraction unit and a power prediction unit; The data collection module is connected to the feature extraction unit in the power prediction module and is used to call the power consumption data corresponding to the electricity-consuming enterprises in the corresponding industries in each historical period from the data center; The data collection module is also connected to the sentiment analysis module and is used to obtain the information corresponding to the corresponding industries in each historical period from various websites; The sentiment analysis module is connected to the feature extraction unit in the power prediction module and is used to perform sentiment analysis on the information corresponding to the corresponding industries received in each historical period by using a pre-set sentiment analysis model and determine the sentiment coefficients of the corresponding industries in each historical period; The feature extraction unit is connected to the power prediction unit and is used to extract features from the received power consumption data corresponding to electricity-consuming enterprises in each historical period by using the received sentiment coefficients corresponding to each historical period and based on a feature extraction model, and generate feature vectors corresponding to the electricity-consuming enterprises in each historical period; The power prediction unit is used to perform power prediction by using a power prediction model and based on the feature vectors corresponding to the electricity-consuming enterprises in each historical period sent by the feature extraction unit, and determine the power prediction results corresponding to one or more target periods.
6. The analysis and optimization system of power data based on natural language processing according to claim 5, characterized in that It further includes: A data output module; The data output module is used to output the power prediction results corresponding to one or more target periods; The power prediction unit is connected to the data output module.