Factory energy consumption prediction method and system based on production behavior recognition, and storage medium

By constructing a three-dimensional coupled feature system of time-series, semantics, and structure, production behavior is identified and operating conditions are classified. This solves the problem of non-coordinated expression of multi-source data in existing technologies, achieves high-precision industrial energy consumption prediction, and improves the stability of the model under operating condition switching and output fluctuation conditions.

CN121882376APending Publication Date: 2026-04-17CHINA TOBACCO ZHEJIANG IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610185734.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing industrial energy consumption prediction technologies fail to effectively identify production behaviors, and multi-source data do not form an effective collaborative expression, making it difficult to accurately reveal the intrinsic driving mechanism of energy consumption changes, resulting in a significant increase in prediction errors under operating condition switching or abnormal conditions.

Method used

By constructing a three-dimensional coupled feature of time-series, semantics, and structure, a multi-source dataset of energy consumption is obtained. Based on the recognition of production behavior, the working conditions are divided, and multiple models are trained to obtain the optimal model for different working conditions. The three-dimensional coupled feature of time-series, semantics, and structure is used to predict the energy consumption of the factory.

Benefits of technology

It significantly improves the expressive power and prediction accuracy of energy consumption characteristics. The prediction determination coefficient under production conditions has increased from 0.62 to 0.92, and the average absolute percentage error under non-production conditions has decreased from 23.5% to 8.3%. The overall prediction error is stably controlled within 10%, meeting the application requirements of industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121882376A_ABST
    Figure CN121882376A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial energy management and intelligent prediction, in particular to a factory energy consumption prediction method and system based on production behavior recognition and a storage medium, and the method comprises the steps: obtaining an energy consumption multi-source data set; constructing a time sequence-semantic-structure three-dimensional coupling feature based on the energy consumption multi-source data set; dividing the energy consumption multi-source data set according to a working condition identifier to obtain data sets of different working conditions; based on the time sequence-semantic-structure three-dimensional coupling features corresponding to the data sets under different working conditions, multi-model training is carried out to obtain optimal models under different working conditions; and predicting the factory energy consumption under the corresponding working conditions by adopting the optimal models of the different working conditions. According to the embodiment of the invention, the time sequence-semantic-structure three-dimensional coupling feature system is provided, and compared with an existing method which only depends on original features or linear combination, collaborative modeling of production behaviors, environmental factors and process structures is realized, and the expression ability and interpretability of energy consumption features are remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial energy management and intelligent forecasting technology, and specifically to a factory energy consumption forecasting method, system and storage medium based on production behavior identification. Background Technology

[0002] Industrial enterprises are placing higher demands on energy efficiency, accuracy of energy consumption prediction, and the level of intelligence in energy management. Industrial energy consumption is characterized by multi-source driving forces, multi-condition switching, and strong time-series dependence. Its changes are not only related to output scale but also influenced by factors such as production cycle time, process linkage, equipment start-up and shutdown strategies, and environmental and meteorological conditions. Against this backdrop, constructing models that can accurately reflect changes in production behavior and predict energy consumption has become an important research direction in the field of industrial energy management.

[0003] Currently, existing industrial energy consumption prediction technologies mainly include statistical analysis-based methods, traditional machine learning methods, and deep learning methods. For example, the method described in the literature [SARSWATULA, Sai Aravind, PUGH, Tanna, et PRABHU, Vittaldas. Modeling energy consumption using machine learning. Frontiers in Manufacturing Technology, 2022, vol. 2, p. 855-208.] involves aggregating data from MES, EMS, and meteorological systems, selecting features such as output and ambient temperature, and using regression or tree models to predict energy consumption. Other studies [WU, Qiong, REN, Hongbo, SHI, Shanshan, et al. Analysis and prediction of industrial energy consumption behavior based on big data and artificial intelligence. Energy Reports, 2023, vol. 9, p. 395-402.] introduce deep models such as Long Short-Term Memory (LSTM) networks to enhance the ability to characterize the time-series characteristics of energy consumption. However, the above methods still have the following shortcomings.

[0004] First, while existing technologies have attempted to incorporate multi-source data, they largely remain at the level of data splicing or simple feature combination. They fail to systematically model the temporal correlations, process semantic relationships, and structural constraints among production, environmental, and energy consumption data. This results in multi-source data failing to form an effective collaborative expression, making it difficult to accurately reveal the intrinsic driving mechanisms of energy consumption changes. Second, existing methods generally use energy consumption or output as the direct modeling object, without explicitly identifying and modeling production behavior states. Changes in production behavior, such as production and non-production states, different output levels, equipment start-up and shutdown, and output transmission across process segments, are often implicit in continuous numerical features. Models struggle to distinguish the differential energy consumption responses under different behavior modes, leading to significantly increased prediction errors during operating condition switching or abnormal situations.

[0005] Therefore, existing technologies still have significant shortcomings in production behavior recognition, multi-source data coupling feature construction, and working condition adaptive energy consumption prediction. There is an urgent need for an industrial energy consumption prediction method and system that can take production behavior recognition as the core, integrate multi-source data, and achieve high-precision energy consumption prediction. Summary of the Invention

[0006] The purpose of this invention is to provide a factory energy consumption prediction method, system, and storage medium based on production behavior recognition, so as to solve the shortcomings of existing technologies in production behavior recognition, multi-source data coupling feature construction, and working condition adaptive energy consumption prediction.

[0007] To achieve the above objectives, embodiments of the present invention provide a factory energy consumption prediction method based on production behavior identification, including: Obtain a multi-source energy consumption dataset; Construct a three-dimensional coupled feature system of time-series, semantics, and structure based on the aforementioned multi-source energy consumption dataset; The energy consumption multi-source dataset is divided according to the operating condition identifier to obtain datasets for different operating conditions; Based on the time-semantic-structural three-dimensional coupled features corresponding to datasets under different working conditions, multiple models are trained to obtain the optimal model for each working condition. The optimal model for different operating conditions is used to predict the factory energy consumption under the corresponding operating conditions.

[0008] Optionally, constructing a time-series-semantic-structural three-dimensional coupled feature based on the energy consumption multi-source dataset includes: By using multi-operator composite mapping, the multi-source energy consumption data is extended into a coupled feature vector containing historical states, which serves as a temporal coupled feature. Construct a high-order interactive feature tensor to map the discrete feature space into a continuous semantic manifold, which serves as a semantic coupling feature; Domain knowledge and physical constraints are embedded into the feature space to construct interpretable proxy variables as structural coupling features; The temporal coupling features, semantic coupling features, and structural coupling features are fused to obtain a three-dimensional coupling feature of temporal-semantic-structural coupling.

[0009] Optionally, through multi-operator composite mapping, the multi-source energy consumption data is extended into a coupled feature vector containing historical states, which serves as temporal coupled features including: Construct a temporal coupling model based on formula (1). (1) in, For the temporal coupling feature space, The original feature sequence, For lag operators, For rolling statistics calculation, It is a difference operator.

[0010] Optionally, the hysteresis operator includes The first-order lag characteristic and the weighted lag characteristic are obtained according to formulas (2) and (3): (2) (3) in, For weighted lag characteristics, for The original feature values ​​at time 1. For the first The normalized weight coefficients at each time step. For the length of the history window, It is an exponentially decaying kernel function. This is the attenuation coefficient.

[0011] Optionally, a higher-order interaction feature tensor is constructed to map the discrete feature space into a continuous semantic manifold, which serves as semantic coupling features, including: Based on the set of production behavior characteristics and the set of environmental factor characteristics, a tensor product representation of basic interaction characteristics is constructed. A polynomial kernel function is used to map the features in the tensor product representation to a high-dimensional space to obtain d-order polynomial interaction terms. An attention mechanism is used to dynamically adjust the weights of interaction features in order to obtain weighted semantic interaction features. Semantic coupling features are obtained based on the tensor product representation, the d-order polynomial interaction terms, and the weighted semantic interaction features.

[0012] Optionally, an attention mechanism is used to dynamically adjust the weights of the interaction features to obtain weighted semantic interaction features, including: The attention score is obtained according to formula (4). (4) in, for( , The importance of ) For the first A behavioral characteristic, For the first One environmental factor, and The projection matrix is ​​learnable. For attention weight vectors, For bias terms, For activation function, The number of behavioral characteristics, The number of environmental factors; According to formula (5), the weighted semantic interaction features are obtained. (5) in, These are the weighted semantic interaction features.

[0013] Optionally, domain knowledge and physical constraints can be embedded into the feature space to construct interpretable proxy variables, which serve as structural coupling features, including: A thermodynamically constrained cooling and heating load model is constructed to obtain an air conditioning load that takes into account asymmetric characteristics. A dehumidification load model for the humidity mass transfer process is constructed to obtain a dehumidification load that takes into account nonlinear effects. Construct a dynamic transmission model of process segment coupling to obtain the process segment coupling strength; Based on the air conditioning load considering asymmetric characteristics, the dehumidification load considering nonlinear effects, and the section coupling strength, the structural coupling characteristics are obtained.

[0014] Optionally, a dynamic transmission model of the process segment coupling is constructed to obtain the process segment coupling strength, including: The coupling strength of the work section is obtained according to formula (6). (6) in, for The coupling strength of the work section at any given time. This is a production level mapping function. for The output of the upstream section at any given time. The output of this section at any given time. and The coupling coefficient is... It is the time delay constant; A feedback adjustment mechanism is introduced based on formula (7). (7) in, for The output of the upstream section at any given time. For feedback gain coefficient, Set a value for the target output.

[0015] On the other hand, the present invention also provides a factory energy consumption prediction system based on production behavior recognition, the system comprising: The data acquisition module is used to collect energy consumption data from multiple sources. The data preprocessing module is used to preprocess the energy consumption multi-source data; The feature engineering module is used to identify production behaviors based on preprocessed data and to construct temporal features, semantic interaction features, and structural coupling features. The working condition segmentation module is used to segment the samples according to the production behavior identification results, forming data subsets under different operating conditions; The model training and evaluation module is used to train and optimize the parameters of the energy consumption prediction model under different operating conditions and to evaluate the performance of different prediction models. The model deployment and prediction module is used to load the target model and predict energy consumption based on real-time or historical data. The processor is connected to the data acquisition module, data preprocessing module, feature engineering module, working condition segmentation module, model training and evaluation module, and model deployment and prediction module, and the processor is configured to perform any of the methods described above.

[0016] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the methods described above.

[0017] The beneficial effects of this invention are: This invention proposes a three-dimensional coupled feature system of temporal-semantic-structural aspects. Compared to existing methods that rely solely on original features or linear combinations, this system achieves collaborative modeling of production behavior, environmental factors, and process structures, significantly enhancing the expressive power and interpretability of energy consumption features. Through multi-scale temporal operators, behavior-environment interaction features, and structural proxy variables based on physical mechanisms, the feature space is expanded from approximately 10 dimensions to over 50 dimensions, resulting in a significant increase in information density.

[0018] The embodiments of this invention achieve high prediction accuracy under various operating conditions. Under production conditions, the prediction determination coefficient increases from 0.62 in the baseline model to 0.92, and the mean absolute percentage error decreases from 23.5% to 8.3%. Under non-production conditions, the prediction accuracy is also significantly improved. The overall prediction error is stably controlled within 10%, meeting the application requirements for energy consumption prediction accuracy in industrial scenarios.

[0019] This invention divides data into operating conditions based on production status and models different operating conditions separately, effectively avoiding the performance fluctuation problem of a single model under multiple operating conditions. Practical results show that operating condition separation modeling can significantly improve the stability of the model under operating condition switching and output fluctuation conditions, providing a reliable basis for formulating differentiated energy-saving strategies.

[0020] The embodiments of this invention construct a complete technical process covering data acquisition, feature engineering, model training and prediction output, support standardized interfaces and automated operation, have low deployment costs and stable operation, and can meet the engineering application needs of industrial enterprises for online prediction and long-term operation.

[0021] In summary, this invention, through production behavior recognition and three-dimensional coupled modeling, achieves high-precision and stable prediction of factory energy consumption while ensuring engineering feasibility, and has high practical value and promotion significance.

[0022] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0023] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 A flowchart of a factory energy consumption prediction method based on production behavior identification according to an embodiment of the present invention; Figure 2 A flowchart illustrating a method for acquiring a multi-source energy consumption dataset according to an embodiment of the present invention; Figure 3 A flowchart illustrating a method for constructing a three-dimensional coupled temporal-semantic-structural feature according to an embodiment of the present invention; Figure 4 A flowchart illustrating a method for constructing semantically coupled features according to an embodiment of the present invention; Figure 5 A flowchart illustrating a method for constructing structural coupling features according to an embodiment of the present invention; Figure 6This is a schematic diagram of a factory energy consumption prediction system architecture based on the coupling of production behavior recognition and environmental factors according to an embodiment of the present invention. Figure 7 This is a schematic diagram illustrating the principle of constructing a three-dimensional coupled temporal-semantic-structural feature according to an embodiment of the present invention. Detailed Implementation

[0024] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0025] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0026] like Figure 1 The diagram shows a flowchart of a factory energy consumption prediction method based on production behavior identification according to an embodiment of the present invention. Figure 1 In this context, the prediction method may include the following steps: In step S10, energy consumption multi-source dataset is obtained; In step S11, a three-dimensional coupled feature of time-series, semantics, and structure is constructed based on the energy consumption multi-source dataset; In step S12, the energy consumption multi-source dataset is divided according to the operating condition identifier to obtain datasets for different operating conditions; In step S13, based on the time-series-semantic-structural three-dimensional coupled features corresponding to the datasets of different working conditions, multiple models are trained to obtain the optimal model for different working conditions. In step S14, the optimal model for different operating conditions is used to predict the factory energy consumption under the corresponding operating conditions.

[0027] In such Figure 1 In the factory energy consumption prediction method based on production behavior recognition shown, step S10 is used to acquire a multi-source energy consumption dataset. In this embodiment, the specific method for acquiring the multi-source energy consumption dataset in step S10 can be of various forms known to those skilled in the art. In one example of the present invention, step S10 may include, for example... Figure 2 The steps shown are described in this. Figure 2 In this context, step S10 may include: In step S20, the data acquisition range and acquisition parameters are determined, and energy consumption multi-source data is acquired. In step S21, the collected multi-source data is preprocessed; In step S22, core behavioral features are constructed; In step S23, a method for calculating auxiliary behavior indicators is constructed.

[0028] In such Figure 2 In the method shown, step S20 is used to define the data collection range and parameters. First, the data collection range needs to be clearly defined to ensure the completeness and relevance of the data entering the system in both time and spatial dimensions. In this example, the time window is a continuous interval from the start date to the end date, which can be represented as... The spatial range is the identifier of the target process section. Based on this, the system collects data from dimensions such as energy consumption, production, environment, equipment, and process coupling. In this example, the energy consumption dimension collects standard coal equivalent indicators such as cooling energy consumption, total energy consumption, steam energy consumption, air conditioning electricity consumption, process electricity consumption, and lighting electricity consumption, and organizes them into... Production dimension data collection of target process segment on date production Environmental dimensions include the collection and prediction of temperature and humidity formation. The equipment dimension records the maintenance date and the duration of maintenance on that day. ; The process section coupling dimension collects the output of the upstream process section. And can construct simple coupling quantities It is used to reflect the impact of output transmission between upstream and downstream processes on energy consumption.

[0029] Step S21 is used to preprocess the collected multi-source data to ensure that the multi-source data meets the modeling requirements. The data preprocessing stage sequentially performs operations such as type standardization, time standardization, outlier filtering, and missing value handling. Regarding type standardization, the original fields... Convert to numerical type Content that cannot be converted is marked as a missing value (NaN); regarding time standardization, date fields from all data sources are uniformly converted to the "year-month-day" format to ensure that different data tables can use dates... Primary key alignment is accurate; for outlier filtering, an outlier date set is pre-established. And remove The record, and at the same time the numerical features Using the three-standard-deviation criterion, when These are then marked as suspected outliers. Regarding missing value handling and data fusion, for numerical features such as temperature and humidity, the average of nearest neighbors can be used. To complete the dataset, for other missing features, data from that day will be removed and not included in the training set; finally, the data will be standardized by date. Primary key, Information such as these are aligned and merged to form a unified original feature. The dataset constitutes the time window This serves as the foundational data for subsequent feature construction and energy consumption prediction modeling.

[0030] Step S22 is used to construct core behavioral features, aiming to identify key behavioral variables reflecting the factory's operating mode, and to characterize energy consumption differences under different production states during the modeling phase. In this example, firstly, a production status indicator is constructed based on output data, with the daily output recorded as... Define production status indicator variables :when season ,when season This is to distinguish between production and shutdown conditions. Secondly, continuous output is categorized into levels, constructing output level variables. For example, let Indicates production stoppage ( ), Indicates low yield ( ), Indicates high yield ( This facilitates the model's learning of energy consumption variations under different load levels. Furthermore, based on the conversion of electricity consumption for lighting into standard coal equivalent... Identify the lighting power status when (like Define lighting status indicator variable when kgce) ,otherwise This is used to characterize changes in workshop lighting load and its consistency with production periods. Finally, the cold load is converted to standard coal equivalent. Steam conversion of heating network to standard coal Electricity consumption for air conditioning converted to standard coal equivalent Three indicators identify the air conditioner's operating status. (like Define the air conditioning operation indicator variable when kgce) ,otherwise The aforementioned air conditioning operating status characteristics are used to reflect the start-up and shutdown status of the refrigeration and heating systems, and are key variables in the coupled analysis of environmental factors and energy consumption.

[0031] Step S23 is used to construct an auxiliary behavior indicator calculation method. This method supplements the description of personnel and equipment operational characteristics, providing explanatory support for core behavior characteristics. Firstly, regarding shift work hour statistics, attendance identifiers for the morning, afternoon, and evening shifts are extracted from the shift management table. , , (Record 1 for working and 0 for not working), and calculate the total number of working shifts for the day. (Value range is 0 to 3), this indicator is used to reflect the intensity of personnel input and its corresponding relationship with equipment running time and energy consumption duration; secondly, in terms of maintenance time summary, all maintenance events of the specified production line are extracted from the equipment maintenance record table, and the date is set. Total on that day Each maintenance record has a duration of [duration]. The cumulative maintenance time for the day is defined as follows: The maintenance duration can be used to identify abnormal shutdown states and to analyze the energy consumption fluctuation characteristics during the maintenance period and the restart phase after maintenance.

[0032] Step S11 is used to construct a three-dimensional coupled feature of time-series, semantics, and structure based on the energy consumption multi-source dataset. In this embodiment, the specific method for constructing the three-dimensional coupled feature of time-series, semantics, and structure in step S11 can be of various forms known to those skilled in the art. In one example of the present invention, step S11 may include, for example... Figure 3 The steps shown are described in this. Figure 3 In this context, step S11 may include: In step S30, the multi-operator composite mapping expands the energy consumption multi-source data into a coupled feature vector containing historical states, which serves as a temporal coupled feature. In step S31, a high-order interactive feature tensor is constructed to map the discrete feature space into a continuous semantic manifold, which serves as the semantic coupling feature. In step S32, domain knowledge and physical constraints are embedded into the feature space to construct interpretable proxy variables as structural coupling features; In step S33, the temporal coupling features, semantic coupling features, and structural coupling features are fused to obtain the temporal-semantic-structural three-dimensional coupling features.

[0033] In such Figure 3 In the method shown, step S30 is used to construct temporal coupling features. Specifically, in this embodiment, a dynamic memory mechanism for energy consumption is constructed through time series analysis theory, enabling the model to capture the transmission effect of historical states on current energy consumption. Let the original feature sequence be... ,in Indicates a time index. Given the observation duration, the temporal coupling feature space... This can be represented as a multi-operator composite mapping: (1) in, For the temporal coupling feature space, The original feature sequence, For lag operators, For rolling statistics calculation, It is a difference operator.

[0034] In this example, the hysteresis operator Used to construct time-delay features and capture historical dependencies at different time scales. Definition The hysteresis characteristic is: (8) in, express Moment hysteresis characteristics, for The original feature values ​​at time 1. This represents the lag order. In this example, we use 1, 7, 14, and 30 to capture the time-series dependencies for daytime, weektime, bi-weekly, and monthly periods.

[0035] Considering the multi-cycle nature of industrial production, a weighted lag feature is introduced to enhance the expressive power of key historical information: (2) (3) in, For weighted lag characteristics, For the length of the history window, For the first Normalized weighting coefficients for each historical moment. This is the decay coefficient, used to control the decay rate of historical information weights over time. The kernel function is the exponentially decaying kernel function, and the denominator is... As a normalization factor, it ensures that the sum of all weights is 1. The exponentially decaying kernel function ensures that recent data has higher weights while preserving long-term trend information. At that time, the weight of the previous day was about 0.35, and the weight of the previous seven days was about 0.15, which achieved a balanced capture of short-term disturbances and medium-term trends.

[0036] In this example, the rolling statistic calculator Robust smoothing of temporal noise is achieved by calculating local statistics through a moving window. The window width is defined as... The rolling mean and rolling standard deviation are: (9) (10) in, for The width of the time window is The rolling mean, For the corresponding rolling standard deviation, For window width, For the first in the window The feature values ​​of each historical moment.

[0037] To enhance the model's sensitivity to abnormal fluctuations, a rolling coefficient of variation is introduced: (11) in, The rolling coefficient of variation, This is a numerical stability term used to prevent numerical overflow caused by a denominator of zero; its value is [value missing]. The coefficient of variation (CV) can quantify the intensity of local fluctuations; a sudden increase in the CV value indicates a structural change in energy consumption patterns.

[0038] Furthermore, an exponentially weighted moving average is introduced to adaptively adjust historical weights: (12) in, for The exponentially weighted moving average at time t, These are the original feature values ​​at the current moment. It is the exponentially weighted moving average of the previous time step. This is a smoothing coefficient used to adjust the relative weights of current observations and historical information in feature representation.

[0039] Adaptive adjustment of historical weights refers to adjusting the smoothing coefficient according to different production conditions, feature types, and time scales. Differentiated settings are implemented to achieve adaptive control over the influence range of historical information. Specifically, under production conditions or when characteristics fluctuate significantly, a larger setting is selected. To enhance the model's responsiveness to short-term changes; in non-production conditions or when characteristic changes are stable, a smaller [size / increase] is selected. To enhance the retention of long-term trend information. In a preferred embodiment of the invention, The value can be selected from the preset range [0.2, 0.5]. When it is necessary to consider both short-term response and long-term stability, the preferred value is... =0.3, in order to suppress noise interference while responding quickly to energy consumption fluctuations.

[0040] In this example, the difference operator Used to capture temporal rate of change and acceleration information. The first-order difference is defined as: (13) in, for The first difference value at time t, The feature value at the current time. The eigenvalue is the feature value from the previous time step.

[0041] To characterize the nonlinear acceleration features of energy consumption changes, a second-order difference is introduced: (14) in, for The second difference value at time t, The first difference value of the previous time step. These are the feature values ​​from the first two time points.

[0042] Based on momentum theory, define the momentum index for energy consumption change: (15) in, for Momentum index at any given moment. The length of the momentum calculation window (preferred in this invention) ), for Feature values ​​from a time point ago, Momentum decay factor (preferred) ), This represents the momentum value at the previous moment. The momentum index combines short-term rate of change with historical momentum information, enabling early identification of trend reversals in energy consumption.

[0043] Considering the multiple periodicities of industrial energy consumption (daily, weekly, and seasonal cycles), this example uses a seasonal decomposition algorithm to decompose the original sequence into a trend term, a seasonal term, and a residual term: (16) in, for The original feature values ​​at time 1. This is a trend term, reflecting the long-term evolution of energy consumption. For seasonal terms, capture cyclical fluctuation patterns. The residual term represents random disturbances and anomalous events.

[0044] The decomposition process iteratively solves for each component using a locally weighted regression algorithm: (17) (18) in, For locally weighted regression functions, This is a trend smoothing parameter that controls the smoothness of the trend term. This is a function for extracting cyclic subsequences. The period length (week period is taken as) ).

[0045] Combining the above operators, the final temporal coupling feature vector can be expressed as:

[0046] (19) in, for The temporal coupled feature vector at each time point comprises 15 dimensions, including original features, lag features, rolling statistical features, difference features, momentum features, and decomposition features. This feature system expands a single point-in-time observation into a 15-dimensional temporal feature space, increasing the feature dimensionality by 15 times and significantly enhancing the model's ability to characterize the dynamic evolution of energy consumption.

[0047] Step S31 is used to construct a high-order interactive feature tensor, mapping the discrete feature space into a continuous semantic manifold, serving as a semantic coupling feature. Factory energy consumption is driven by multiple factors, with complex nonlinear coupling relationships between the variables. The semantic coupling layer, by constructing a high-order interactive feature tensor, maps the discrete feature space into a continuous semantic manifold, achieving deep fusion of multi-source heterogeneous information. In this embodiment, the specific method for constructing the semantic coupling feature in step S31 can be of various forms known to those skilled in the art. In one example of this invention, step S31 may include, for example... Figure 4 The steps shown are described in this. Figure 4 In this context, step S31 may include: In step S40, a tensor product representation of the basic interaction features is constructed based on the set of production behavior features and the set of environmental factor features. In step S41, a polynomial kernel function is used to map the features in the tensor product representation to a high-dimensional space to obtain d-order polynomial interaction terms. In step S42, an attention mechanism is used to dynamically adjust the weights of the interaction features in order to obtain weighted semantic interaction features. In step S43, semantic coupling features are obtained based on tensor product representation, d-order polynomial interaction terms, and weighted semantic interaction features.

[0048] In such Figure 4 In the method shown, step S40 is used to construct the tensor product representation of the basic interaction features. Specifically, in this example, let the set of production behavior features be... This is used to characterize the state of factory production behavior, including at least output and a set of environmental factor features that identify production behavior indicators derived from output and electricity consumption status. This is used to characterize external environmental conditions, including at least temperature and humidity-related environmental features. The basic semantic interaction features can then be represented as a second-order tensor product: (20) in, It is a second-order tensor product, and the tensor contains Each interactive item, each element Characterizing the first The first behavioral characteristic and the second Synergistic effects of various environmental factors.

[0049] Step S41 is used to obtain the d-th order polynomial interaction terms. To capture higher-order nonlinear relationships, a polynomial kernel is introduced to map the features to a high-dimensional space. Definition The interaction term of the order polynomial is: ,(twenty one) in, For free parameters, The order is the polynomial. This invention preferably uses the following: After unfolding, we get: ,(twenty two) This expansion includes square terms, cross terms, and constant terms, and can simultaneously characterize the nonlinear effects of a single variable, the synergistic effects of two variables, and the benchmark bias.

[0050] Step S42 is used to dynamically adjust the interaction feature weights using an attention mechanism to obtain weighted semantic interaction features. Considering the significant differences in the contribution of different interaction items to energy consumption, an attention mechanism is introduced to dynamically adjust the interaction feature weights. The attention scoring function is defined as: ,(twenty three) in, for( , The importance of ) For the first A behavioral characteristic, For the first One environmental factor, For activation function, The number of behavioral characteristics, This refers to the number of environmental factors. , The projection matrix is ​​learnable. For attention weight vectors, For bias terms, This refers to the hidden layer dimension. In this example, the present invention preferably... .

[0051] The weighted semantic interaction features are: ,(twenty four) in, These are weighted semantic interaction features. This mechanism enables the model to automatically identify key interaction patterns and suppress interference from redundant features.

[0052] Step S43 is used to obtain semantic coupling features based on the tensor product representation, the d-th order polynomial interaction terms, and the weighted semantic interaction features. The final semantic coupling feature space can be represented as: (25) in, The basic second-order interaction relationship between production behavior characteristics and environmental factors can be described by formula (20). For further characterization of the nonlinear combination effect of the interaction features, see formula (22). To adaptively weight different semantic interaction features in order to highlight key interaction patterns that contribute significantly to energy consumption prediction, see formula (24).

[0053] Step S32 is used to construct structural coupling features. The structural coupling layer explicitly embeds domain knowledge and physical constraints into the feature space, constructing interpretable proxy variables and realizing a hybrid modeling paradigm from data-driven to mechanism-guided. In this embodiment, the specific method for constructing structural coupling features in step S32 can be of various forms known to those skilled in the art. In one example of the present invention, step S32 may include, for example... Figure 5 The steps shown are described in this. Figure 5 In this context, step S32 may include: In step S50, a thermodynamically constrained heating and cooling load model is constructed to obtain an air conditioning load that takes into account asymmetric characteristics. In step S51, a dehumidification load model of the humidity mass transfer process is constructed to obtain a dehumidification load that takes into account nonlinear effects. In step S52, a dynamic transmission model of the process section coupling is constructed to obtain the process section coupling strength; In step S53, the structural coupling characteristics are obtained based on the air conditioning load considering asymmetric characteristics, the dehumidification load considering nonlinear effects, and the section coupling strength.

[0054] In such Figure 5 In the method shown, step S50 is used for thermodynamically constrained cooling and heating load modeling. Specifically, in this example, based on the first law of thermodynamics, workshop energy consumption is closely related to the heat balance equation, and the cooling / heating load calculation is based on the principle of energy conservation, defined as: (26) in, for Total load of the air conditioning system at any given time The density of air (taken as 1.293 kg / m³ under standard conditions) is... For the workshop volume, For the isobaric specific heat capacity (taken as 1.005 kJ / (kg·K) for dry air), Indoor temperature, To set the temperature, The system energy efficiency ratio (COP) It represents the absolute value of the temperature difference and is used to uniformly characterize cooling and heating loads.

[0055] Considering that indoor temperature is difficult to obtain in real time in actual measurements, outdoor temperature is used as a proxy variable in this example, and a correction coefficient is introduced to compensate for the conduction delay effect of the indoor-outdoor temperature difference: (27) in, For the simplified proxy variables of heating and cooling loads, This is a function to indicate the start / stop status of the air conditioner (1 for on, 0 for off). Outdoor temperature As a benchmark for human comfort temperature, This is an empirical correction coefficient used to adjust the intensity of the impact of outdoor temperature on indoor load. It is obtained through least squares regression of historical data, and is preferably used in this invention. .

[0056] Furthermore, a piecewise linear function is introduced to distinguish the asymmetric energy efficiency characteristics of cooling and heating: (28) in, To account for asymmetrical air conditioning loads, The coefficient of performance (COP) reflects the marginal impact of a 1°C increase in temperature on cooling energy consumption. The heating efficiency coefficient reflects the marginal impact of a 1°C decrease in temperature on heating energy consumption. Both efficiency coefficients are obtained by fitting historical energy consumption data using the least squares method, accurately characterizing the differentiated response patterns of cooling and heating loads.

[0057] Step S51 is used to construct a dehumidification load model for the humidity mass transfer process to obtain a dehumidification load considering nonlinear effects. Specifically, based on the principles of mass transfer, the dehumidification load is closely related to the air humidity gradient and the latent heat of phase change. The dehumidification power calculation is based on the laws of mass conservation and energy conservation, and is defined as follows: (29) in, for Dehumidification power at any time The air mass flow rate (unit: kg / s) through the dehumidification system. Moisture content of imported air (unit: kg water / kg dry air) This refers to the moisture content of the outlet air. The latent heat of vaporization of water is approximately 2500 kJ / kg under standard conditions. Formula (29) shows that the energy consumption for dehumidification is proportional to the product of the amount of condensate and the latent heat of vaporization.

[0058] Using the approximate relationship between humidity and moisture content ,in The saturated vapor pressure (Pa) Atmospheric pressure (101325 Pa under standard conditions), For relative humidity (dimensionless), equation (29) can be rewritten in a simplified form based on measurable humidity: (30) in, As a proxy variable for dehumidification load, This represents the outdoor relative humidity (percentage). The comprehensive dehumidification coefficient, which incorporates the combined effects of airflow, enthalpy difference, system efficiency, and unit conversion, is calibrated through regression analysis using historical data.

[0059] To accurately capture the nonlinear growth characteristics of dehumidification load under high humidity conditions, an exponential correction term is introduced in this example to characterize the sharp response after the humidity exceeds a threshold: (31) in, To account for dehumidification load with nonlinear effects, The intensity coefficient of the exponential term controls the amplitude of the nonlinear response. The exponential decay coefficient controls the steepness of the nonlinear response. The humidity threshold is used. The index correction mechanism can accurately reflect the physical phenomenon that when the humidity exceeds 70%, the condensation dehumidification load rises sharply due to the rapid drop in dew point temperature, significantly improving the prediction accuracy under high humidity conditions.

[0060] Step S52 is used to construct a dynamic transmission model of process segment coupling to obtain the coupling strength between process segments. Based on the principles of material balance and energy conservation in a production line, a state-space model of cross-process segment coupling is established in this example to quantify the energy consumption transmission effect between upstream and downstream process segments. Let the output of the upstream process segment be... The output of this section is The coupling strength of a work section is defined as a linear combination of steady-state coupling terms and dynamic transmission terms: (6) in, for The coupling strength of the work section at any given time. This is a production level mapping function used to discretize continuous production output into equipment load levels. for The output of the upstream section at any given time. The output of this section at any given time. This represents the time-varying rate of upstream output, reflecting the intensity of fluctuations in the upstream feed rate. This is the time delay constant (unit: hours), reflecting the time delay effect of material transport from upstream to downstream. The coupling coefficients are determined using multivariate regression of historical data. In this example, the yield grade mapping function... In the system, 0 corresponds to production stoppage, 1 to low production, and 2 to high production.

[0061] The first term in formula (6) The steady-state coupling effect is characterized by the product of upstream material supply and downstream load level, reflecting the continuous driving effect of upstream output on downstream energy consumption under stable production conditions; the second term in formula (6) Characterizing the dynamic transmission effect, when the upstream output changes rapidly, the downstream needs to adjust the equipment operation status in advance or in advance to adapt to the fluctuation of material supply, thereby generating additional start-up and shutdown energy consumption, preheating energy consumption or idling energy consumption.

[0062] Furthermore, in this example, a feedback regulation mechanism is introduced to describe the reverse impact of the downstream on the upstream: (7) in, For the upstream output at the next moment, The feedback gain coefficient (dimensionless) controls the degree to which downstream deviation affects the upstream adjustment speed. A target output value is set. The feedback loop creates a closed-loop regulation mechanism for energy consumption between work sections. When downstream output is lower than the target value, upstream automatically reduces material supply to avoid material accumulation; when downstream output is higher than the target value, upstream increases material supply to meet production needs, thereby improving the overall stability and energy efficiency of the system.

[0063] Step S53 is used to obtain structural coupling characteristics based on the air conditioning load considering asymmetric characteristics, the dehumidification load considering nonlinear effects, and the section coupling strength. The final structural coupling feature vector is expressed as: (32) in, For structural coupling feature vectors, As a thermodynamically constrained proxy variable for cooling and heating loads, it reflects the driving effect of temperature on air conditioning energy consumption. As a proxy variable for dehumidification load based on the mass transfer process, it reflects the nonlinear impact of humidity on dehumidification energy consumption. The section coupling strength, based on a dynamic transmission model, reflects the transmission effect of upstream and downstream production fluctuations on energy consumption. The feature space integrates three physical mechanisms: thermodynamic constraints, mass transfer processes, and dynamic coupling. This explicitly embeds domain knowledge into the data model, significantly improving the interpretability and generalization ability of the predictions.

[0064] Step S33 is used to fuse the temporal coupling features obtained in step S30, the semantic coupling features obtained in step S31, and the structural coupling features obtained in step S32 to obtain three-dimensional coupling features of temporal-semantic-structure.

[0065] Step S12 is used to divide the multi-source energy consumption dataset according to the operating condition identifier to obtain datasets for different operating conditions. In this embodiment, the operating condition division is based on the production status, with "whether production is in progress" as the primary criterion. The operating condition determination function is defined as follows: (33) in, for Operating status indicators at any given time. for Output at any given time. When output is greater than zero, it is considered a production condition; when output is equal to zero, it is considered a non-production condition. This classification method can effectively distinguish between two basic modes of energy consumption: during production, energy consumption is mainly driven by equipment operation and output load, while during non-production, energy consumption is mainly composed of basic loads such as air conditioning, lighting, and standby systems.

[0066] Under production conditions, output and equipment operating status are the dominant factors influencing energy consumption changes, while environmental factors play a secondary moderating role. Under non-production conditions, however, output-related characteristics become ineffective, and energy consumption is primarily affected by environmental conditions such as temperature and humidity, as well as the operation of the basic system. Therefore, the difference in energy consumption driver weights between the two types of conditions can be expressed as follows: (34) in, The weighting percentage of output characteristics. The weighting of environmental features is used. This work condition division based on production status simplifies model complexity, avoids the problem of sample sparsity under multiple combinations, and can significantly improve the stability and generalization ability of the model.

[0067] Based on the determination of whether production occurred, all samples were divided into two subsets: production and non-production. Feature engineering and model training were then performed on each subset separately. An adaptive strategy was adopted for the test set proportion. (35) in, The proportion of the test set. This represents the total number of samples for this operating condition. This strategy ensures that when the sample size is sufficient, the most recent 30% of the data is selected as the test set; if the sample size is small, the proportion of the test set is dynamically adjusted to ensure the representativeness of the validation samples. The sample division follows the principle of temporal order to maintain temporal causality. Finally, the sample size and energy consumption statistics (mean, standard deviation, maximum value, minimum value) for the two types of operating conditions are recorded respectively, providing a basis for model evaluation and diagnosis.

[0068] Step S13 involves training multiple models based on the temporal-semantic-structural three-dimensional coupled features corresponding to datasets under different operating conditions to obtain the optimal model for each condition. For energy consumption prediction tasks characterized by small to medium sample sizes, high-dimensional features, nonlinear relationships, and temporal dependencies, this example may involve comparing four mainstream algorithms: XGBoost (Extreme Gradient Boosting), CatBoost (Class Boosting), RandomForest, and PyTorch Neural Networks. The model can be uniformly represented as: (36) in, To predict energy consumption, For the input feature vector, For model mapping function, These are the model parameters.

[0069] Furthermore, to ensure optimal performance of each model under different operating conditions, in this example, the optimal hyperparameters are found using a grid search method: (37) in, For the optimal combination of hyperparameters, The loss function is used for the validation set. Key parameters include the number of trees, depth, and learning rate of the tree model, as well as the layer structure and learning rate of the neural network.

[0070] After training multiple models, the models need to be evaluated to obtain the optimal model for different operating conditions. In this example, an evaluation mechanism combining cross-validation and independent validation sets is used. Finally, the coefficient of determination is selected. The highest-performing model is considered the optimal model. (38) in, For actual energy consumption, To predict energy consumption, This represents the average energy consumption. For neural network models, an early stopping mechanism is introduced to avoid overfitting.

[0071] After training is completed, the optimal model for each operating condition will be persistently saved for subsequent energy consumption prediction and model updates.

[0072] Step S14 is used to predict the factory energy consumption under the corresponding operating conditions using the optimal model for different operating conditions.

[0073] On the other hand, this invention also provides a factory energy consumption prediction system based on production behavior recognition. The system includes a data acquisition module, a data preprocessing module, a feature engineering module, a working condition segmentation module, a model training and evaluation module, a model deployment and prediction module, and a processor. The data acquisition module collects multi-source energy consumption data, including energy consumption data, production data, and environmental data from the enterprise's production system, energy management system, and environmental information sources. The data preprocessing module preprocesses the multi-source energy consumption data, including cleaning, format standardization, anomaly handling, and time alignment of the collected data to form standardized input data. The feature engineering module identifies production behavior based on the preprocessed data and constructs temporal features, semantic interaction features, and structural coupling features. The working condition segmentation module segments the samples according to the production behavior recognition results, forming data subsets under different operating conditions. The model training and evaluation module trains and optimizes the parameters of the energy consumption prediction model under different working conditions and evaluates the performance of different prediction models, selecting the target model that meets the prediction requirements. The model deployment and prediction module loads the target model and performs energy consumption prediction on real-time or historical data. A processor configured to perform any of the methods described above.

[0074] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the methods described above.

[0075] An embodiment of the present invention is as follows: like Figure 6 As shown, this invention provides a factory energy consumption prediction system based on the coupling of production behavior recognition and environmental factors. The system comprises a data acquisition module, a data preprocessing module, a feature engineering module, a working condition classification module, a model training module, a model evaluation module, a result output module, a database interaction module, a model deployment and monitoring module, a user interaction and configuration module, a real-time historical / relational database, a data analysis server, a field operation terminal, and an industrial Ethernet network.

[0076] This embodiment selects a specific process segment of a manufacturing enterprise for verification. The data collection period is 180 consecutive days, and after data cleaning, 168 days of valid samples are retained. The data acquisition module collects multi-source data from the enterprise's MES system, energy management system, and meteorological data interface via industrial Ethernet. This includes energy consumption data (cooling energy converted to standard coal equivalent, heating network steam converted to standard coal equivalent, air conditioning electricity converted to standard coal equivalent, process electricity converted to standard coal equivalent, production lighting electricity converted to standard coal equivalent, and total energy consumption converted to standard coal equivalent), production data (output and shift information), environmental data (predicted average temperature, predicted maximum temperature, predicted minimum temperature, predicted average humidity, predicted maximum humidity, and predicted minimum humidity), and process segment coupling data (output of adjacent process segments). The collection frequency is once a day, and the data is summarized by date.

[0077] The data preprocessing module executes a three-layer cleaning mechanism. The first layer utilizes statistical detection... The first layer identifies outliers. For example, if the total energy consumption (converted to standard coal equivalent) on a certain day is 78.5 tons, while the historical average is 46.3 tons and the standard deviation is 8.2 tons, since 78.5 > 46.3 + 3 × 8.2 = 70.9, this sample is marked as a suspected outlier. The second layer filters contradictory samples. For example, if the output on a certain day is zero but the total energy consumption (converted to standard coal equivalent) is 52.3 tons, this significantly deviates from the normal level for non-production conditions (usually 10-15 tons). The third layer uses manual verification combined with an outlier date database. For example, if production was suspended on June 15th due to equipment maintenance, although the energy consumption data is complete, the sample for that day is excluded. For missing values, such as a missing predicted average temperature for a certain day, interpolation using the mean of the two days before and after is used. After three layers of cleaning, 12 days of abnormal samples were removed from the 180 days of raw data, and 168 days of valid samples were retained for use in the feature engineering module.

[0078] The feature engineering module first identifies production behavior. For a sample of a certain day, according to Formula 31: the output is 850 units, the electricity consumption for production lighting (equivalent to 1.8 tons of standard coal) is 1.8 tons (exceeding the threshold of 1.2 tons), and the electricity consumption for air conditioning (equivalent to 8.5 tons of standard coal) is 8.5 tons (exceeding the threshold of 2.0 tons). Then, a behavior feature vector is generated: whether production is underway = 1 (because output > 0), output level = 2 (because 850 exceeds the high-output threshold of 720), whether production lighting electricity is used = 1, and whether air conditioning is on = 1. These four types of behavior identifiers, together with the output itself, constitute the five-dimensional basic behavior features, serving as the basis for subsequent coupled feature construction.

[0079] Secondly, construct temporal coupling features. For example... Figure 7As shown, time-series coupling captures the historical dependence and dynamic variation patterns of energy consumption through lag operators (Equations 8 and 2-3), rolling statistical operators (Equations 9-12), difference operators (Equations 13-14), momentum operators (Equation 15), and seasonal decomposition operators (Equations 16-18). Taking total energy consumption converted to standard coal equivalent as an example, assuming the current time t, the value of t is 46.5 tons, and the historical values ​​for the previous 14 days are as follows: [45.2,47.1,44.8,46.3,48.5,43.7,44.9,46.8,45.5,47.3,44.2,46.0,48.1,45.8] tons. Historical values ​​are directly extracted using lag characteristics. Tons (the previous day) Tons (same period last week). Tons (first two weeks).

[0080] The weighted lag feature uses exponentially decaying weights with a decay coefficient λ = 0.1, and the weight sum is calculated as follows: The weight of the most recent day is ; ton.

[0081] Rolling statistical features are used to calculate the mean and standard deviation over the most recent 7 days: ton; ton; Exponentially weighted moving averages use a smoothing coefficient. Current value tons (assuming the previous day) (45.91 tons).

[0082] Calculate the difference between adjacent days using differential features: ton, ton.

[0083] Momentum indicators use windows and attenuation factor The calculation is as follows: (Assume the momentum of the previous day was 0.18).

[0084] Seasonal decomposition uses the STL algorithm to decompose the original sequence into a trend term, a seasonal term, and a residual term. For example, after decomposing 90 consecutive days of data, the trend term for a certain day is T=46.8 tons (reflecting a long-term upward trend), the seasonal term is S=1.5 tons (reflecting a weekly cyclical pattern where Monday to Friday are higher than the weekend), and the residual term... Tons (reflecting random fluctuations on the same day).

[0085] These 15 time-series features significantly enhance the model's ability to perceive historical dependence and dynamic changes in energy consumption.

[0086] Semantic coupling features construct interaction terms between behavior and environment through tensor products, multinomial kernel functions, and attention mechanisms. For example... Figure 7 As shown, for the same sample (yield 850 units, predicted average temperature 32℃, predicted average humidity 75%), the tensor product interaction first constructs a second-order interaction term: "yield × predicted average temperature" = 850 × 32 = 27200. , Polynomial kernel function expansion generates higher-order terms: Product squared = 850 2 =722500, predicted average temperature squared =32 2 =1024, The attention mechanism dynamically weights these interactions. Assuming a projection dimension of 8, the query vector Q and key vector K are calculated using the projection matrix obtained during training, and then the attention score matrix A is calculated. For high-temperature, high-yield samples, the attention mechanism using Equation 18 automatically increases the weights of key interactions: the attention weight for "yield × predicted average temperature" is 0.18, for "yield × predicted average humidity" it is 0.15, and for "yield² × predicted average temperature" it is 0.10, with the first five weights accounting for a cumulative 64%. For low-temperature, low-yield samples (yield 320 units, predicted average temperature 15℃), the attention weights for the same interactions decrease to 0.08, 0.06, and 0.04, respectively, indicating that the model can adaptively adjust feature weights according to sample characteristics. These semantically coupled features total more than 30 dimensions, effectively revealing the impact of multi-factor synergy on energy consumption.

[0087] The structural coupling characteristics are based on physical mechanisms to construct three types of proxy variables. The cooling and heating load proxy is calculated using a piecewise linear function formula (21), with the comfort temperature benchmark set at 22℃ and the cooling energy efficiency coefficient... Heating energy efficiency coefficient For the summer sample (yield 850 units, predicted average temperature 32℃), since 32 > 22, cooling mode is activated. For the winter sample (yield 680 units, predicted average temperature 8℃), since 8 < 22, heating mode is activated. The dehumidification load agent uses Formula 26 to introduce an exponential correction term to characterize the nonlinear characteristics of high humidity environments. The comfort humidity benchmark is set at 60%, and the dehumidification energy efficiency coefficient is... Exponential correction coefficient For the high humidity sample (yield 850 units, predicted average humidity 80%)... When humidity increases from 70% to 80%, the dehumidification load agent increases from 33.7 to 64.26, a growth rate of 90.7%, significantly higher than the linear growth rate of humidity (14.3%). The section coupling agent includes both steady-state and dynamic components, calculated using formulas 27 and 28. The steady-state coupling coefficient... Dynamic transmission coefficient Assuming the upstream process output is 1200 units, compared to 1150 units the previous day, and the downstream load level is 2 (high load), then: The dynamic transmission term contributes 85.7%, indicating a significant dynamic impact of upstream output fluctuations on downstream energy consumption. These three types of proxy variables, totaling three dimensions, explicitly link production processes with energy consumption responses, enhancing the physical interpretability of the model.

[0088] After processing by the feature engineering module, the original 10+ dimensional features were expanded to approximately 50 dimensional coupled features. The operating condition segmentation module divided the 168 days of samples into 120 samples for production conditions and 48 samples for non-production conditions based on the "whether in production" label. For production conditions, the samples were divided into a training set of 96 samples (80%) and a test set of 24 samples (20%) in chronological order. For non-production conditions, due to the smaller sample size, the proportion of the test set was increased to 30%, i.e., 34 samples for the training set and 14 samples for the test set, to ensure the representativeness of the validation samples.

[0089] The model training module trains four algorithms for each working condition. Taking the XGBoost model for the production working condition as an example, the hyperparameter search space includes the number of trees {100, 200, 300}, maximum tree depth {3, 5, 7}, and learning rate {0.01, 0.05, 0.1}. Each parameter combination is evaluated using 3-fold cross-validation. For example, the parameter combination (number of trees = 200, maximum depth = 5, learning rate = 0.05) has a negative mean squared error of -4.21 on the training set, which is better than other combinations and is selected as the optimal parameters. After retraining on the full training set, the model performs as follows on the test set: coefficient of determination R² = 0.92, mean absolute percentage error (MAPE) = 8.3%, root mean square error (RMSE) = 2.14 tons of standard coal equivalent, mean absolute error (MAE) = 1.68 tons of standard coal equivalent, and training time is 9.2 minutes. Compared to the baseline model before the improvement (R²=0.62, MAPE=23.5%), the coefficient of determination is improved by approximately 48%, the mean absolute percentage error is reduced by approximately 65%, and the prediction accuracy far exceeds the industrial application standard (the industry standard requires MAPE<20%). The neural network model adopts a hidden layer structure [128,64,32] with a learning rate of 0.005. After 150 iterations on the training set, the validation set loss decreased from the initial 28.5 to 3.8. Training was terminated at the 150th iteration through an early stopping mechanism to avoid overfitting. The model achieved R²=0.89 and MAPE=9.7% on the test set, slightly lower than XGBoost. The test set R² values ​​for CatBoost and RandomForest models were 0.88 and 0.85, respectively, both lower than XGBoost. Therefore, the system automatically selected XGBoost as the optimal model for production conditions and persisted it.

[0090] For non-production conditions, the neural network model performed optimally, with R²=0.73 and MAPE=9.8% on the test set. Compared to the baseline model before the improvement (R²=0.35, MAPE=85.3%), the coefficient of determination increased by approximately 109%, and the mean absolute percentage error decreased by approximately 88%. Feature importance analysis showed that the dominant features in this condition were significantly different from those in the production condition: the Top-5 features were "whether production lighting electricity was used" (weight 11.5%), "predicted average temperature_lag1" (weight 9.8%), "predicted average temperature" (weight 8.7%), "cooling load proxy" (weight 7.9%), and "predicted average humidity" (weight 7.3%), while the weight of production-related features approached zero. Statistical analysis indicated that the weight of production features was significantly lower in the production condition. Environmental feature weighting The weighting of environmental characteristics under non-production conditions. The weights of production features approach 0. The Euclidean distance between the feature weight vectors of the two types of operating conditions is 0.73, confirming the necessity of operating condition separation modeling: the dominant driving factors of energy consumption under different operating conditions are fundamentally different. Through operating condition grouping modeling, the MAPE of the test set decreased by an average of 2.7 percentage points, and R² was significantly improved (from 0.83 to 0.92 for production conditions and from 0.59 to 0.73 for non-production conditions), with an energy saving effect of about 15% compared to the global model.

[0091] The model evaluation module makes predictions and calculates errors for each sample in the test set. Taking a sample from a specific production day as an example, the actual total energy consumption was 48.3 tons, while the XGBoost model predicted 46.9 tons, with an absolute error of [missing value]. tons, with a relative error of After analyzing all test set samples, 92.5% of the samples under production conditions had a relative error within 10%, with only two samples having a relative error exceeding 15% (17.2% and 19.8%, respectively). The system automatically filtered out these two high-error samples and correlated them with detailed energy consumption data: The first sample (July 8th) had an actual total energy consumption of 52.1 tons and a predicted total of 44.6 tons. A review of the detailed energy consumption data revealed that air conditioning electricity consumption, converted to standard coal equivalent, was 15.2 tons (significantly higher than the normal value of 8-10 tons). The reason for this was that the weather forecast temperature was 28℃, but the actual temperature reached 34℃. The second sample (September 15th) had an actual total energy consumption of 39.2 tons and a predicted total of 47.1 tons. The detailed energy consumption data showed that process electricity consumption was only 18.5 tons (normal is 26-28 tons). The reason for this was that the production line was temporarily shut down for 2 hours that day for equipment debugging. This diagnostic information provided a basis for model optimization and production management decisions.

[0092] Feature importance analysis extracted the gain values ​​of the XGBoost model. The top-5 features for production conditions were: output (weight 9.2%), predicted average temperature (weight 8.1%), the interaction term "output × predicted average temperature" (weight 7.8%), output_lag1 (weight 8.5%), and predicted average temperature_lag7 (weight 5.3%), with a cumulative weight of 38.9%. This indicates that under production conditions, output is the most significant energy consumption driver (directly contributing 9.2%). The time-series feature output_lag1 entering the top-3 (weight 8.5%) reflects historical dependence. Ambient temperature has a significant impact on energy consumption through its direct effect (8.1%) and its interaction with output (7.8%). The cooling load proxy, as a structural feature based on physical mechanisms, jumped to the top-6 (weight 7.1%) under high-temperature summer conditions, validating the effectiveness of structural coupling features. For the neural network model, predicted average temperature_lag1 ranked second in feature importance (weight 7.2%), further validating the crucial role of time-series coupling features. Under high temperature and humidity conditions, the weight of the "output × predicted average temperature" interaction term increased to 6.3%, and the attention mechanism increased the weight concentration of key interaction modes to 45% of the top 10 terms. When upstream output fluctuates by more than 15%, the weight of the process coupling agent increases from the usual 7.1% to 8.2%, indicating an enhanced dynamic transmission effect. Comparatively, without constructing three-dimensional coupling features, the baseline model trained using only the original 10-dimensional features has an R² of only 0.62 and a MAPE as high as 23.5%, while the three-dimensional coupling features expand the input feature dimensions to more than 50 dimensions, increasing the information density by more than 5 times, verifying the crucial role of feature engineering.

[0093] The results output module generates an Excel report and visualization charts. The model results summary table summarizes the optimal model name and performance indicators for each operating condition: XGBoost (R²=0.92, MAPE=8.3%, hyperparameters: 200 trees, depth 5, learning rate 0.05) is the optimal model for production conditions, and Neural Network (R²=0.73, MAPE=9.8%, hyperparameters: hidden layers [128,64,32], learning rate 0.005) is the optimal model for non-production conditions. The performance comparison table displays detailed indicators for the four algorithms; for example, under production conditions, XGBoost's R² is 0.04 higher than CatBoost, 0.07 higher than RandomForest, and 0.03 higher than Neural Network. The error details table lists the date, predicted value, actual value, error, and detailed energy consumption items (calculated per ton of standard coal for each of the following: cooling capacity, heating network steam, air conditioning electricity, process electricity, and production lighting electricity). The visualization charts include the training loss curve (showing that the loss decreased from 28.5 to 3.8 during the training of the neural network), the prediction comparison scatter plot (showing that the predicted value and the actual value are highly fitted to the diagonal), the feature importance bar chart (showing the weight distribution of the top-20 features), and the error distribution histogram (showing that more than 90% of the sample errors are concentrated in the range of ±5 tons of standard coal).

[0094] The model deployment and monitoring module supports online prediction and performance tracking. After the data acquisition module obtains data for a new day, the system automatically determines the operating condition (based on whether the output is zero) and calls the corresponding optimal model for prediction. For example, if the daily output is 780 units, it is determined to be a production condition. The XGBoost model is loaded, and a 50-dimensional feature vector (including behavioral features, temporal features, semantic interaction features, and structural proxy features) is input, outputting a predicted energy consumption of 45.2 tons. The actual energy consumption that day was 44.8 tons, with a relative error of 0.9%, which is within the normal range. The model monitoring submodule calculates the rolling MAPE for the past 7 days as 8.7% and the cumulative MAPE for the past 30 days as 9.1%, both below the alarm threshold of 15%, indicating stable model performance. If the relative error suddenly rises to 18.5% on a certain day, an alarm mechanism is triggered. The system notifies maintenance personnel via email and WeChat: "The energy consumption prediction error on October 15, 2024, is 18.5%, exceeding the threshold of 15%. It is recommended to investigate equipment status or process changes."

[0095] Based on the energy consumption drivers identified by the model, the company formulated targeted energy-saving measures. The feature importance indicator shows that the interaction term "output × predicted average temperature" has a weight of 7.8%, indicating a surge in energy consumption under high-temperature, high-output conditions. Accordingly, the company adjusted its production plan: during the high-temperature period in summer (predicted average temperature > 30℃), some high-energy-consuming production tasks were shifted to the early morning or nighttime hours when temperatures are lower. Actual measurements show an average energy consumption reduction of approximately 8% in July and August. The cooling load proxy weight of 7.1% indicates that the air conditioning system is a significant energy source. The company implemented a pre-cooling strategy and optimized the air conditioning start-stop strategy: starting the air conditioning two hours in advance for pre-cooling, utilizing the low-temperature, low-electricity-price period at night for cold storage, and appropriately increasing the set temperature during the high-temperature period during the day. Actual energy consumption was 12% lower than the predicted baseline, demonstrating a significant reduction in air conditioning energy consumption. "Whether production lighting is used" has a weight of 11.5% under non-production conditions. The company implemented zoned lighting control: turning off lighting in some areas during low-output periods and keeping only duty lighting on rest days, resulting in a significant reduction in lighting energy consumption. After implementing the above measures, the average annual energy consumption of this process section decreased by 10.3%, saving approximately 1.8 million yuan in energy costs annually, and shortening the investment payback period from the expected 3 years to 2.1 years. High-precision forecasting allows enterprises to accurately estimate energy consumption needs 2-3 days in advance, saving large enterprises with an annual energy consumption of 100,000 tons of standard coal equivalent approximately 5-10 million yuan. The model triggers alarms when energy consumption anomalies are identified, quickly locating equipment faults or process deviations, reducing anomaly losses by approximately 8% annually, and shortening fault response time from 4-6 hours to within 1 hour.

[0096] The system displays modeling results and predictive analysis in real time through on-site operation terminals. The terminal interface displays the optimal model, performance indicators, feature importance ranking, high-error sample list, and energy-saving measure suggestions for each operating condition. Energy management personnel can adjust parameters (such as modifying the abnormal date database, adjusting the hyperparameter search space, and updating operating condition classification rules) through user interaction and configuration modules, triggering model retraining. Non-technical personnel can operate independently after 2 hours of training. The system supports RESTful API interfaces with a response time of less than 100ms, facilitating seamless integration with enterprise MES, ERP, and energy management platforms, enabling online energy consumption prediction without modifying the existing system architecture. The system automatically generates multi-dimensional Excel reports and visualization charts, with report generation time less than 30 seconds, directly usable for management decision-making. The data analysis server uses a task queue mechanism to schedule parallel modeling tasks across multiple process stages, avoiding resource conflicts. The system is compatible with both CPU and GPU computing modes. Small and medium-sized enterprises can quickly deploy on ordinary servers, with tree model training time less than 15 minutes; large enterprises can accelerate neural network training using GPUs, reducing training time to 6-10 minutes. The entire process, from data loading to model saving, takes approximately 25 minutes (3 minutes for data loading, 5 minutes for preprocessing, 7 minutes for feature generation, approximately 10 minutes for model training, and approximately 1 minute for evaluation). Compared to traditional manual modeling methods (which take 2-3 weeks), this reduces the time to 3-5 days, improving efficiency by approximately 80%. The system has a built-in model monitoring mechanism that automatically triggers an alarm when the rolling MAPE exceeds 15% for three consecutive days or the cumulative MAPE exceeds 18%, suggesting retraining or parameter adjustment to ensure long-term stable operation. The system boasts high operational stability, with an average annual failure rate of less than 1%, meeting the requirements for continuous 24 / 7 operation.

[0097] The beneficial effects of this invention are: This invention proposes a three-dimensional coupled feature system of temporal-semantic-structural aspects. Compared to existing methods that rely solely on original features or linear combinations, this system achieves collaborative modeling of production behavior, environmental factors, and process structures, significantly enhancing the expressive power and interpretability of energy consumption features. Through multi-scale temporal operators, behavior-environment interaction features, and structural proxy variables based on physical mechanisms, the feature space is expanded from approximately 10 dimensions to over 50 dimensions, resulting in a significant increase in information density.

[0098] The embodiments of this invention achieve high prediction accuracy under various operating conditions. Under production conditions, the prediction determination coefficient increases from 0.62 in the baseline model to 0.92, and the mean absolute percentage error decreases from 23.5% to 8.3%. Under non-production conditions, the prediction accuracy is also significantly improved. The overall prediction error is stably controlled within 10%, meeting the application requirements for energy consumption prediction accuracy in industrial scenarios.

[0099] This invention divides data into operating conditions based on production status and models different operating conditions separately, effectively avoiding the performance fluctuation problem of a single model under multiple operating conditions. Practical results show that operating condition separation modeling can significantly improve the stability of the model under operating condition switching and output fluctuation conditions, providing a reliable basis for formulating differentiated energy-saving strategies.

[0100] The embodiments of this invention construct a complete technical process covering data acquisition, feature engineering, model training and prediction output, support standardized interfaces and automated operation, have low deployment costs and stable operation, and can meet the engineering application needs of industrial enterprises for online prediction and long-term operation.

[0101] In summary, this invention, through production behavior recognition and three-dimensional coupled modeling, achieves high-precision and stable prediction of factory energy consumption while ensuring engineering feasibility, and has high practical value and promotion significance.

[0102] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0103] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0104] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0106] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0107] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0108] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0109] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0110] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A factory energy consumption prediction method based on production behavior recognition, characterized in that, The prediction method includes: Obtain a multi-source energy consumption dataset; Construct a three-dimensional coupled feature system of time-series, semantics, and structure based on the aforementioned multi-source energy consumption dataset; The energy consumption multi-source dataset is divided according to the operating condition identifier to obtain datasets for different operating conditions; Based on the time-semantic-structural three-dimensional coupled features corresponding to datasets under different working conditions, multiple models are trained to obtain the optimal model for each working condition. The optimal model for different operating conditions is used to predict the factory energy consumption under the corresponding operating conditions.

2. The prediction method according to claim 1, characterized in that, The construction of time-series-semantic-structural three-dimensional coupled features based on the aforementioned multi-source energy consumption dataset includes: By using multi-operator composite mapping, the multi-source energy consumption data is extended into a coupled feature vector containing historical states, which serves as a temporal coupled feature. Construct a high-order interactive feature tensor to map the discrete feature space into a continuous semantic manifold, which serves as a semantic coupling feature; Domain knowledge and physical constraints are embedded into the feature space to construct interpretable proxy variables as structural coupling features; The temporal coupling features, semantic coupling features, and structural coupling features are fused to obtain a three-dimensional coupling feature of temporal-semantic-structural coupling.

3. The prediction method according to claim 2, characterized in that, By using multi-operator composite mapping, energy consumption multi-source data is extended into coupled feature vectors containing historical states, which serve as temporal coupled features, including: Construct a temporal coupling model based on formula (1). ,(1) in, For the temporal coupling feature space, The original feature sequence, For lag operators, For rolling statistics calculation, It is a difference operator.

4. The prediction method according to claim 3, characterized in that, The hysteresis operator includes The first-order lag characteristic and the weighted lag characteristic are obtained according to formulas (2) and (3): ,(2) ,(3) in, For weighted lag characteristics, for The original feature values ​​at time 1. For the first The normalized weight coefficients at each time step. For the length of the history window, It is an exponentially decaying kernel function. This is the attenuation coefficient.

5. The prediction method according to claim 2, characterized in that, Construct a high-order interaction feature tensor to map the discrete feature space into a continuous semantic manifold, which serves as semantic coupling features, including: Based on the set of production behavior characteristics and the set of environmental factor characteristics, a tensor product representation of basic interaction characteristics is constructed. A polynomial kernel function is used to map the features in the tensor product representation to a high-dimensional space to obtain d-order polynomial interaction terms. An attention mechanism is used to dynamically adjust the weights of interaction features in order to obtain weighted semantic interaction features. Semantic coupling features are obtained based on the tensor product representation, the d-order polynomial interaction terms, and the weighted semantic interaction features.

6. The prediction method according to claim 5, characterized in that, An attention mechanism is used to dynamically adjust the weights of interaction features to obtain weighted semantic interaction features, including: The attention score is obtained according to formula (4). ,(4) in, for( , The importance of ) For the first A behavioral characteristic, For the first One environmental factor, and The projection matrix is ​​learnable. For attention weight vectors, For bias terms, For activation function, The number of behavioral characteristics, The number of environmental factors; According to formula (5), the weighted semantic interaction features are obtained. ,(5) in, These are the weighted semantic interaction features.

7. The prediction method according to claim 2, characterized in that, Domain knowledge and physical constraints are embedded into the feature space to construct interpretable proxy variables, which serve as structural coupling features, including: A thermodynamically constrained cooling and heating load model is constructed to obtain an air conditioning load that takes into account asymmetric characteristics. A dehumidification load model for the humidity mass transfer process is constructed to obtain a dehumidification load that takes into account nonlinear effects. Construct a dynamic transmission model of process segment coupling to obtain the process segment coupling strength; Based on the air conditioning load considering asymmetric characteristics, the dehumidification load considering nonlinear effects, and the section coupling strength, the structural coupling characteristics are obtained.

8. The prediction method according to claim 1, characterized in that, Constructing a dynamic transmission model of process segment coupling to obtain the process segment coupling strength includes: The coupling strength of the work section is obtained according to formula (6). ,(6) in, for The coupling strength of the work section at any given time. This is a production level mapping function. for The output of the upstream section at any given time. The output of this section at any given time. and The coupling coefficient is... It is the time delay constant; A feedback adjustment mechanism is introduced based on formula (7). ,(7) in, for The output of the upstream section at any given time. For feedback gain coefficient, Set a value for the target output.

9. A factory energy consumption prediction system based on production behavior recognition, characterized in that, The system includes: The data acquisition module is used to collect energy consumption data from multiple sources. The data preprocessing module is used to preprocess the energy consumption multi-source data; The feature engineering module is used to identify production behaviors based on preprocessed data and to construct temporal features, semantic interaction features, and structural coupling features. The working condition segmentation module is used to segment the samples according to the production behavior identification results, forming data subsets under different operating conditions; The model training and evaluation module is used to train and optimize the parameters of the energy consumption prediction model under different operating conditions and to evaluate the performance of different prediction models. The model deployment and prediction module is used to load the target model and predict energy consumption based on real-time or historical data. The processor is connected to the data acquisition module, data preprocessing module, feature engineering module, working condition segmentation module, model training and evaluation module, and model deployment and prediction module, and the processor is configured to perform the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 8.