Industrial added value growth rate prediction method, system, electronic device and storage medium

By using Stacking integrated algorithm in industrial value-added prediction combined with multiple machine learning algorithms, an industrial value-added growth rate prediction model is built, which solves the problems of prediction inaccurate prediction and data lag in the existing technology, and achieves higher accuracy and accuracy prediction.

CN116805182BActive Publication Date: 2025-05-13ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID NINGXIA ELECTRIC POWER COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310849641.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-11
Publication Date
2025-05-13
Estimated Expiration
2043-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to utilize a large number of explanatory variables in industrial value-added predictions, resulting in inaccurate predictions and the lag of traditional statistics cannot reflect economic activities in real time.

Method used

Using a method based on Stacking integrated algorithm, combining linear regression algorithm, decision tree algorithm, support vector machine, k-nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm, an industrial added value growth prediction model is built, and power data and economic data are used to predict.

Benefits of technology

By combining multiple machine learning algorithms, the sensitivity of a single algorithm is reduced, the accuracy and accuracy of industrial added value growth prediction can be improved, and real-time information of power data and economic data can be better utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116805182B_ABST
    Figure CN116805182B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, electronic device and storage medium for predicting the growth rate of industrial added value, which belongs to the technical field of industrial added value growth rate prediction. The method includes: obtaining sample data, the sample data includes annual parameter data and annual industrial added value, the annual parameter data includes annual economic data and annual power data; preprocessing each of the sample data and inputting it into a sample set; based on the sample set, based on the Stacking algorithm, combined with the linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm, the industrial added value growth rate prediction modeling is performed to obtain a prediction model; the industrial added value growth rate is predicted using the prediction model; and the prediction model is evaluated using volatility and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial value-added growth rate prediction, and in particular to an industrial value-added growth rate prediction method, system, electronic equipment and storage medium. Background Art

[0002] At present, there are relatively few studies on industrial value-added forecasting, and they are mainly based on a number of selected traditional statistical indicators, such as using the total amount of social consumer snacks and broad money data to forecast industrial value-added, and forecasting industrial value-added based on the constructed private enterprise credit spread index. In the case of increasingly complex macroeconomic transmission mechanisms and more diverse factors affecting industrial value-added, the appropriate forecasting model should contain a large number of explanatory variables to make full use of valuable information. It is difficult to obtain accurate forecasts using only a few variables. In addition, the lag of traditional statistical data makes it difficult to use current information on economic activities when forecasting, while power big data is closely related to economic activities and available in real time, which can provide current information on economic activities to make up for the shortcomings of traditional statistical data. Therefore, the comprehensive use of power data and traditional statistical data may obtain more accurate industrial value-added forecast results. The power industry is a basic energy industry of the national economy and plays a vital supporting role in the development of other industrial sectors. Electricity consumption is one of the important inputs in the production process, and there is no inventory phenomenon itself. It can also be seen that there must be a certain correlation between electricity consumption and industrial output value.

[0003] With the development of short-term power load forecasting technology, the factors affecting short-term power load are considered more comprehensively, and the relationship between the factors and the load is not a simple linear relationship, which makes the traditional and classical forecasting methods show great disadvantages, and the processing of large sample data is also a huge challenge for traditional and classical forecasting methods. Some machine learning algorithms have shown excellent performance with their strong learning and adaptive capabilities. The essence of applying machine learning algorithms to load forecasting is to first assume a model, and then learn to solve the model parameters that minimize the loss function. Commonly used machine learning algorithms include artificial neural network method, support vector machine method, random forest, gradient boosting decision tree (GBDT), ridge regression, etc. These methods have significant performance in improving the accuracy of power load forecasting. However, the above are all single load forecasting methods, and the data distribution characteristics will affect the algorithm results. This is caused by the defects of the single algorithm itself. Therefore, for the same set of test data, the calculation results will be quite different, and the prediction accuracy is not high. Summary of the invention

[0004] The technical solution adopted by the embodiment of the present invention to solve the technical problem is:

[0005] A method for predicting the growth rate of industrial added value, comprising:

[0006] Step S1, obtaining sample data, wherein the sample data includes annual parameter data and annual industrial added value, and the annual parameter data includes annual economic data and annual power data;

[0007] Step S2, pre-processing each of the sample data and inputting them into a sample set;

[0008] Step S3, based on the sample set, based on the Stacking algorithm, combined with the linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm, perform industrial added value growth rate forecast modeling to obtain a forecast model;

[0009] Step S4, using the prediction model to predict the growth rate of industrial added value;

[0010] Step S5, evaluating the prediction model using volatility and accuracy.

[0011] Preferably, the sample data is recorded as (x i ,y i ), where x i is the annual parameter data, y i is the industrial added value of the year, i is the year, and the characteristic vector x i =[x i A,x i B,...,x i G],x i A,x i B,...,x i G respectively represents annual electricity data A, annual residents’ income and consumption data B, annual socio-economic data C, annual basic information of industrial development zones D, annual industrial output value E, annual industrial product price index F, and annual average price of industrial raw materials G.

[0012] Preferably, the prediction model includes a primary learning model and a secondary learning model, and step S3 includes:

[0013] Step S31, dividing the sample set into a training set and a test set, wherein the test set is the sample data of this year, and the training set is the sample data of other years;

[0014] Step S32, establishing the primary learning model S X, and use the output of each primary learning model as the input of the secondary learning model, the primary learning models include a linear regression algorithm model S1, a decision tree algorithm model S2, a support vector machine model S3, a k-nearest neighbor algorithm model S4, a random forest algorithm model S5, an AdaBoost algorithm model S6, a gradient regression algorithm model S7 and a time series analysis algorithm model S8, each of the primary learning models is trained by the training set and tested using the test set;

[0015] Step S33, establishing the secondary learning model, the secondary learning model is used to 2 The method and genetic algorithm assign weight configuration to the output result of the primary learning model and set the output result of the secondary learning model, and the output result of the secondary learning model is the output result of the prediction model.

[0016] Preferably, the step S33 includes:

[0017] Step S331, with R X The minimum is the goal, and the genetic algorithm is used to find the appropriate individual. The objective function f(x) of the genetic algorithm is:

[0018] f(x)=min(R1+R2.+Rg)

[0019]

[0020] Among them, R X is the primary learning model S X The correction coefficient of the output result, x∈[1,8], n is the element x in the training set i number, R2 is the linear regression determination coefficient, p is the number of variables, and S is the number of variables through the primary learning model X Calculate the feature vector x in the training set i The annual industrial added value forecast obtained is based on the training set and the primary learning model S X Calculate the mean of all the above values;

[0021] Step S332, according to the R in f(x) obtained by the genetic algorithm X Calculate each of the primary learning models S X The maximum weight value Q X :

[0022]

[0023] Step S333, the output result f(y) of the secondary learning model is defined as:

[0024]

[0025] The calculation process of step S331 includes:

[0026] Step S331a, population initialization, read in the original data, and convert (x i ,y i ) and each of the primary learning models S X The output industrial added value forecast is converted into a genetic population P, in which multiple unevolved chromosomes j are set i , j i By the primary learning model S X The output industrial added value forecast and a set of actual values ​​(x i ,y i ), set the population size N, the maximum genetic generation N max and mutation rate a;

[0027] Step S331b, calculate the gene fitness, using the R X The minimum sum is the objective function, and the fitness function Z(x) is the inverse of the objective function:

[0028]

[0029] in:

[0030]

[0031] Step S331c, genetic selection, calculate each of the unevolved chromosomes ji, bring the result into the fitness function Z(x) to obtain the corresponding fitness value; loop n times to obtain n fitness values, sort the obtained n fitness values ​​from large to small, replace the last 1 / 3 with the first 1 / 3, and re-form n daughter chromosomes; thereby, traverse any two genes in any two chromosomes in the genetic population P to perform gene crossover to obtain evolved daughter chromosomes, and the daughter chromosomes constitute the daughter genetic population P'; repeat the genetic selection process until the number of daughter chromosomes in the daughter genetic population P' is the same as the number of unevolved chromosomes in the daughter genetic population P';

[0032] Step S331d, stop evolution, when the fitness of the offspring chromosome in the offspring genetic population P' is greater than or equal to the recombined offspring chromosome or high fitness chromosome, or the current evolution number reaches the maximum genetic generation number N max When , the optimization is stopped. At this time, f(x) reaches the minimum. The primary learning models S are calculated according to the minimum f(x). X The maximum weight value Q X .

[0033] The present invention also provides an industrial added value growth rate prediction system based on stacking integrated algorithm, comprising:

[0034] An acquisition module, used for acquiring sample data, wherein the sample data includes annual parameter data and annual industrial added value, wherein the annual parameter data includes annual economic data and annual power data;

[0035] A preprocessing module, used for preprocessing each of the sample data and inputting the preprocessing data into a sample set;

[0036] A model building module is used to perform industrial added value growth rate forecasting modeling based on the sample set, based on the Stacking algorithm, combined with the linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm to obtain a forecasting model;

[0037] A prediction module, used to predict the growth rate of industrial added value using the prediction model;

[0038] An evaluation module is used to evaluate the forecasting model using volatility and accuracy.

[0039] Preferably, the sample data is recorded as (x i ,y i ), where x i is the annual parameter data, y i is the industrial added value of the year, i is the year, and the characteristic vector x i =[x i A,x i B,...,x i G],x i A,x i B,...,x i G respectively represents annual electricity data A, annual residents’ income and consumption data B, annual socio-economic data C, annual basic information of industrial development zones D, annual industrial output value E, annual industrial product price index F, and annual average price of industrial raw materials G.

[0040] Preferably, the prediction model includes a primary learning model and a secondary learning model, and the model building module includes:

[0041] A sample set processing unit, which divides the sample set into a training set and a test set, wherein the test set is the sample data of this year, and the training set is the sample data of other years;

[0042] Establishing unit, establishing the primary learning model S X, and use the output of each primary learning model as the input of the secondary learning model, the primary learning models include a linear regression algorithm model S1, a decision tree algorithm model S2, a support vector machine model S3, a k-nearest neighbor algorithm model S4, a random forest algorithm model S5, an AdaBoost algorithm model S6, a gradient regression algorithm model S7 and a time series analysis algorithm model S8, each of the primary learning models is trained by the training set and tested using the test set;

[0043] The establishing unit establishes the secondary learning model, and the secondary learning model is used to 2 The method and genetic algorithm assign weight configuration to the output result of the primary learning model and set the output result of the secondary learning model, and the output result of the secondary learning model is the output result of the prediction model.

[0044] It can be seen from the above technical scheme that the industrial added value growth rate prediction method based on the stacking integrated algorithm provided by the embodiment of the present invention first obtains sample data, wherein the sample data includes annual parameter data and annual industrial added value, and the annual parameter data includes annual economic data and annual power data; pre-processes each sample data and inputs it into the sample set; according to the sample set, based on the Stacking algorithm, combined with the linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm, the industrial added value growth rate prediction model is performed to obtain the prediction model; the prediction model is used to predict the industrial added value growth rate; and the prediction model is evaluated using volatility and accuracy. By combining the prediction method, the prediction is completed, the sensitivity of a single algorithm is reduced, and the load prediction accuracy is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flowchart of the industrial added value growth rate prediction method based on stacking ensemble learning algorithm;

[0046] Figure 2 It is the system framework diagram;

[0047] Figure 3 It is a block diagram of an electronic device used to implement the industrial added value growth rate prediction method based on the stacking integrated algorithm of this embodiment.

[0048] In the figure, 300, electronic device; 301, computing unit; 302, ROM; 303, RAM; 304, bus; 305, I / O interface; 306, input unit; 307, output unit; 308, storage unit; 309, communication unit. DETAILED DESCRIPTION

[0049] The technical scheme and technical effects of the present invention are further elaborated in detail below in conjunction with the accompanying drawings of the present invention.

[0050] The purpose of the present invention is to provide an industrial added value prediction method based on the fusion of multiple machine learning algorithms. By collecting economic data and power data, and cleaning the data, the first layer model is constructed using a linear regression algorithm, a decision tree algorithm, a support vector machine, a k-nearest neighbor algorithm, a random forest algorithm, an AdaBoost algorithm, a gradient regression algorithm and a time series analysis algorithm, the model is trained and the prediction results are tested, and then the prediction accuracy of the algorithm is combined to construct a comprehensive prediction model using an improved stacking ensemble learning algorithm.

[0051] like Figure 1 As shown, the industrial added value growth rate prediction method based on stacking ensemble learning algorithm provided by the present invention comprises the following steps:

[0052] Step S1, obtaining sample data, the sample data including annual parameter data and annual industrial added value, the annual parameter data including annual economic data and annual power data;

[0053] Step S2, preprocessing each sample data and inputting it into the sample set;

[0054] Step S3, based on the sample set, based on the Stacking algorithm, combined with the linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm, perform industrial added value growth rate forecast modeling to obtain a forecast model;

[0055] Step S4, using the forecasting model to forecast the growth rate of industrial added value;

[0056] Step S5, evaluating the forecasting model using volatility and accuracy.

[0057] In the embodiment of the present invention, the sample data is recorded as (x i ,y i ), where x i is the annual parameter data, y i is the annual industrial added value, i is the year, and the characteristic vector x i =[x i A,x iB,...,xiG], xiA,xiB,...,xiG are annual electricity data A, annual resident income and consumption data B, annual social and economic data C, annual basic information of industrial development zones D, annual industrial output value E, annual industrial product price index F, and annual average price of industrial raw materials G; sample data is the basis for the construction of machine learning models, and each sample contains the data required for model training. When collecting samples, in order to meet the model's evaluation of relevance and multi-dimensionality, it is not enough to only collect the data to be evaluated, and as much relevant data as possible should be collected.

[0058] Step S2 preprocesses the data. Since the data are not all structured data, and some data may be missing or delayed, the data needs to be processed. The processing methods are mainly to delete duplicate values ​​and the associated filling method; duplicate or invalid data in the sample can be eliminated; for incomplete, erroneous or inconsistent data, because the required data are basically structured data, approximate data that is highly correlated with the target data can be found, and the data correlation can be used to complete the filling. For example, the growth rates of residents' income and regional GDP are usually consistent. When the residents' income data is missing, the residents' income data can be filled according to the growth of regional GDP data, and the two can complement each other. In addition, the time of the data needs to be consistent. Because the required data are highly correlated with time, all data are sorted according to the same time dimension.

[0059] The prediction model includes a primary learning model and a secondary learning model.

[0060] In the above step S3, economic data, power generation, etc. are used as independent variables, and industrial added value is used as dependent variable, and the industrial added value model is established one by one by using linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm. The specific process of establishing the prediction model includes:

[0061] Step S31, dividing the sample set into a training set and a test set, wherein the test set is the sample data of this year, and the training set is the sample data of other years;

[0062] Step S32, establish a primary learning model S X , and use the output of each primary learning model as the input of the secondary learning model, the primary learning models include linear regression algorithm model S1, decision tree algorithm model S2, support vector machine model S3, k nearest neighbor algorithm model S4, random forest algorithm model S5, AdaBoost algorithm model S6, gradient regression algorithm model S7 and time series analysis algorithm model S8, each primary learning model is trained by the training set and tested by the test set;

[0063] Step S33, establish a secondary learning model, the secondary learning model is used to 2 The method and genetic algorithm allocate weight configuration to the output result of the primary learning model and set the output result of the secondary learning model, and the output result of the secondary learning model is the output result of the prediction model.

[0064] The improved stacking algorithm is used to fuse multiple machine learning models to improve the overall prediction ability. The traditional stacking algorithm is usually designed in two layers. The first layer is composed of multiple sub-algorithms, and the second layer has only one meta-model. The predicted values ​​and true values ​​obtained by various algorithms in the first layer are trained and assigned corresponding weights to obtain the final result. However, the shortcomings of the stacking algorithm are also prominent, and it is easy to cause overfitting problems. The improved stacking algorithm avoids the problem of overfitting by limiting the maximum value of the weight of the first-layer sub-algorithm. Therefore, the purpose of establishing a secondary learning model in step S33 is to perform algorithm integration, specifically:

[0065] Step S331, with R X The minimum is the goal, and the genetic algorithm is used to find the appropriate individual. The objective function f(x) of the genetic algorithm is:

[0066] f(x)=min(R+R2..+R)

[0067]

[0068] Among them, R X is the primary learning model S X The correction coefficient of the output result, x∈[1,8], n is the element x in the training set i Number, this patent uses the data of the past five years as the training set, and there are 12 sets of monthly data each year. Therefore, if the monthly data granularity is used, then n = 12 * 5 = 60, if the annual data granularity is used, then n = 5, R2 is the linear regression determination coefficient, p is the number of variables, this patent has a total of 7 variables, namely annual electricity data A, annual resident income and consumption data B, annual social and economic data C, annual industrial development zone basic information D, annual industrial output value E, annual industrial product price index F, annual industrial raw material average price G, so this patent p = 7, To pass the primary learning model S X Calculate the annual industrial added value forecast obtained by the feature vector xi in the training set, Based on the training set, the primary learning model S X Calculate all the means;

[0069] Step S332, according to the R in f(x) obtained by the genetic algorithm XCalculate each of the primary learning models S X The maximum weight value Q X :

[0070]

[0071] Step S333, the output result f(y) of the secondary learning model is defined as:

[0072]

[0073] The calculation process of step S331 includes:

[0074] Step S331a, population initialization, read in the original data, and convert (x i ,y i ) and each primary learning model S X The output industrial added value forecast is converted into a genetic population P, in which multiple unevolved chromosomes j are set i , j i The predicted industrial added value output by the primary learning model SX and a set of actual values ​​(x i ,y i ), set the population size N, the maximum genetic generation N max and mutation rate a. Here, the population size can be set to 40, the maximum genetic generation to 50, and the mutation rate to 0.1;

[0075] Step S331b, calculate the gene fitness, using the R X The minimum sum is the objective function, and the fitness function Z(x) is the inverse of the objective function:

[0076]

[0077] in:

[0078]

[0079] Step S331c, genetic selection, for each of the non-evolved chromosomes j i Calculate and bring the result into the fitness function Z(x) to get the corresponding fitness value; loop n times to get n fitness values, sort the obtained n fitness values ​​from large to small, replace the last 1 / 3 with the first 1 / 3, and re-form n daughter chromosomes; thus, traverse any two genes in any two chromosomes in the genetic population P to perform gene crossover to obtain evolved daughter chromosomes, and the daughter chromosomes constitute the daughter genetic population P'; repeat the genetic selection process until the number of daughter chromosomes in the daughter genetic population P' is the same as the number of unevolved chromosomes in the daughter genetic population P';

[0080] Step S331d, stop evolution, when the fitness of the offspring chromosome in the offspring genetic population P' is greater than or equal to the recombined offspring chromosome or high fitness chromosome, or the current evolution number reaches the maximum genetic generation number N max When , the optimization stops. At this time, f(x) reaches the minimum. According to the minimum f(x), each primary learning model S is calculated. X The maximum weight value Q X .

[0081] Step S5 evaluates the output of the prediction model using both accuracy and volatility, where:

[0082] Accuracy = accurate number / total number

[0083] Volatility = ∑(predicted result - actual result) 2

[0084] The sample set is divided into a training set and a validation set. The training set is used to train the model, and the validation set is used to verify the accuracy of the model. In the validation set, the error range between the predicted results and the actual results is within 5%, which can be considered accurate, and the rest is inaccurate.

[0085] After testing, the accuracy of the combined forecasting model is higher than that of most single models, and the volatility of the results is the smallest. It can be used as a method to predict industrial added value.

[0086] The present invention utilizes the advantages and avoids the disadvantages of a combined prediction method. The combined prediction method combines different algorithms by weighting to jointly complete the prediction, thereby reducing the sensitivity of a single algorithm and improving the load prediction accuracy.

[0087] The present invention also provides an industrial added value growth rate prediction system based on stacking integrated algorithm, which is used to implement Figure 1 The method shown in the system includes:

[0088] An acquisition module is used to acquire sample data, the sample data includes annual parameter data and annual industrial added value, and the annual parameter data includes annual economic data and annual power data;

[0089] A preprocessing module is used to preprocess each sample data and input it into the sample set;

[0090] The model building module is used to perform industrial added value growth rate forecasting modeling based on the sample set, based on the Stacking algorithm, combined with the linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm to obtain a forecasting model;

[0091] The forecasting module is used to forecast the growth rate of industrial added value using the forecasting model;

[0092] Evaluation module for evaluating forecasting models using volatility and accuracy.

[0093] Among them, the sample data is recorded as (x i ,y i ), where x i is the annual parameter data, y i is the annual industrial added value, i is the year, and the characteristic vector x i =[x i A,x i B,...,x i G],x i A,x i B,...,x i G are annual electricity data A, annual resident income and consumption data B, annual social and economic data C, annual industrial development zone basic information D, annual industrial output value E, annual industrial product price index F, and annual average price of industrial raw materials G. The prediction model includes primary learning model and secondary learning model.

[0094] The model building module includes:

[0095] The sample set processing unit divides the sample set into a training set and a test set, wherein the test set is the sample data of this year and the training set is the sample data of other years;

[0096] Establish units and build primary learning models X , and use the output of each primary learning model as the input of the secondary learning model, the primary learning models include linear regression algorithm model S1, decision tree algorithm model S2, support vector machine model S3, k nearest neighbor algorithm model S4, random forest algorithm model S5, AdaBoost algorithm model S6, gradient regression algorithm model S7 and time series analysis algorithm model S8, each primary learning model is trained by the training set and tested by the test set;

[0097] Establish a unit and establish a secondary learning model, which is used to learn through R 2 The method and genetic algorithm allocate weight configuration to the output result of the primary learning model and set the output result of the secondary learning model, and the output result of the secondary learning model is the output result of the prediction model.

[0098] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0099] Figure 3 The following examples can be used to implement Figure 1Schematic block diagram of an example electronic device 300 of the illustrated method. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0100] like Figure 3 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. In RAM 303, various programs and data required for the operation of the device 300 can also be stored. The computing unit 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0101] A number of components in the device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0102] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 301 performs the various methods and processes described above, such as the pruning method of the machine learning model. For example, in some embodiments, the pruning method of the machine learning model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the pruning method of the machine learning model described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform the pruning method of the machine learning model in any other appropriate manner (e.g., by means of firmware).

[0103] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0104] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0105] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0107] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0108] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0109] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0110] What is disclosed above is only a preferred embodiment of the present invention, which certainly cannot be used to limit the scope of rights of the present invention. A person skilled in the art can understand that all or part of the processes of the above embodiments and equivalent changes made according to the claims of the present invention still fall within the scope of the invention.

Claims

1. A method for predicting the growth rate of industrial added value, characterized in that: include: Step S1, obtaining sample data, wherein the sample data includes annual parameter data and annual industrial added value, and the annual parameter data includes annual economic data and annual power data; Step S2, pre-processing each of the sample data and inputting them into a sample set; Step S3, based on the sample set, based on the Stacking algorithm, combined with the linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm, perform industrial added value growth rate forecast modeling to obtain a forecast model; Step S4, using the prediction model to predict the growth rate of industrial added value; Step S5, evaluating the prediction model using volatility and accuracy; The sample data is recorded as (x i ,y i ), where x i is the annual parameter data, y i is the industrial added value of the year, i is the year, and the characteristic vector xi = [x i A,x i B,...,x i G],x i A,x i B,...,x i G respectively represent annual electricity data A, annual resident income and consumption data B, annual social and economic data C, annual basic information of industrial development zones D, annual industrial output value E, annual industrial product price index F, and annual average price of industrial raw materials G; The prediction model includes a primary learning model and a secondary learning model, and step S3 includes: Step S31, dividing the sample set into a training set and a test set, wherein the test set is the sample data of this year, and the training set is the sample data of other years; Step S32, establish a primary learning model S X , and using the output of each primary learning model as the input of the secondary learning model, the primary learning models include a linear regression algorithm model S1, a decision tree algorithm model S2, a support vector machine model S3, a k-nearest neighbor algorithm model S4, a random forest algorithm model S5, an AdaBoost algorithm model S6, a gradient regression algorithm model S7 and a time series analysis algorithm model S8, each of the primary learning models is trained by the training set and tested using the test set; Step S33, establishing the secondary learning model, the secondary learning model is used to 2 The method and genetic algorithm assign weight configuration to the output result of the primary learning model and set the output result of the secondary learning model, and the output result of the secondary learning model is the output result of the prediction model; The step S33 comprises: Step S331, with R X The minimum is the goal, and the genetic algorithm is used to find the appropriate individual. The objective function f(x) of the genetic algorithm is: f(x)=min(R1+R2...+R8) Among them, R X is the primary learning model S X The correction coefficient of the output result, x∈[1,8], n is the element x in the training set i Number, R 2 is the linear regression coefficient of determination, p is the number of variables, To pass the primary learning model S X Calculate the annual industrial added value forecast obtained by the feature vector xi in the training set, Based on the training set, the primary learning model S X Calculate all the means; Step S332, according to the R in f(x) obtained by the genetic algorithm X Calculate each of the primary learning models S X The maximum weight value Q X : Step S333, the output result f(y) of the secondary learning model is defined as: The calculation process of step S331 includes: Step S331a, population initialization, read in the original data, and convert (x i ,y i ) and each of the primary learning models S X The output industrial added value forecast is converted into a genetic population P, in which multiple unevolved chromosomes j are set i , j i By the primary learning model S X The output industrial added value forecast and a set of actual values ​​(x i ,y i ), set the population size N, the maximum genetic generation N max and mutation rate a; Step S331b, calculate the gene fitness, using the R X The minimum sum is the objective function, and the fitness function Z(x) is the inverse of the objective function: in: Step S331c, genetic selection, for each of the non-evolved chromosomes j i Perform calculations, and bring the results into the fitness function Z(x) to obtain the corresponding fitness value; repeat n times to obtain n fitness values, sort the obtained n fitness values ​​from large to small, and replace the last 1 / 3 with the first 1 / 3 to re-form n daughter chromosomes; thereby, traverse any two genes in any two chromosomes in the genetic population P to perform gene crossover to obtain evolved daughter chromosomes, and the daughter chromosomes constitute the daughter genetic population P'; repeat the genetic selection process until the number of daughter chromosomes in the daughter genetic population P' is the same as the number of unevolved chromosomes in the daughter genetic population P'; Step S331d, stop evolution, when the fitness of the offspring chromosome in the offspring genetic population P' is greater than or equal to the recombined offspring chromosome or high fitness chromosome, or the current evolution number reaches the maximum genetic generation number N max When , the optimization is stopped. At this time, f(x) reaches the minimum. The primary learning models S are calculated according to the minimum f(x). X The maximum weight value Q X .

2. An industrial value-added growth rate prediction system using the industrial value-added growth rate prediction method according to claim 1, characterized in that: include: An acquisition module, used for acquiring sample data, wherein the sample data includes annual parameter data and annual industrial added value, wherein the annual parameter data includes annual economic data and annual power data; A preprocessing module, used for preprocessing each of the sample data and inputting the preprocessing data into a sample set; A model building module is used to perform industrial added value growth rate forecasting modeling based on the sample set, based on the Stacking algorithm, combined with the linear regression algorithm, decision tree algorithm, support vector machine, k nearest neighbor algorithm, random forest algorithm, AdaBoost algorithm, gradient regression algorithm and time series analysis algorithm to obtain a forecasting model; A prediction module, used to predict the growth rate of industrial added value using the prediction model; An evaluation module is used to evaluate the forecasting model using volatility and accuracy.

3. The industrial added value growth rate prediction system according to claim 2 is characterized in that: The sample data is recorded as (x i ,y i ), where x i is the annual parameter data, y i is the industrial added value of the year, i is the year, and the characteristic vector xi = [x i A,x i B,...,x i G],x i A,x i B,...,x i G respectively represents annual electricity data A, annual residents’ income and consumption data B, annual socio-economic data C, annual basic information of industrial development zones D, annual industrial output value E, annual industrial product price index F, and annual average price of industrial raw materials G.

4. The industrial added value growth rate prediction system according to claim 3 is characterized in that: The prediction model includes a primary learning model and a secondary learning model, and the model building module includes: A sample set processing unit, which divides the sample set into a training set and a test set, wherein the test set is the sample data of this year, and the training set is the sample data of other years; Establishing unit, establishing the primary learning model S X , and use the output of each primary learning model as the input of the secondary learning model, the primary learning models include a linear regression algorithm model S1, a decision tree algorithm model S2, a support vector machine model S3, a k-nearest neighbor algorithm model S4, a random forest algorithm model S5, an AdaBoost algorithm model S6, a gradient regression algorithm model S7 and a time series analysis algorithm model S8, each of the primary learning models is trained by the training set and tested using the test set; The establishing unit establishes the secondary learning model, and the secondary learning model is used to 2 The method and genetic algorithm assign weight configuration to the output result of the primary learning model and set the output result of the secondary learning model, and the output result of the secondary learning model is the output result of the prediction model.

5. A storage medium, characterized in that: The storage medium includes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method of claim 1.

6. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can execute the method of claim 1.

Citation Information

Patent Citations

  • Lithium battery remaining service life prediction method and system based on ensemble learning

    CN114881246A