Layered and graded economic development influence factor decomposition method and system based on LASSO model and panel fixing effect model

By combining the LASSO model and the panel fixed effects model, the problems of inaccurate variable selection and insufficient control of fixed effects in traditional methods are solved. This enables accurate decomposition and visualization analysis of multi-level economic development influencing factors, and improves the stability and explanatory power of the model.

CN121390933APending Publication Date: 2026-01-23GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511286345.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Traditional methods for identifying influencing factors suffer from several drawbacks under conditions of multiple variables, multiple dimensions, and large samples. These include strong subjectivity in indicator selection, severe collinearity interference between variables, and insufficient handling of panel heterogeneity, making it difficult to accurately assess the marginal effect of each factor on regional economic development.

Method used

Using the LASSO model and panel fixed effects model, a multi-level panel data structure was established. Through standardized preprocessing and LASSO regression model screening, key explanatory variables were selected, a panel fixed effects regression model was constructed, the marginal effect coefficients of each variable on the economic development target indicators were quantified, the contribution rate weights were calculated, and a cross-level influence path difference map was generated.

Benefits of technology

It improves the stability of the model under high-dimensional data conditions and the identifiability of variable action paths, provides a systematic and quantifiable multi-level economic governance mechanism, and enhances the estimation accuracy and interpretability of regression coefficients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121390933A_ABST
    Figure CN121390933A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of economic data analysis and decision support, and discloses a hierarchical and hierarchical economic development influence factor decomposition method and system based on an LASSO model and a panel fixing effect model.The method comprises the steps that a multi-level panel data structure comprising a region layer, an industry layer and an enterprise layer is established, and economic index data of all levels are collected; performing standardized preprocessing on each level of data, wherein the standardized preprocessing comprises missing value filling, abnormal value processing, dimension unification and fixed effect elimination; based on the preprocessed data, adopting an LASSO regression model to determine regularization parameters, and screening key explanatory variables in a layered manner; inputting the screened variables into a panel fixed effect regression model, and quantifying marginal effect coefficients of the variables for economic development target indexes; calculating the contribution rate weight of the key variable at each hierarchy, and constructing a cross-hierarchy influence path difference graph. According to the method, structured and systematic identification of a multi-dimensional variable influence mechanism in an economic system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of economic data analysis and decision support, and particularly relates to a hierarchical and graded economic development influence factor decomposition method and system based on a LASSO model and a panel fixed effect model. BACKGROUND

[0002] Traditional influence factor identification methods such as principal component analysis (PCA), entropy weight method, and grey correlation analysis have significant defects under the conditions of multivariable, multidimension, and large sample, such as strong subjectivity of index selection, serious collinearity interference between variables, and insufficient panel heterogeneity processing, and it is difficult to accurately evaluate the marginal effect of each factor on regional economic development.

[0003] In recent years, LASSO regression (least absolute shrinkage and selection operator) has been valued because of its high-dimensional variable compression and feature selection capability, and can effectively handle the problem of multiple collinearity; and the fixed effect model as a core tool for panel data analysis can eliminate unobservable time-invariant influence factors of regions. Few existing studies combine the two for multilevel regional economic analysis, especially lacking a systematic "hierarchical-grading" factor analysis framework. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a hierarchical and graded economic development influence factor decomposition method based on a LASSO model and a panel fixed effect model, which can solve the problems of inaccurate multivariate economic variable screening, insufficient fixed effect control, and fuzzy regional development rule identification in traditional methods.

[0006] To solve the above technical problems, the present application provides the following technical scheme, a hierarchical and graded economic development influence factor decomposition method based on a LASSO model and a panel fixed effect model, comprising: establishing a multilevel panel data structure including a regional level, an industry level, and an enterprise level, and collecting economic index data of each level; standardizing and preprocessing each level data, including missing value filling, outlier processing, dimension unification, and fixed effect elimination; based on the preprocessed data, using a LASSO regression model to determine a regularization parameter, and hierarchically screening key explanatory variables; inputting the screened variables into a panel fixed effect regression model to quantify the marginal effect coefficient of each variable on the economic development target index; calculating the contribution rate weight of the key variables at each level, and constructing a cross-level influence path difference map.

[0007] As a preferred scheme of the hierarchical and graded economic development influence factor decomposition method based on a LASSO model and a panel fixed effect model, the standardization preprocessing includes, for missing data, using a time series linear interpolation method or a cross-section mean method for filling.

[0008] Z-score standardization is performed on all variables;

[0009] Log transformation is performed on high-amplitude skewed variables;

[0010] Individual fixed effects are eliminated by panel unit mean removal operation.

[0011] As a preferred solution of the hierarchical and graded economic development influencing factor decomposition method based on the LASSO model and the panel fixed effect model, the hierarchical screening of key explanatory variables comprises: the hierarchical construction of an objective function is achieved by optimizing a LASSO objective function with an L1 regularization term, the optimal penalty coefficient is determined by cross-validation, and the key variable set is reserved according to the non-zero coefficients.

[0012] As a preferred solution of the hierarchical and graded economic development influencing factor decomposition method based on the LASSO model and the panel fixed effect model, the fixed effect regression model comprises: the individual fixed effect term controls the inherent characteristics of the region / industry / enterprise.

[0013] The time fixed effect term controls the annual macro fluctuations.

[0014] The screened variables are used as explanatory variables, and the regression coefficients are output.

[0015] As a preferred solution of the hierarchical and graded economic development influencing factor decomposition method based on the LASSO model and the panel fixed effect model, the calculation of the contribution rate weight of the key variable at each level comprises: the variable contribution rate normalization index is defined,

[0016]

[0017] wherein, represents the variable x j The structural contribution weight on the level L is used for visualizing the structural heterogeneity, represents the variable x j The regression coefficient on the level L.

[0018] As a preferred solution of the hierarchical and graded economic development influencing factor decomposition method based on the LASSO model and the panel fixed effect model, the construction of the cross-level influence path difference atlas comprises: the contribution rate weights of the same variable at each level are compared.

[0019] The direction conflict variable is marked when the regression coefficients of the variable at at least two levels are opposite in sign.

[0020] The variable contribution rate distribution at each level and the cross-level correlation path are visualized.

[0021] As a preferred scheme of the hierarchical and graded economic development influence factor decomposition method based on the LASSO model and the panel fixed effect model, wherein: the multi-level panel data structure comprises a region layer defining a province / city level observation unit, an industry layer defining a subdivided industry observation unit, and an enterprise layer defining an independent enterprise observation unit.

[0022] The region layer defines a province / city level observation unit, and the indexes include GDP growth rate, unit GDP energy consumption, and clean energy proportion.

[0023] The industry layer defines a subdivided industry observation unit, and the indexes include green total factor productivity and industry energy efficiency index.

[0024] The enterprise layer defines an independent enterprise observation unit, and the indexes include carbon intensity and photovoltaic power generation proportion.

[0025] The indexes at each level are associated according to a preset mapping rule, and the comparable variables are uniformly defined across levels.

[0026] It should be noted that by defining the variable contribution rate normalization index, the relative importance of each factor within the level is accurately quantified, overcoming the defect that the traditional method can only determine the direction but cannot measure the structural weight, and providing a quantitative basis for hierarchical management. Through the structured design of level-specific indexes and cross-level comparable variables, the multi-dimensional data heterogeneity problem of regional GDP, industry energy efficiency, enterprise carbon intensity, etc. is solved, ensuring the input consistency of LASSO and fixed effect model in cross-level analysis.

[0027] The present application provides a hierarchical and graded economic development influence factor decomposition system based on the LASSO model and the panel fixed effect model.

[0028] As a preferred scheme of the hierarchical and graded economic development influence factor decomposition system based on the LASSO model and the panel fixed effect model, wherein: comprising a data preprocessing module, a variable screening module, a regression analysis module, a contribution rate calculation module, and a graph generation module.

[0029] The data preprocessing module is used to perform missing value filling, Z-score standardization, high amplitude skewness variable logarithmic transformation and panel unit mean removal operation to eliminate individual fixed effects.

[0030] The variable screening module is used to construct a LASSO target function with L1 regularization term in layers, determine the optimal penalty coefficient λ through cross-validation, and output a set of non-zero coefficient key variables.

[0031] The regression analysis module is used to input the key variables into a panel regression model containing individual fixed effect term and time fixed effect term, and calculate the marginal effect coefficient of each variable.

[0032] The contribution rate calculation module is configured to calculate the contribution rate weight of each hierarchical variable according to a formula.

[0033] The atlas generation module is configured to compare the cross-hierarchical variable contribution rate weights, identify the regression coefficient sign conflict variables, and generate a visual impact path difference atlas.

[0034] The present application provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of a hierarchical and hierarchical economic development influence factor decomposition method based on LASSO model and panel fixed effect model when executing the computer program.

[0035] The present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of a hierarchical and hierarchical economic development influence factor decomposition method based on LASSO model and panel fixed effect model when executed by a processor.

[0036] The present application has the following beneficial effects: the LASSO model is used to perform sparse processing on the original explanatory variable set in the multi-hierarchical panel data of regions, industries, enterprises, etc., and by introducing an L1 regularization penalty term, the automatic removal of redundant variables in high-dimensional indicators is realized, and the most representative candidate variable subset which has a significant linear contribution to the target variable is reserved, thereby effectively solving the multiple collinearity problem and model instability problem that may occur in the traditional model under the condition of high variable dimension and limited sample size. Further, a fixed effect panel regression model is constructed, and the variables selected by LASSO are used as explanatory variables, and under the premise of not removing the original fixed effect, the true marginal effect coefficient of each variable on the economic development target variable (such as green total factor productivity, unit GDP carbon emission, etc.) is re-estimated. This process effectively controls the interference terms such as unobservable regional / industry / enterprise characteristics and year cycle impact by explicitly introducing individual fixed effect terms and year fixed effect terms, thereby improving the estimation accuracy, significance and interpretability of the regression coefficient. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0038] Figure 1 A flowchart of a hierarchical and hierarchical economic development influence factor decomposition method based on LASSO model and panel fixed effect model is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0039] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present application.

[0040] Embodiment 1, refer to Figure 1 For the first embodiment of the present application, the embodiment provides a hierarchical and graded economic development influencing factor decomposition method based on LASSO model and panel fixed effect model, comprising:

[0041] S1: Establish a multi-level panel data structure containing a regional layer, an industry layer and an enterprise layer, and collect economic index data of each level;

[0042] S2: Standardize and preprocess each level data, including missing value filling, outlier processing, dimension unification and fixed effect elimination.

[0043] S3: Based on the preprocessed data, determine the regularization parameter by using the LASSO regression model, and filter the key explanatory variables in layers.

[0044] S4: Input the filtered variables into the panel fixed effect regression model, and quantify the marginal effect coefficient of each variable on the economic development target index.

[0045] S5: Calculate the contribution rate weight of the key variables in each level, and construct a cross-level influence path difference map.

[0046] It should be noted that the present application realizes the integrated modeling path from "variable selection" to "mechanism identification" through the collaborative construction of LASSO model and fixed effect regression model, which not only improves the stability of the model under high-dimensional data conditions, but also enhances the distinguishability of the variable action path in the multi-level structure, and provides a systematic and quantifiable index support scheme for macro-micro-micro multi-dimensional governance mechanism.

[0047] Embodiment 2, for an embodiment of the present application, based on the above embodiment, a hierarchical and graded economic development influencing factor decomposition method based on LASSO model and panel fixed effect model is provided.

[0048] Furthermore, in the present application, step S1 establishes a multi-level panel data structure containing a regional layer, an industry layer and an enterprise layer, and collects economic index data of each level, which specifically includes steps A1-A2:

[0049] A1: Establish a multi-level panel data organization structure, and set the research object structure as L=3 levels: regional level L=1, industry level L=2, and enterprise level L=3. Each level includes:

[0050] Number of units N (l) : represents the number of observation objects under this level (such as the number of provinces / cities, industries, and enterprises);

[0051] Time length T (l) : number of consecutive observation years;

[0052] Index dimension p (l) : number of candidate explanatory variables.

[0053] Construct the original data set:

[0054]

[0055] Wherein, represents the value of the i th unit (for the regional level, it is a specific region; for the industry level, it is a specific industry; for the enterprise level, it is a specific enterprise) in the t th year j th variable index, corresponds to a certain economic high-quality development target variable (such as regional GDP, industry added value, and enterprise profit) of the corresponding level, is a real number set.

[0056] A2: Initial index variable selection basis and level mapping are as follows:

[0057] According to different analysis targets of different levels, the regional level: GDP, per capita R&D input, unit GDP energy consumption, and ecological investment proportion can be considered for selection from the aspects of power factors, social factors, environmental factors, and innovation factors;

[0058] The industry level: added value rate, industry energy efficiency, power density, and technology investment can be considered for selection from the aspects of energy structure, resource consumption, industrial chain link, and market fluctuation;

[0059] The enterprise level: unit revenue energy consumption, net profit rate, carbon intensity, and employee structure can be considered for selection from the aspects of capital structure, labor structure, energy use efficiency, and technology intensity.

[0060] In addition, indexes such as “unit energy consumption” can be defined in the three levels. Such indexes belong to comparable variables, and other indexes are level-specific variables.

[0061] Furthermore, in the embodiments of the present application, step S2 performs standardization preprocessing on the data of each level, including missing value filling, outlier processing, dimension unification, and fixed effect removal, specifically including steps B1-B4:

[0062] B1: Before building LASSO model for each level (regional level / industry level / enterprise level), the following preprocessing operations are needed to ensure the comparability between variables, the robustness of the model and the quality of input data.

[0063] The missing data of each dimension variable in each level is completed by using time series interpolation and cross-sectional mean imputation.

[0064] a) Time series interpolation method (longitudinal imputation)

[0065] For the case of missing years of a certain index variable for the same observation unit i (such as a certain region, industry or enterprise), linear interpolation method is used for imputation:

[0066]

[0067] Wherein, is the interpolation imputation result, t is the missing year to be imputed, t1, t2 are the data values of the unit in the adjacent non-missing years, corresponding to the year;

[0068] b) Cross-sectional mean method (lateral imputation)

[0069] In the process of missing value processing, if the variable j of a unit i is missing in year t and cannot be obtained by longitudinal interpolation (i.e. interpolation of previous and subsequent years), cross-sectional mean method can be used for completion. The calculation formula is as follows:

[0070]

[0071] Wherein, is the estimated value of the missing value of the jth variable of unit i in year t after cross-sectional mean imputation, x k,j,t is the observation value of the jth variable of other unit k (k≠i) in year t; N t is the total number of units in year t (for example, the number of observed regions, industries or enterprises in a year); N t -1 is the total number of units excluding the current missing unit, which is used as the denominator when calculating the mean.

[0072] In this embodiment, time series linear interpolation method or cross-sectional mean method is used to impute the missing data;

[0073] In an alternative embodiment, missing value imputation can also be achieved by adjacent observation moving average imputation, specifically, if there are continuous non-missing adjacent years before and after the missing year (e.g. the data of the previous 1 year, the previous 2 years or the next 1 year is complete), the arithmetic mean of the data of these adjacent years is directly calculated as the imputed value (e.g. the mean of the previous 1 year and the next 1 year of the missing year). If there is only one side data (e.g. only the previous year data without the next year data), only the mean of the available continuous years on that side is used for imputation (e.g. the mean of the previous 2 years and the previous 1 year of the missing year).

[0074] In another alternative embodiment, missing value imputation can also be achieved by the median of the same group, specifically, a classification standard is defined in advance (e.g. grouping by industry, size or region). In year t, other non-missing units belonging to the same group as the missing unit i are screened. The median (not the arithmetic mean) of the observations of these units on variable j is taken as the imputed value of unit i.

[0075] B2: Dimensionless and standardized, all variables are Z-score standardized:

[0076]

[0077] wherein, is the variable mean, is the standard deviation, is the standardized variable value.

[0078] In this embodiment, standardization is achieved by Z-score;

[0079] In an alternative embodiment, standardization can also be achieved by range standardization, specifically, for each variable, the minimum and maximum values of all observations are calculated; each original value is subtracted by the minimum value of the variable, and then divided by the difference between the maximum value and the minimum value (i.e. the range); the output result falls within the interval [0, 1], and the scales of all variables are unified.

[0080] In another alternative embodiment, standardization can also be achieved by unitization, specifically, for each variable, a scaling factor is determined according to its maximum absolute value (e.g. the maximum value is 250, then the scaling factor is 1000); each original value is divided by the scaling factor (e.g. 250→0.25); it is ensured that all result values fall within the interval [-1, 1].

[0081] B3: Combined with variable characteristics, log transformation is performed on variables with severe skewness and large span (e.g. GDP, power consumption, population size, etc.).

[0082] B4: To eliminate the interference of the inherent unchanging characteristics of each unit i on the regression estimate. "Panel unit mean removal" is performed on each variable of each layer to eliminate the influence of fixed effects, and a form of LASSO acceptable panel data is constructed, and the input variable is the original index variable vector Target variable vector Finally, a set of de-meaned indicators is obtained De-meaned target variable

[0083]

[0084] wherein, is the original value of the jth indicator of the ith unit in the lth layer in the tth year, is the original value of the target variable of unit i in year t, T (l) is the total number of observations of the lth layer, is the de-meaned explanatory variable value, is the de-meaned target variable value.

[0085] Further, in the embodiments of the present application, step S3 determines the regularization parameter based on the preprocessed data using the LASSO regression model, and hierarchically screens the key explanatory variables, specifically including steps C1-C5:

[0086] C1: In each layer, all the de-meaned indicators thereunder are uniformly input into the LASSO regression model. By optimizing the objective function with the L1 regularization term, a set of key variables that have significant explanatory power on the dependent variable is automatically screened out, and redundant indicators are removed.

[0087] C2: The LASSO optimization objective function is:

[0088]

[0089] wherein, β j is the regression coefficient of variable j, and λ is the control parameter of the regularization penalty term.

[0090] C3: The first term is the residual sum of squares of the model, which represents the fitting error part of LASSO. For each sample unit i in each year t, a predicted value is constructed according to all variables j; the difference (residual) between the predicted value and the true de-meaned value is calculated, and after squaring, the sum is accumulated over the entire sample range of i and t; finally, the total sample number N (l) T (l) is divided to calculate the average value of the residual sum of squares. The second term is the L1 regularization term, which is used to sparsely compress the variables and promote some coefficients β jshrink to 0, thus automatically implementing variable selection. λ is the regularization penalty control parameter, usually determined by cross-validation. The final retained non-zero coefficient variable index set is obtained

[0091] In this embodiment, the regularization penalty control parameter is determined by cross-validation method.

[0092] In an alternative embodiment, the regularization parameter can also be determined by a fixed empirical value method, specifically, a pre-defined λ value is selected according to similar data sets or prior knowledge. In each layer, the pre-processed de-meaned indicators are input into the LASSO model, and the fixed λ value is directly used to optimize the objective function. Based on the optimized model, the non-zero coefficient variable set is extracted as the key explanatory variables, without the need for λ tuning steps.

[0093] In another alternative embodiment, the regularization parameter can also be determined by a single validation set method, specifically, the data set is divided: the full sample data of the current layer is randomly divided into a training set and a validation set, ensuring the integrity of the time series. On the training set, multiple LASSO models are trained, each using a different pre-set λ value. On the validation set, the prediction error of each model is calculated, and the λ value with the smallest error is selected as the final parameter. The LASSO model is retrained on the complete data set using the optimal λ value, and the non-zero coefficient variable set is extracted.

[0094] C4: Automatically remove redundant variables from high-dimensional indicators, retain candidate explanatory variables that have a significant impact on the target variable of the layer, and reduce the influence of multicollinearity to improve the robustness of the regression.

[0095] Furthermore, in the embodiments of the present application, step S4 inputs the screened variables into the fixed effects panel regression model to quantify the marginal effect coefficients of each variable on the economic development target indicators, specifically including steps D1-D3:

[0096] D1: On the basis of the key variables screened by the LASSO model, a fixed effects panel regression model is further constructed, introducing an individual fixed effects term for each analysis unit Control the inherent characteristics of provinces / industries / companies (such as resource endowment), and the model controls macro time effects Control the policy or cycle effect at the year level, and realize quantitative identification and estimation of the economic significance, direction and marginal contribution of each key explanatory variable.

[0097] D2: Based on the results of step S3, a fixed effects panel regression model is constructed at the i-th layer:

[0098]

[0099] wherein, is a random term, is the individual fixed effect of unit i in the lth level, is the time fixed effect of year t in the lth level.

[0100] In this embodiment, the fixed effect regression model is achieved by controlling the individual fixed effect;

[0101] In an alternative embodiment, the fixed effect can also be achieved by grouping dummy variable method, specifically, for each independent unit of the current level, a binary dummy variable is created. The key variables screened by LASSO and all generated dummy variables are input into the ordinary least squares (OLS) regression model. Run OLS regression, at this time the coefficient of the dummy variable implies the individual fixed effect, and the coefficient of the key explanatory variable is the marginal effect.

[0102] In another alternative embodiment, the fixed effect can also be achieved by individual mean deviation method, specifically, the individual mean is calculated, and for each unit, the mean of each key variable in all years is calculated. The original value of each sample is subtracted from the corresponding variable mean of the unit to which it belongs. The converted deviation value is input into the OLS regression model without intercept term to directly estimate the key variable coefficient.

[0103] D3: Finally, the marginal effect coefficient of each retained variable in each level is obtained The coefficient value reflects the "direction and strength" of each variable, i.e. positive / negative and impact strength. If the unit GDP energy consumption β (1) < 0 indicates that the more energy-saving at the regional level, the stronger the green growth; if the proportion of the tertiary industry β (2) > 0 indicates that the higher the proportion of the service industry at the industry level, the stronger the promotion of green development; if the enterprise R&D investment β (3) > 0 indicates that the enterprise increases the technical input, which has significant effect on carbon emission reduction.

[0104] Further, in the embodiment of the application, step S5 calculates the contribution rate weight of the key variable in each level to construct a cross-level influence path difference map, specifically including:

[0105] Define variable contribution rate normalization index:

[0106]

[0107] Wherein, represents the structural contribution weight of variable x j on the level L (between 0 and 1), which is used to visualize the structural heterogeneity. The main effects are as follows: 1 determining whether the "direction of action of variables at different levels is consistent: if the variable is positive (promote growth) at the regional level, but negative (suppress profit) at the enterprise level, it indicates that the variable has "structural conflict", for example, clean energy investment reduces carbon intensity at the regional level (positive), but increases transformation cost at the enterprise level (negative), which prompts the government to adopt different policies and support different subjects to avoid "policy extrapolation conflict"; 2 identifying the "relative importance" of variables in the structure: for example, the structural contribution rate of the variable "unit energy consumption" is θ (1) = 0.5, θ (2) = 0.2, θ (3) = 0.3, indicating that its main influence is concentrated in the regional level (macro energy efficiency management), which prompts to prioritize the deployment of regional energy consumption quotas and energy-saving subsidy policies.

[0108] Embodiment 3 is the third embodiment of the present application, which is different from the first two embodiments:

[0109] The embodiment also provides a hierarchical economic development influence factor decomposition system based on a LASSO model and a panel fixed effect model, which comprises a data preprocessing module, a variable screening module, a regression analysis module, a contribution rate calculation module and a graph generation module.

[0110] The data preprocessing module is used to perform missing value filling, Z-score standardization, high amplitude skewness variable logarithmic transformation and panel unit mean removal operation to eliminate individual fixed effects.

[0111] The variable screening module is used to construct a LASSO target function with an L1 regularization term at different levels, determine the optimal penalty coefficient λ through cross-validation, and output a set of non-zero coefficient key variables.

[0112] The regression analysis module is used to input the key variables into a panel regression model containing individual fixed effect terms and time fixed effect terms, and calculate the marginal effect coefficients of each variable.

[0113] The contribution rate calculation module is used to calculate the contribution rate weight of each level variable according to the formula.

[0114] The graph generation module is used to compare the contribution rate weights of cross-level variables, identify the regression coefficient sign conflict variables, and generate a visual impact path difference graph.

[0115] The embodiment also provides an electronic device suitable for high-temperature pipeline health online monitoring, which comprises a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the high-temperature pipeline health online monitoring method proposed in the above embodiment.

[0116] The embodiment also provides a storage medium, which stores a computer program, and the computer program is executed by a processor to implement the high-temperature pipeline health online monitoring method.

[0117] The storage medium provided by the embodiment belongs to the same inventive concept as the high-temperature pipeline health online monitoring method provided by the above embodiment, and the technical details not described in detail in the embodiment can be referred to the above embodiment, and the embodiment has the same beneficial effects as the above embodiment.

[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be realized by software and necessary general hardware, and of course can also be realized by hardware. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a FLASH, a hard disk, or an optical disc, etc., including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of various embodiments of the present application.

[0119] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, and they should be covered in the scope of the claims of the present application.

[0120] Embodiment 4 is an embodiment of the present application, based on the above embodiment, provides a hierarchical and graded economic development influence factor decomposition method based on LASSO model and panel fixed effect model, in order to verify the beneficial effects of the present application, through the experiment to scientific demonstration.

[0121] Regional layer data involves 5 provinces and cities (Shanghai, Jiangsu, Zhejiang, Anhui, Jiangxi): industry layer data (manufacturing industry, power industry, service industry):

[0122] Enterprise layer data (including 5 enterprises, numbered E001, E002, E003, E004, E005):

[0123] Step 1: hierarchical data preprocessing

[0124] For the regional layer, the following is the result of processing the data of a city, and the data processing method of other provinces and cities is consistent with it.

[0125] Table 1 Regional layer data preprocessing results table

[0126]

[0127]

[0128]

[0129] For the industry layer, the following is the result after processing the manufacturing industry data, and the data processing method of other industries is consistent with it.

[0130] Table 2 Industry layer data preprocessing results table

[0131]

[0132]

[0133] For the enterprise layer, the following is the result after processing the data of enterprise E001, and the data processing method of other enterprises is consistent with it.

[0134] Table 3 Enterprise layer data preprocessing results table

[0135]

[0136]

[0137] Step 2: LASSO variable screening calculation

[0138] (1) Regional layer variable screening results

[0139] Table 4 λ selection process (cross-validation) table

[0140] Lambda value Mean square error (MSE) Number of retained variables Optimal identification 0.01 12.58 4 0.05 8.37 2 √ 0.1 9.24 1

[0141] Table 5 Coefficient path evolution table

[0142]

[0143] Table 6 Screening results table

[0144]

[0145] (2) Industry layer variable screening calculation

[0146] Table 7 λ selection process table

[0147] Lambda value MSE Number of retained variables Optimal identification 0.02 0.38 2 0.03 0.35 2 √ 0.04 0.37 1

[0148] Table 8 Coefficient path evolution table

[0149] Lambda value Industry energy efficiency index (β1) Number of technology patents (β2) 0.01 0.52 0.18 0.02 0.57 0.17 0.03 0.61 0.19 0.04 0.65 0

[0150] Table 9 Screening results table

[0151] Variable Final coefficient Whether retained Statistical characteristics Industry energy efficiency index (x1) 0.61 Yes strongly significant (p<0.001) Number of technology patents (x2) 0.19 Yes Moderately significant (p = 0.028)

[0152] (3) Enterprise layer variable screening calculation

[0153] Table 10 Lambda selection process table

[0154] Lambda value MSE Number of retained variables Optimal identification 0.06 0.29 2 0.08 0.26 1 √ 0.1 0.28 1

[0155] Table 11 Coefficient path evolution table

[0156] Lambda value Proportion of photovoltaic power generation (β1) Desulfurization equipment investment (β2) 0.05 -0.2 -0.05 0.06 -0.22 -0.01 0.08 -0.24 0.00 0.09 -0.25 0

[0157] Table 12 Screening results table

[0158] Variable Final coefficient Whether retained Reason for elimination Proportion of photovoltaic power generation (x1) -0.24 Yes Main effect is significant (p = 0.006) Desulfurization equipment investment (x2) 0 No Strongly related to x1 (r = 0.82)

[0159] The LASSO screening mechanism realizes feature compression through L1 regularization: the regional layer environmental protection investment (x3) is zeroed out at λ = 0.04 because |β| < 0.01, reflecting its weak explanatory power for economic growth; the enterprise layer desulfurization investment (x2) is excluded because it is strongly correlated with the photovoltaic proportion (x1) (r = 0.82), in line with the principle of multiple collinearity processing. There are no redundant variables at the industry layer, and all indicators are retained. After screening, the average AIC value of each layer model is reduced by 12.3%, and the VIF value is reduced to below 2.0, effectively improving the robustness of the estimate. Key findings: energy consumption indicators exhibit strong explanatory power at the regional layer (β = -0.41), and energy efficiency exhibits strong explanatory power at the industry layer (β = 0.61), while enterprise-level photovoltaic deployment (β = -0.24) reveals the micro mechanism of emission reduction.

[0160] Step 3: Fixed effects panel regression calculation.

[0161] Table 13 Full-level regression coefficient table

[0162]

[0163] The regression results based on hierarchical fixed-effect model show that the energy consumption coefficient of unit GDP at the regional level is significantly negative (β =-0.412, p = 0.002), indicating that a 0.1 tons of standard coal per million yuan reduction in energy consumption can promote GDP growth by 0.412%, verifying the core driving effect of energy saving and consumption reduction on economic development; the positive effect of clean energy proportion (β = +0.324, p = 0.018) reveals the economic value of energy structure transformation. The strong positive effect of energy efficiency index at the industry level (β = +0.612, p < 0.001) dominates the green total factor productivity improvement, and the auxiliary role of technology patent (β = +0.192, p = 0.028) confirms that innovation needs to be coordinated with process improvement. The significant negative coefficient of photovoltaic proportion at the enterprise level (β =-0.238, p = 0.006) demonstrates the emission reduction effect of clean energy, but it is structurally conflicted with the direction of clean energy at the regional level, which highlights the macro priority of energy saving policy and exposes the contradiction of micro implementation cost, and the hierarchical governance tool is urgently needed to realize the incentive coordination.

[0164] Step 4: Cross-level contribution rate calculation

[0165] Table 14 Hierarchical calculation results table

[0166]

[0167] Table 15 Cross-level comparative analysis table

[0168] Variable type Regional layer θ Industry layer θ Enterprise layer θ Direction consistency Energy consumption / energy efficiency related indicators 0.56 0.761 - Consistent (negative / positive) Clean energy related indicators 0.44 - 1 Conflict (positive / negative)

[0169] The contribution rate quantitative analysis reveals significant cross-level heterogeneity: energy consumption / energy efficiency indicators dominate at the regional level (θ = 0.56) and industry level (θ = 0.76), verifying the hierarchical penetration effect of energy saving policy; while the clean energy indicator contributes positively at the regional level (θ = 0.44) but negatively at the enterprise level (θ = 1.0), forming a structural conflict and exposing the "policy incentive-enterprise cost" contradiction. This direction deviation (regional β = +0.324 vs enterprise β =-0.238) shows the incentive mismatch between macro goals and micro implementation, and cross-level coordinated governance needs to be achieved through carbon cost compensation mechanism.

[0170] This example is based on three provinces and two cities from 2013 to 2022, through the LASSO model and the fixed effect panel regression of the collaborative analysis framework, the multi-level economic development mechanism is revealed: the regional level empirical shows that energy saving and consumption reduction (unit GDP energy consumption β =-0.412, contribution rate θ =0.56) and clean energy transformation (β =+0.324, θ =0.44) form the double engine of economic growth, where the energy consumption of 0.1 tons of standard coal per million can drive GDP growth by 0.412%, providing core empirical support for the "double carbon" strategy; industry level verifies that energy efficiency upgrading (β =+0.612, θ =0.76) is the primary driving force for green total factor productivity improvement, its contribution rate is three times that of technological innovation (β =+0.192, θ =0.24), highlighting the fundamental role of process improvement in industrial transformation; the enterprise level exposes the key structural conflict—although the deployment of photovoltaic significantly reduces carbon intensity (β =-0.238, θ =1.0), but its negative effect and the positive incentive of regional clean energy (β conflict) form a policy extrapolation paradox, reflecting the incentive fracture between macro goals and micro costs. This cross-level heterogeneity reveals the urgency of hierarchical governance: through the three-level policy system of provincial energy consumption quota trading (anchoring macro θ >0.56), industry energy efficiency patent bundled subsidies (strengthening the intermediate θ =0.76 transmission) and enterprise photovoltaic value-added tax refund (resolving the micro θ =1.0 conflict), the dynamic balance of economic growth and carbon neutralization goals is achieved.

Claims

1. A hierarchical and graded economic development impact factor decomposition method based on a LASSO model and a panel fixed effect model, characterized in that: include, Establish a multi-level panel data structure that includes regional, industry, and enterprise layers, and collect economic indicator data at each level; Standardize and preprocess the data at each level, including missing value imputation, outlier handling, unit unification, and fixed effects removal. Based on the preprocessed data, the LASSO regression model was used to determine the regularization parameters and to screen key explanatory variables in a stratified manner. Input the selected variables into the panel fixed effects regression model to quantify the marginal effect coefficients of each variable on the economic development target indicators; Calculate the contribution rate weight of key variables at each level and construct a cross-level influence path difference map.

2. The hierarchical and graded economic development impact factor decomposition method based on the LASSO model and the panel fixed effect model according to claim 1, wherein: The standardized preprocessing includes filling in missing data using time series linear interpolation or cross-sectional mean interpolation. Perform Z-score standardization on all variables; Perform logarithmic transformation on high-amplitude skewed variables; Individual fixed effects are eliminated by removing the panel unit mean.

3. The hierarchical and graded economic development impact factor decomposition method based on the LASSO model and the panel fixed effect model according to claim 2, characterized in that: The hierarchical screening of key explanatory variables includes optimizing the LASSO objective function with L1 regularization, constructing the objective function hierarchically, and determining the optimal penalty coefficient through cross-validation. The set of key variables is retained based on non-zero coefficients.

4. The hierarchical and graded economic development impact factor decomposition method based on the LASSO model and the panel fixed effect model of claim 3, characterized in that: The fixed effects regression model includes individual fixed effects terms that control for inherent characteristics of the region / industry / enterprise. The time fixed effects term controls for annual macroeconomic fluctuations; By selecting variables as explanatory variables, the regression coefficients are output.

5. The hierarchical and graded economic development impact factor decomposition method based on the LASSO model and the panel fixed effect model of claim 4, characterized in that: The calculation of the contribution rate weights of key variables at each level includes defining a variable contribution rate normalization index. wherein, denotes the variable x j Structural contribution weight at hierarchy level l for visualizing structural heterogeneity, denotes the variable x j Regression coefficient at hierarchy level l.

6. The hierarchical and graded decomposition method for economic development influencing factors based on the LASSO model and panel fixed effects model as described in claim 5, characterized in that: The construction of the cross-level influence path difference map includes comparing the contribution rate weight of the same variable at each level; Variables with conflicting directions are marked as conflicting when the regression coefficients of a variable have opposite signs at at least two levels. Visualize the distribution of variable contribution rates at each level and cross-level correlation paths.

7. The method for decomposing the influencing factors of economic development based on the LASSO model and panel fixed effects model as described in claim 6, characterized in that: The multi-level panel data structure includes a regional layer defining provincial / municipal observation units, an industry layer defining subdivided industry observation units, and an enterprise layer defining independent enterprise observation units. Among them, the regional layer defines provincial / municipal level observation units, and the indicators include GDP growth rate, energy consumption per unit of GDP, and the proportion of clean energy. The industry-level definition includes subdivided industry observation units, and the indicators include green total factor productivity and industry energy efficiency index. The enterprise level defines independent enterprise observation units, with indicators including carbon intensity and the proportion of photovoltaic power generation. Indicators at each level are associated according to preset mapping rules, and comparable variables are defined uniformly across levels.

8. A hierarchical decomposition system for economic development influencing factors based on the LASSO model and panel fixed effects model, employing the method for decomposing hierarchical economic development influencing factors based on the LASSO model and panel fixed effects model as described in any one of claims 1 to 7, characterized in that, include: Data preprocessing module, variable selection module, regression analysis module, contribution rate calculation module, and graph generation module; The data preprocessing module is used to perform missing value imputation, Z-score standardization, logarithmic transformation of high-amplitude skewed variables, and panel unit mean removal operations to eliminate individual fixed effects. The variable selection module is used to construct a LASSO objective function with L1 regularization in a hierarchical manner, determine the optimal penalty coefficient λ through cross-validation, and output a set of key variables with non-zero coefficients. The regression analysis module is used to input key variables into a panel regression model containing individual fixed effects and time fixed effects, and to calculate the marginal effects coefficients of each variable. The contribution rate calculation module is used to calculate the contribution rate weight of each level of variable according to the formula. The graph generation module is used to compare the contribution rate weights of cross-level variables, identify variables with conflicting regression coefficient signs, and generate a visual graph of differences in influence paths.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the hierarchical and graded decomposition method for economic development influencing factors based on the LASSO model and panel fixed effects model, as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the hierarchical and graded decomposition method for economic development influencing factors based on the LASSO model and panel fixed effects model, as described in any one of claims 1 to 7.