A dynamic model design method for predicting long-term proportion change of industry share

CN122779360APending Publication Date: 2026-09-18INTERNATIONAL COLLEGE OF RENMIN UNIVERSITY OF CHINA (SUZHOU RESEARCH INSTITUTE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610977463.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0003]在研究中发现,上述基于“份额恒定”假设的现有技术方法,在应用于中长期行业份额预测时,存在以下技术问题:其一,基于份额恒定的静态假设,无法识别技术革命、产业政策转型引发的行业份额结构性突变,对行业跳跃式份额变化无响应机制,预测偏差极大;其二,静态时序模型无法适配行业份额时间序列的趋势性、波动聚集等非平稳特性,预测精度随预测周期延长大幅衰减,且对所有行业采用统一模型结构,未区分行业数据特性差异;其三,未构建多维外生因素统一量化建模框架,无法将资本流向、技术创新、市场结构及非结构化政策文本信息有效纳入模型,预测结果解释力与准确性不足;其四,未针对新兴、成熟、衰退不同生命周期行业构建差异化建模与模型自适应切换机制,无法适配不同行业的数据特征与演化规律;其五,缺乏产业政策非结构化文本信息的量化与融合手段,难以捕捉政策驱动下的行业结构变化,无法适配新能源、高端制造等政策敏感型行业的长期预测需求,因此,本发明提出一种预测行业份额长期占比变化的动态模型设计方法以解决现有技术中存在的问题

Benefits of technology

[0022] 1. This invention constructs a hybrid prediction framework of static baseline + dynamic adjustment factor. The time-decay weighted static baseline ensures the basic stability and economic interpretability of share prediction. The exponential dynamic adjustment factor incorporates multi-dimensional driving factors, which not only ensures that the adjusted share is always positive, but also makes the impact of various input features on the share proportional. This effectively adapts to the non-stationary characteristics of the industry share sequence, significantly improves the response to structural changes, and reduces medium- and long-term prediction errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779360A_ABST
    Figure CN122779360A_ABST
Patent Text Reader

Abstract

The application provides a dynamic model design method for predicting long-term proportion change of industry share, relates to the technical field of macroeconomic data processing and computer-aided prediction, and comprises the following steps: collecting historical output data and multi-source characteristic data of each industry, calculating the industry share of each cycle through pretreatment, and simultaneously constructing three types of characteristic data sets of macro-level variables, industry-level variables and industrial policy text features; based on the historical sequence of the industry share, a static baseline share is calculated by using time attenuation weighting; the application constructs a hybrid prediction framework of static baseline + dynamic adjustment factor, uses the time attenuation weighted static baseline to guarantee the basic stability and economic interpretability of the share prediction, and incorporates multi-dimensional driving factors through the exponential dynamic adjustment factor, so that the adjusted share is always positive, and the impact of various input characteristics on the share is proportional, thereby effectively adapting to the non-stationary characteristics of the industry share sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of macroeconomic data processing and computer-aided forecasting technology, and in particular to a dynamic model design method for predicting long-term changes in industry share. Background Technology

[0002] In the field of macroeconomic research and industry forecasting, researchers typically use the "industry share method," which decomposes macroeconomic aggregate indicators into individual industries based on their output share, thereby predicting the future output scale of each industry. The core assumption of this method is that the share of each industry in total macroeconomic output remains relatively stable over a certain period. Therefore, the average share or a fixed proportion from historical periods can be used as a basis for estimating future shares, and the macroeconomic aggregate forecast can be proportionally decomposed into each industry accordingly.

[0003] The study found that existing technical methods based on the "constant market share" assumption have the following technical problems when applied to medium- and long-term industry market share forecasting: First, the static assumption of constant market share cannot identify structural changes in industry market share caused by technological revolutions and industrial policy transformations, and it lacks a response mechanism for sudden changes in industry market share, resulting in significant prediction bias. Second, static time series models cannot adapt to the non-stationary characteristics of industry market share time series, such as trends and fluctuation clusters. Prediction accuracy decreases significantly with the extension of the prediction period, and the use of a uniform model structure for all industries fails to distinguish the differences in industry data characteristics. Third, a unified quantitative modeling framework for multidimensional exogenous factors has not been constructed, making it impossible to incorporate... The existing technology suffers from several shortcomings. First, it fails to effectively incorporate capital flows, technological innovation, market structure, and unstructured policy text information into the model, resulting in insufficient explanatory power and accuracy in the prediction results. Second, it lacks differentiated modeling and adaptive switching mechanisms for industries with different life cycles (emerging, mature, and declining), making it unable to adapt to the data characteristics and evolutionary patterns of different industries. Third, it lacks methods for quantifying and integrating unstructured industrial policy text information, making it difficult to capture policy-driven changes in industry structure and unable to meet the long-term prediction needs of policy-sensitive industries such as new energy and high-end manufacturing. Therefore, this invention proposes a dynamic model design method for predicting long-term changes in industry share to address the problems existing in the prior art. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a dynamic model design method for predicting long-term changes in industry market share. This method constructs a hybrid prediction framework consisting of a static baseline and a dynamic adjustment factor. The time-decay weighted static baseline ensures the basic stability and economic interpretability of the market share prediction. By incorporating multi-dimensional driving factors through an exponential dynamic adjustment factor, it ensures that the adjusted market share remains positive while proportionalizing the impact of various input features on the market share. This effectively adapts to the non-stationary characteristics of the industry market share sequence, significantly improves the responsiveness to structural changes, and reduces medium- and long-term prediction errors.

[0005] To achieve the objectives of this invention, the invention is implemented through the following technical solution: a dynamic model design method for predicting long-term changes in industry market share, comprising the following steps:

[0006] S1. Collect historical output data and multi-source feature data from various industries, calculate the industry share for each period after preprocessing, and construct three types of feature datasets: macro-level variables, industry-level variables, and industrial policy text features.

[0007] S2. Calculate the static baseline share based on the historical sequence of industry share using a time decay weighted method;

[0008] S3. Based on the static baseline share, introduce three types of feature datasets to construct a multi-source feature dynamic adjustment factor, and calculate the adjusted unnormalized industry share.

[0009] S4. Identify the industry's life cycle stage based on multi-dimensional indicators, adaptively select a matching prediction model according to the life cycle stage, and predict the adjusted industry share to obtain the initial predicted share.

[0010] S5. Normalize the initial forecast share of each industry to obtain the final forecast share that meets the total share constraint.

[0011] S6. Perform rolling updates according to preset trigger conditions, iteratively optimize model parameters, and output the updated prediction results.

[0012] Further improvements are made in the following ways: In S1, the industry share is the ratio of the output of a single industry to the total output of all statistical industries; the macroeconomic variables include the growth rate of the total macroeconomy, interest rates, inflation rates, and the industrial policy intensity index; the industry-level variables include the industry's own output growth rate, industry activity index, capacity utilization rate and its rate of change, market concentration index, and cyclical fluctuation index; the industrial policy text features are structured feature vectors obtained by semantic encoding, time-weighted aggregation, and dimensionality reduction of the industrial policy texts within the cycle.

[0013] The further improvement lies in the following: In S1, the process of constructing the features of industrial policy texts is as follows: collect industrial policy texts periodically and complete text preprocessing; use a pre-trained language model to extract the semantic vector of a single text; perform exponential decay weighted averaging according to the release time to obtain a periodic aggregate vector; and obtain the policy text feature vector corresponding to the industry through dimensionality reduction; further, map it to an interpretable policy strength score through supervised learning, with positive values ​​representing encouragement strength and negative values ​​representing restriction strength.

[0014] Further improvements are made in the following aspects: In S1, the industry activity index is obtained by standardizing four sub-indicators: capital investment growth rate, new orders index, technological innovation index, and market transaction activity index, and then extracting the first principal component through principal component analysis; the cyclical fluctuation index is the rolling standard deviation of the industry share in the fourth quarter; and the market concentration index is the sum of the market share of the top four companies in the industry.

[0015] A further improvement is made in S2, where the static baseline share is calculated using an exponential weighted average method. The historical sequence of industry shares over the past N time periods is taken, and weights that decrease with increasing time intervals are assigned. The weighted sum is then used to obtain the static baseline share for the current period. The time decay coefficient of the weights ranges from 0 to 1.

[0016] Further improvements are made in S3, where the multi-source feature dynamic adjustment factor is constructed using an exponential adjustment equation. The adjusted unnormalized industry share is the product of the static baseline share and the index term. The index term is a linear combination of industry-level variables, macro-level variables, and industrial policy text features with the corresponding industry adjustment coefficients. The adjustment coefficients are calibrated by regression of historical data using the maximum likelihood estimation method or the Bayesian estimation method.

[0017] Further improvements are made in S4, which constructs a life cycle quantitative index LCI-Score based on six categories of indicators: industry output growth rate, industry activity index, capacity utilization rate and its change rate, market concentration index, cyclical fluctuation index, and policy intensity score. After standardization, the index is mapped to the range of [-1,1]. The weight coefficients are determined by logistic regression calibration of historical samples or by the analytic hierarchy process.

[0018] Further improvements are made in S4, where industries are divided into three life cycle stages—emerging, mature, and declining—based on LCI-Score values ​​and continuous trends. When making a determination, the corresponding threshold conditions are met and the preset number of cycles is maintained continuously. Correspondingly, mature industries are predicted using the ARIMA model, emerging industries are predicted using the Long Short-Term Memory (LSTM) network model, and declining industries are predicted using a state-space model combined with Kalman filtering.

[0019] Further improvements are made in the following: In S5, the normalization process adopts proportional normalization or Softmax normalization with temperature parameters, so that the sum of the final predicted shares of each industry in the same period is always 1; at the same time, a minimum share lower limit is set. When the share of an industry after normalization is lower than the lower limit, it is corrected to the lower limit value, and the difference is deducted from the shares of the remaining industries proportionally before being normalized again.

[0020] Further improvements are made in S6, where the triggering conditions for rolling updates include at least one of fixed-time-period timed triggering and event-driven triggering when the industry lifecycle label changes; after triggering, all feature data are updated, the adjustment coefficient is recalibrated, and the share calculation, prediction, and normalization process is repeated.

[0021] The beneficial effects of this invention are as follows:

[0022] 1. This invention constructs a hybrid prediction framework of static baseline + dynamic adjustment factor. The time-decay weighted static baseline ensures the basic stability and economic interpretability of share prediction. The exponential dynamic adjustment factor incorporates multi-dimensional driving factors, which not only ensures that the adjusted share is always positive, but also makes the impact of various input features on the share proportional. This effectively adapts to the non-stationary characteristics of the industry share sequence, significantly improves the response to structural changes, and reduces medium- and long-term prediction errors.

[0023] 2. This invention establishes a quantitative fusion path for unstructured industrial policy texts. By extracting semantic features through a pre-trained language model and combining time decay weighted aggregation and dimensionality reduction, policy guidance information is transformed into computable structured features, which are further mapped into interpretable policy intensity scores. This makes up for the shortcomings of existing technologies in integrating policy text information, strengthens the model's ability to perceive policy-driven industry structural changes, and especially improves the prediction accuracy of policy-sensitive industries.

[0024] 3. This invention constructs an LCI-Score lifecycle quantification system covering six core indicators, which can automatically identify the emerging, mature, and declining lifecycle stages of an industry and match them with three different prediction models: ARIMA, LSTM, and state-space Kalman filter. This achieves accurate adaptation of the model structure to the characteristics of industry data, overcomes the accuracy loss caused by "one-size-fits-all" modeling, and can simultaneously take into account the prediction needs of industries at different development stages.

[0025] 4. This invention sets up a rolling update mechanism that combines timed triggering with lifecycle state change event triggering. This mechanism can periodically iterate model parameters to adapt to the gradual evolution of the industry structure, and can also adjust the model architecture in real time when the industry development stage leaps, effectively avoiding structural drift of the model and ensuring the continuity and robustness of long-term prediction results.

[0026] 5. This invention strictly ensures the economic constraint that the sum of the shares of all industries within the same period is always 1 through proportional normalization or Softmax normalization mechanism with temperature coefficient; at the same time, it sets a minimum share lower limit to avoid numerical collapse in very small industries during the normalization process, ensuring that the prediction results conform to basic economic laws and improving the usability and reliability of the prediction results in actual decision-making scenarios. Attached Figure Description

[0027] Figure 1 This is an overall flowchart of the dynamic prediction method for long-term changes in industry market share according to the present invention.

[0028] Figure 2 This is a schematic diagram of the system functional modules for implementing the dynamic prediction method of the present invention;

[0029] Figure 3 This is a flowchart illustrating step S4 of the present invention. Detailed Implementation

[0030] To enhance understanding of the present invention, the present invention will be further described in detail below with reference to embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the scope of protection of the present invention.

[0031] Example 1

[0032] according to Figure 1 , 2 As shown in Figure 3, this embodiment proposes a dynamic model design method for predicting long-term changes in industry market share, including the following steps:

[0033] Step S1: Data Acquisition and Preprocessing

[0034] Collect historical output data from various industries and calculate the industry share of each industry during historical periods:

[0035]

[0036] in, For industry code, For discrete time (preferably quarterly, but can also be set to monthly or annual), For the industry In time The output or added value, At the same time The sum of output from all industries included in the statistics (approximately equal to total macroeconomic output, such as GDP).

[0037] Simultaneously, the following three types of data were collected and corresponding feature construction was completed:

[0038] (1) Macro-level variables This includes macroeconomic growth rate, interest rate, inflation rate, and industrial policy intensity index, etc.

[0039] (2) Industry-level variables These include, but are not limited to, capital investment growth rate, technological innovation indicators, capacity utilization rate, market competition structure indicators, and cyclical fluctuation indicators. The preferred calculation method is as follows:

[0040] Industry's own output growth rate:

[0041]

[0042] Industry Activity Index Collect capital investment growth rates separately New orders or New order sub-item Technological innovation indicators (such as the growth rate of patent applications or authorizations) Market trading activity (such as total industry turnover or turnover rate) The four sub-indicators were standardized using z-scores, and principal component analysis (PCA) was used to extract the first principal component as the industry activity index.

[0043]

[0044] As an alternative implementation method, when the requirement for model interpretability is higher than the requirement for accuracy, a weighted summation of the four standardized sub-indices can also be used to obtain the model. .

[0045] Capacity utilization rate and its rate of change:

[0046]

[0047]

[0048] When nominal capacity data is missing, it is preferable to use the available capacity data disclosed in the company's financial statements or the estimate from the industry association as a substitute.

[0049] Market concentration index (the sum of the market share of the top 4 companies) (For example)

[0050]

[0051] Indicates industry Inner Market share of large enterprises; when it is difficult to obtain data for all enterprises in the industry, it is preferable to use weighted data of representative enterprises in the industry for estimation.

[0052] Cyclical volatility indicator (4-quarter rolling standard deviation):

[0053]

[0054] This represents the average industry share over the most recent four quarters.

[0055] (3) Characteristics of Industrial Policy Texts. Central and local industrial policy texts (including policy documents, announcements, and news) related to various industries are collected at a preset cycle (preferably quarterly). After preprocessing such as removing webpage tags, sentence segmentation, word segmentation, and removing stop words, they are encoded using any of the following methods:

[0056] The first implementation method (dictionary weighted method): establish a sentiment dictionary containing keywords such as "encouragement", "restriction", "subsidy" and "punishment", and sum the word frequency weights for each policy text to obtain the sentiment score of the text;

[0057] The second implementation method (pre-trained language model method, preferred): Use a pre-trained language model (such as BERT, RoBERTa or Chinese BERT variant) to extract the semantic vector corresponding to the classification tag bit (CLS) for each policy text;

[0058] For multiple policy text vectors within the same statistical period, an exponentially decaying weighted average is applied according to the release time to obtain the aggregated policy text vector for that period, thus transforming unstructured policy information into structured predictive variables.

[0059]

[0060]

[0061] For the first The semantic vector corresponding to each policy text. This represents the time interval between the current statistical period and the current text. This is the time decay coefficient (preferably 0.95, with lower weighting for policy texts published a long time ago).

[0062] The aggregated vectors for each statistical period are further reduced to 5 to 20 dimensions using methods such as principal component analysis or UMAP, resulting in industry-specific vectors. ,time The policy text feature vector, denoted as ;

[0063] Furthermore, a supervised learning method (using the actual industry growth during the period corresponding to the historical policy text as labels) can be used to train the mapping model, which will... This is mapped to an interpretable scalar—the policy strength score. (Incentive intensity is positive, restriction intensity is negative), this scalar is used to determine the industry life cycle in the subsequent step S4.

[0064] Finally, missing value processing, time alignment, and standardization are performed on all the above variables to obtain a unified feature dataset.

[0065] Step S2: Static baseline calculation

[0066] Based on the past One time period (preferred) The static baseline share is calculated using an index-weighted average method based on the historical industry share series (corresponding to the corresponding number of cycles over the past 5 years).

[0067]

[0068]

[0069] The time decay coefficient ( This is used to decrease the weight as the time interval increases, thereby reflecting the relative stability of the recent industry structure while maintaining the interpretability of the share.

[0070] Step S3: Modeling of dynamic adjustment factors for multi-source features

[0071] Based on the static baseline share, industry-level variables, macroeconomic variables, and policy text feature vectors are introduced to construct the following exponential dynamic adjustment equation, yielding the adjusted (unnormalized) industry share:

[0072]

[0073] in, Industry The adjustment coefficient vectors of the corresponding industry-level variables, macro-level variables, and policy text feature vectors are preferably determined by regression calibration of historical data using maximum likelihood estimation or Bayesian estimation. The technical advantage of employing an exponential function structure is that it ensures, on the one hand, the adjusted share... It is always positive and requires no additional constraints; on the other hand, it makes the impact of various input features on the share proportional, which facilitates the compatibility of input features of different scales and the economic interpretation of the coefficients.

[0074] As a preferred implementation method, It can also include market concentration indicators. and cyclical fluctuation indicators This is to improve the sensitivity of dynamic adjustment factors to changes in market structure.

[0075] Step S4: Industry Lifecycle Identification and Adaptive Model Selection

[0076] This step is one of the core steps of this invention. Its technical purpose is to automatically select a prediction model that matches the data characteristics of the current life cycle stage of the industry. Specifically, it includes the following sub-steps:

[0077] Sub-step S401: Based on the six types of indicators obtained in step S1, namely industry output growth rate Industry Activity Index Capacity utilization rate indicators and its rate of change, market concentration indicators Cyclical fluctuation indicators and policy strength score Construct a unified lifecycle quantitative index (LCI-Score):

[0078]

[0079]

[0080] The preset weighting coefficients are preferably determined by logistic regression or analytic hierarchy process (AHP) on historical industry lifecycle label samples. The obtained scores are then standardized and mapped to the interval [-1, 1].

[0081] Sub-step S402: Based on the LCI-Score value and its continuous trend, the industry is divided into three life cycle stages: emerging, mature, and declining. The criteria for determining each stage (meeting at least three of the following conditions is considered to meet the criteria for that stage) and the corresponding LCI-Score thresholds are shown in the table below:

[0082]

[0083] The aforementioned thresholds were determined through historical sample backtesting and grid search to achieve optimal accuracy in lifecycle classification.

[0084] When an industry fails to meet any of the above criteria in the current cycle (e.g., during a transitional period of phase change), it is preferable to maintain the lifecycle label and corresponding prediction model of the previous cycle to avoid frequent model switching due to short-term indicator fluctuations.

[0085] Based on the lifecycle labels determined in sub-step S402, the corresponding prediction model pairs are adaptively selected. Perform smoothing or incremental forecasting:

[0086] For mature industries, the ARIMA(p,d,q) model is used to predict the adjusted share series. Specifically, the most recent... A time period (preferably T=20, i.e., 20 quarters) The historical sequence is used to traverse all parameter combinations within a predefined parameter search space (preferably with p, d, and q each taking values ​​in the range {0, 1, 2}). An ARIMA model is fitted to each combination, and the corresponding Akaike Information Criterion (AIC) value is calculated. The parameter combination with the smallest AIC value is selected as the optimal model structure, and this model is used to predict future... Forecasting is performed over a time period (preferably H=8, i.e., 8 quarters), and the forecast result is denoted as... Preferably, a smoothing process or a trend correction function may be further applied to the predicted sequence. ,Right now This is to ensure the continuity and stability of the prediction results and eliminate high-frequency noise.

[0087] For emerging industries, a Long Short-Term Memory (LSTM) network model is used to dynamically predict the adjusted share increment. Specifically, the share sequence over the past 12 time periods is used. As the model input sequence (input layer dimension 12×1), after passing through an LSTM layer with an optimal number of hidden units of 32 and a random dropout layer with an optimal dropout ratio of 0.2, the fully connected layer outputs the dynamic incremental adjustment amount for the next cycle. Thus, the predicted share is obtained:

[0088]

[0089] The model is preferably trained using the Adam optimizer, with a learning rate preferably set to 0.001 and a loss function preferably using mean squared error (MSE). An early stopping strategy is also used to prevent overfitting.

[0090] For declining industries, a combination of state-space modeling and Kalman filtering is used for prediction to more accurately characterize the data characteristics of a long-term downward trend coupled with slow fluctuations. Specifically, the following state equation and observation equation are constructed:

[0091]

[0092]

[0093] in, This reflects the true (potential) state of industry share. To reflect process noise that reflects long-term trend changes, To reflect observation noise due to short-term observation errors, the variance parameters of both are preferably estimated using the expectation-maximization (EM) algorithm with maximum likelihood estimation. The state update formula based on Kalman filtering is:

[0094]

[0095]

[0096] The covariance of the state estimate, To generate the process noise covariance, the future values ​​are progressively generated by rolling forward using the above recursive formula. Predicted values ​​for each period .

[0097] It should be noted that the above three types of models are designed for the data characteristics of industries in three different life cycle stages. Those skilled in the art can replace them with other models with similar functions without departing from the technical concept of this invention (for example, replacing ARIMA with an exponential smoothing model, replacing LSTM with a gated recurrent unit (GRU) model, or replacing Kalman filtering with particle filtering or other state estimation methods). Such replacements should all fall within the protection scope of this invention.

[0098] Step S5: Share normalization processing

[0099] Due to the industry forecast share output in step S4 The economic constraint that "the sum of the shares of all industries is 1" may not be satisfied, therefore normalization is required. The preferred method is proportional normalization (or Softmax normalization):

[0100]

[0101] When it is necessary to further enhance the stability of the normalization results, a temperature parameter T can be introduced, and a Softmax function with a temperature coefficient can be used:

[0102]

[0103] Furthermore, to prevent numerical collapse (i.e., the share approaching 0 leading to abnormal subsequent calculations) in extremely small industries during the normalization process, it is preferable to set a minimum share threshold. When the normalized market share of a certain industry is lower than At that time, its share was set as The resulting difference in market share will be deducted proportionally from the market share of other industries and then normalized again to ensure that the sum of the market share of each industry remains constant at 1 after adjustment.

[0104] Step S6: Rolling Update

[0105] Rolling updates are performed at a preset frequency (preferably quarterly or annually). Triggering conditions include at least one of the following two methods: (i) timed triggering at a fixed time period; (ii) event-driven triggering when step S4 determines that the life cycle label of an industry has changed (e.g., from a declining industry to a mature industry). After triggering the rolling update, the system automatically performs the following operations: updates industry characteristic data, policy text data, and macroeconomic variable data; and adjusts the adjustment coefficients. The model is retrained or updated online using a Bayesian method; steps S2 to S5 are returned for recalculation; and the industry share prediction result for the next period is output. This rolling update mechanism enables the prediction model to respond promptly to gradual or sudden changes in industry structure, avoiding structural drift in the model.

[0106] like Figure 1 As shown, the dynamic prediction method for long-term changes in industry share provided by the present invention is executed by computer equipment in the following steps: S1 data acquisition and preprocessing, S2 static baseline calculation, S3 multi-source feature dynamic adjustment factor modeling, S4 industry life cycle identification and model adaptive selection, S5 share normalization processing, and S6 rolling update. The specific implementation methods of each step have been described in detail in the foregoing invention content section and will not be repeated here.

[0107] like Figure 2 As shown, the system implementing the above method can be functionally divided into a data acquisition and preprocessing module, a static baseline calculation module, a dynamic adjustment factor calculation module (including industry-level feature sub-modules, macro-level feature sub-modules, and policy text feature extraction sub-modules), an industry life cycle identification module, a model adaptive selection and prediction module (including ARIMA prediction units, LSTM prediction units, and state-space / Kalman filter prediction units), a share normalization module, and a rolling update control module. Each module is further divided into... Figure 2 The data flow shown is processed sequentially. When the triggering conditions are met, the rolling update control module feeds back the updated data to the data acquisition and preprocessing module and the static baseline calculation module, driving the system to re-execute the subsequent calculation process.

[0108] like Figure 3As shown, the specific execution flow of step S4 is as follows: First, sub-step S401 is executed to calculate six indicators: industry output growth rate G(i,t), industry activity index A(i,t), capacity utilization rate CU(i,t), market concentration index CR4(i,t), cycle fluctuation index σ4Q(i,t), and policy intensity score P(i,t); second, sub-step S402 is executed to calculate the LCI-Score and complete the standardization; then, sub-steps S403, S405, and S407 are executed sequentially to determine whether an industry meets the corresponding life cycle determination criteria in the order of emerging, mature, and declining; if the emerging industry determination criteria are met (sub-step S403 determines "yes"), then the process is executed. In sub-step S404, the industry is marked as an emerging industry and the LSTM prediction unit is called for prediction. If the emerging industry condition is not met but the mature industry condition is met (sub-step S405 judges "yes"), then sub-step S406 is executed, the industry is marked as a mature industry and the ARIMA prediction unit is called for prediction. If the mature industry condition is not met but the declining industry condition is met (sub-step S407 judges "yes"), then sub-step S408 is executed, the industry is marked as a declining industry and the state space / Kalman filter prediction unit is called for prediction. If none of the above three conditions are met, then sub-step S409 is executed, maintaining the life cycle label of the industry in the previous cycle and the corresponding prediction model.

[0109] Based on the above description and in conjunction with publicly available computational tools (the ARIMA prediction unit can be implemented based on the time series analysis module in a statistical computing software package; the LSTM prediction unit can be built and trained based on a general deep learning framework; and the state space / Kalman filter prediction unit can be programmed based on the standard recursive formula of state space modeling and Kalman filtering), the complete technical solution of the present invention can be implemented.

[0110] Example 2

[0111] according to Figure 1 , 2 As shown in Figure 3, this embodiment proposes a dynamic model design method for predicting long-term changes in industry share. Taking the prediction of industry share in the chemical industry as an example, it illustrates the specific application process of the above method in a mature industry. The chemical industry is a typical cyclical and mature industry. Its industry share usually changes cyclically with macroeconomic fluctuations, with a relatively stable long-term trend and a slower rate of structural change compared to emerging industries. This embodiment achieves dynamic prediction of the chemical industry share through "static baseline + exponential dynamic adjustment factor + ARIMA prediction model," corresponding to the implementation method described in sub-step (a) of step S4. The specific calculation process is as follows:

[0112] (1) Call the adjusted share sequence obtained after processing in steps S2 and S3. (In the code, it is represented by a pandas time series object, with the index being the time point of each quarter).

[0113] (2) Select the most recent T quarters (T is preferably 20 in this embodiment) of the historical sequence as the model fitting window, and optionally perform a stationarity test (such as ADF test) on the sequence to help judge the sequence characteristics and provide a reference for model selection;

[0114] (3) Traverse all 27 parameter combinations in the parameter search space p, d, q∈{0,1,2}, fit the ARIMA(p,d,q) model respectively, calculate the AIC value of each model, and select the parameter combination with the smallest AIC value as the optimal model structure.

[0115] (4) Using the optimal ARIMA model determined in step S03, predict the next H quarters (H is preferably 8 in this embodiment) and output the predicted mean sequence and the corresponding confidence interval.

[0116] After execution, the output will include the optimal ARIMA model order selected using the AIC criterion, the corresponding AIC value, the stationarity test p-value, and the predicted mean series and confidence intervals for the next 8 quarters (H=8). This predicted mean series is... Figure 1 Input of step S5 After the proportional normalization process in step S5, the predicted share of the industry shares, together with the predicted shares of other industries, constitutes the final industry share prediction result Sfinal(i,t) that satisfies the constraint that "the sum of the shares is 1".

[0117] Since the market share time series of the chemical industry typically exhibits a stable trend or tends to stabilize after differencing, the ARIMA prediction method described in this embodiment can effectively eliminate high-frequency noise while maintaining low model complexity, and capture the medium- to long-term stable trend of the chemical industry's market share as it fluctuates with the macroeconomic cycle. This verifies the feasibility and effectiveness of the method described in this invention in mature and cyclical industries.

[0118] Example 3

[0119] according to Figure 1 , 2 As shown in Figure 3, this embodiment proposes a dynamic model design method for predicting long-term changes in the industry share. The new energy vehicle industry is a typical emerging industry, and its industry share is usually affected by multiple factors such as technological innovation, industrial policies and changes in market demand. It is characterized by rapid growth, large fluctuations and frequent structural changes.

[0120] After calculating indicators such as industry output growth rate, industry activity index, capacity utilization rate change rate and policy intensity score in step S401, the LCI-Score value is calculated.

[0121] Within a certain statistical period: industry output growth rate G = 15.2%, industry activity A increased by 8.7%, capacity utilization rate change rate ΔCU = 5.4%, and policy intensity score P = 0.38. According to the life cycle determination rules in step S4: more than 3 conditions for emerging industry determination are met. Simultaneously, LCI-Score = 0.46, and the conditions are met continuously for 10 statistical periods; therefore, it is determined to be an emerging industry.

[0122] The LSTM prediction unit is then used for prediction, with the adjusted industry share sequence for the most recent 12 quarters as the model input: The input dimension is 12×1, the number of hidden units in the LSTM layer is set to 32, the Dropout ratio is set to 0.2, and the model is trained using the Adam optimizer with a learning rate of 0.001. The model outputs the industry share increment for the next quarter. The final result is:

[0123]

[0124] The prediction results are then normalized in step S5.

[0125] Since the new energy vehicle industry has obvious non-linear growth characteristics, the LSTM model can effectively capture the complex time series relationships in the process of rapid expansion of industry share, thereby improving prediction accuracy.

[0126] Example 4

[0127] according to Figure 1 , 2 As shown in Figure 3, this embodiment proposes a dynamic model design method for predicting long-term changes in the industry share. The coal industry is a typical declining industry, and its industry share is affected by factors such as energy structure transformation, environmental protection policies and declining market demand, showing a long-term downward trend.

[0128] After calculating the relevant indicators in step S401: industry output growth rate G = -5.6%, industry activity A decreased by 6.2%, capacity utilization rate change rate ΔCU = -3.5%, and policy intensity score P = -0.31. According to the life cycle determination rules in step S4: more than 3 conditions for a declining industry are met. Simultaneously, the LCI-Score = -0.42, and this condition has been met for 10 consecutive statistical periods; therefore, the industry is determined to be in decline.

[0129] Then, the state-space prediction unit is invoked for prediction. Let the true state of the industry share be... The observed value is Establish the state equation and observation equation:

[0130]

[0131]

[0132] Among them, process noise, which reflects long-term trend changes, To reflect observation noise due to short-term observation errors, a Kalman filter recursive algorithm is used to estimate future industry share status. Through the state prediction and observation update process, the predicted industry share values ​​for the next H quarters are obtained. The prediction results are then normalized in step S5.

[0133] Since the market share of the coal industry typically exhibits a long-term downward trend accompanied by short-term fluctuations, using a state-space model and Kalman filtering can effectively separate long-term trends from random noise, thereby improving the stability and robustness of the prediction results.

[0134] Verification example:

[0135] To further verify the quantitative technical effect of the present invention compared with the prior art, and to verify the independent contributions of the policy text feature module and the life cycle identification and model adaptive switching module in the present invention, the following comparative experiments and ablation experiments were carried out based on historical data.

[0136] Comparative experiment: Quantization error comparison with existing technologies

[0137] Experimental Setup: Experimental data are sourced from published industry value-added data, Wind database industry indicator data, and publicly available industrial policy texts. The sample covers 60 quarterly data periods from 2010 to 2024, selecting 15 sub-sectors including manufacturing, energy, and information technology as research objects. Five sub-sectors are categorized as emerging, mature, and declining. Historical quarterly data is used as the data source, divided chronologically into training and test sets. A rolling multi-step forecasting method is employed to backtest and predict industry share for the next H quarters (H=8). The evaluation metric used is the Mean Absolute Percentage Error (MAPE), defined as follows:

[0138]

[0139] This represents the actual observed industry share for the corresponding quarter. This refers to the predicted value output by the corresponding method.

[0140] The comparison methods are set as follows: Method A is the fixed proportion method in the prior art (i.e., the method described in the background section, which uses the average share of historical periods as a fixed estimate of the future share); Method B is the single ARIMA model method in the prior art (it does not distinguish between industry life cycle stages, and uniformly applies a set of ARIMA(p,d,q) models to the share sequences of all industries, and does not include the multi-source dynamic adjustment and life cycle adaptive switching mechanism described in steps S1 to S4 of this invention); Method C is the complete technical solution of this invention (i.e., steps S1 to S6). The MAPE comparison results of the three methods in industries at different life cycle stages are shown in the table below:

[0141]

[0142] As can be seen from the table above, the method C of the present invention has achieved better prediction results than the existing methods A and B in all three types of industries at different life cycle stages.

[0143] The experimental results above demonstrate that the prediction framework constructed in this invention, consisting of "static baseline + dynamic adjustment factor + life cycle adaptive model selection," can effectively improve the accuracy of medium- and long-term industry share predictions. Specifically, the prediction capabilities of traditional fixed-proportion methods and unified ARIMA models are limited for emerging and declining industries due to their strong non-stationarity and structural change characteristics. This invention, by introducing an industry life cycle identification mechanism and employing prediction models that match the industry's development stage, further reduces prediction errors.

[0144] For mature industries, where market share changes are relatively stable, the ARIMA model itself already possesses good predictive capabilities. Therefore, the improvement of this invention compared to a single ARIMA model is relatively limited. However, experimental results still show that by introducing multi-source information such as macroeconomic variables, industry characteristics, and policy text features, this invention can further improve prediction accuracy in mature industry scenarios, verifying the effectiveness of the multi-source dynamic adjustment mechanism.

[0145] Based on the full sample results, the overall MAPE of the method of this invention is 9.0%, which is 46.4% lower than the fixed ratio method and 27.4% lower than the single ARIMA model. This shows that the invention can not only be applied to specific industry types, but also maintain stable predictive performance across industries at different life cycle stages.

[0146] Ablation Experiment: Validation of the Independent Contributions of the Policy Text Feature Module and the Lifecycle Recognition Switching Module

[0147] To verify the independent contributions of the two innovations in step S3 (policy text features) and step S4 (lifecycle identification and model adaptive switching mechanism), two ablation variants were set up using the same dataset and evaluation metrics as the comparative experiment above:

[0148] Variant D (Removing Policy Text Features): Removing the policy text feature vector in the dynamic adjustment equation of step S3. The corresponding term, i.e., the adjustment equation, simplifies to The remaining steps (including step S4 lifecycle identification and model adaptive switching) remain unchanged;

[0149] Variant E (removing life cycle identification and adaptive model switching): skips the determination of industry life cycle stage in step S4, and uses the ARIMA model for prediction for all industries (i.e., retains the multi-source dynamic adjustment factors and policy text features from steps S1 to S3, but does not distinguish between emerging, mature, and declining stages and adaptively switches between LSTM and state-space Kalman filter models).

[0150] The following table shows the MAPE comparison results of variants D and E with the complete method C of this invention in three types of industries at different life cycle stages:

[0151]

[0152] The above ablation experiment results further verify the independent contributions of the policy text feature module and the life cycle identification and model adaptive switching module to the prediction performance.

[0153] First, after removing policy text features (Variant D), the MAPE of emerging industries increased from 10.2% to 11.9%, a rise of 15.7%; mature industries increased from 5.0% to 5.5%, a rise of 10.0%; and declining industries increased from 11.8% to 12.8%, a rise of 8.5%. This result indicates that industrial policy information can provide additional explanatory power for changes in industry share, especially in emerging industries that are heavily influenced by policy, where its contribution is more significant.

[0154] Secondly, after removing the life cycle identification and model adaptive switching mechanism (Variant E), the MAPE for emerging industries increased from 10.2% to 12.7%, a rise of 24.5%; for declining industries, it increased from 11.8% to 14.1%, a rise of 19.5%; while for mature industries, it only increased from 5.0% to 5.2%, a rise of 4.0%. This result indicates that different industry life cycle stages have different data distribution characteristics and evolutionary patterns. Using a unified prediction model is difficult to meet the prediction needs of various industries, while automatically selecting a matching prediction model based on the life cycle stage can further improve prediction performance.

[0155] In summary, the policy text feature module primarily enhances the model's ability to perceive external driving factors, while the lifecycle identification and model adaptive switching mechanism primarily enhance the model's adaptability to industry heterogeneity. These two mechanisms improve prediction accuracy from two independent dimensions: "quantitative inclusion of multi-dimensional driving factors" and "matching model structure with data characteristics," respectively, together forming the important technical foundation for the improved prediction accuracy achieved in this invention.

[0156] Further explanation of the technical contribution of this invention: the technical obstacles in constructing a unified LCI-Score index system and overcoming model switching barriers.

[0157] The ARIMA, LSTM, and Kalman filter models used in this invention are all publicly available prediction algorithms. However, the technical contribution of this invention is not a simple selection or superposition of the above-mentioned existing algorithms, but rather to solve the following unique technical obstacles faced when integrating the above heterogeneous algorithms into a unified prediction system that can automatically identify industry status and switch seamlessly:

[0158] First, there are obstacles to the unified quantification and integration of heterogeneous indicators. The LCI-Score constructed in step S4 of this invention needs to incorporate multiple types of indicators with different dimensions, statistical characteristics, and update frequencies—including continuous macroeconomic and industry economic indicators (such as industry output growth rate and capacity utilization rate) and discretized, weakly labeled policy intensity scores derived from unstructured policy texts and mapped through language models—into a comparable and weighted quantitative scoring system. Furthermore, it requires determining a combination of weight coefficients and judgment thresholds that can simultaneously adapt to different industries and different development stages without causing misjudgments (such as the judgment rule of "meeting at least three conditions and continuously meeting them for 10 cycles"). The construction of this indicator system involves the standardized integration of multi-source heterogeneous data, weight calibration, and robust threshold design, which cannot be directly achieved by any single algorithm or simple weighted averaging in existing technologies.

[0159] Secondly, there is the continuity barrier in switching between model categories. ARIMA models are based on linear parameterization statistical assumptions, LSTM models are based on data-driven nonlinear learning mappings, and Kalman filters are based on latent state-space assumptions. These three types of models differ fundamentally in their mathematical structure, parameter space, and requirements for historical data length. When an industry's life cycle stage changes (e.g., from a declining industry to a mature industry), triggering a switch from a state-space / Kalman filter unit to an ARIMA unit, directly switching models without processing can easily lead to abrupt discontinuities in prediction results at the switching point, impairing the usability of the prediction results. This invention achieves a smooth transition in the model switching process through the robust design in step S4 ("meeting the judgment condition and continuously satisfying it for several cycles to confirm the life cycle label change") and the gradual update method of the adjustment coefficient in the rolling update mechanism in step S6. This solves the continuity problem of switching in multi-model hybrid systems at the engineering implementation level. This problem is not something that existing ARIMA, LSTM, or Kalman filter algorithms can solve themselves; rather, it is a unique technical solution proposed by this invention for the specific application scenario of industry share.

[0160] In summary, for the specific technical scenario of long-term industry share prediction, a holistic technical solution has been invented that can uniformly quantify multi-source heterogeneous information (especially unstructured policy text information) into computable indicators, and automatically identify the industry life cycle stage and smoothly switch matching prediction models accordingly. This is not a simple combination or conventional selection of existing public algorithms, but rather overcomes the aforementioned technical obstacles that have long existed in the existing technology and achieves the technical effect of improved prediction accuracy as verified by the experimental data above.

[0161] This invention constructs a hybrid prediction framework combining a static baseline and dynamic adjustment factors. The time-decay weighted static baseline ensures the fundamental stability and economic interpretability of share predictions. The exponential dynamic adjustment factor incorporates multi-dimensional driving factors, guaranteeing that the adjusted share remains positive while proportionalizing the impact of various input features on the share. This effectively adapts to the non-stationary nature of industry share sequences, significantly improving the responsiveness to structural changes and reducing medium- to long-term prediction errors. Furthermore, this invention establishes a quantitative fusion path for unstructured industrial policy texts. Semantic features are extracted through a pre-trained language model, and combined with time-decay weighted aggregation and dimensionality reduction, policy guidance information is transformed into computable structured features, further mapped to interpretable policy intensity scores. This overcomes the shortcomings of existing technologies in integrating policy text information, strengthens the model's ability to perceive policy-driven industry structural changes, and particularly improves the prediction accuracy for policy-sensitive industries. This invention constructs an LCI-Score lifecycle quantification system covering six core indicators. It can automatically identify the emerging, mature, and declining lifecycle stages of an industry and match them with three differentiated prediction models: ARIMA, LSTM, and state-space Kalman filter. This achieves precise adaptation of the model structure to the characteristics of industry data, overcoming the accuracy loss caused by "one-size-fits-all" modeling and simultaneously addressing the prediction needs of industries at different development stages. This invention employs a rolling update mechanism combining timed triggering and lifecycle state change event triggering. This allows for periodic iteration of model parameters to adapt to the gradual evolution of the industry structure and real-time adjustment of the model architecture during leaps in industry development stages, effectively preventing structural drift and ensuring the continuity and robustness of long-term prediction results. This invention strictly ensures the economic constraint that the sum of the market shares of all industries within the same cycle is always 1 through proportional normalization or Softmax normalization with a temperature coefficient. Simultaneously, a minimum market share lower limit is set to prevent numerical collapse in extremely small industries during normalization, ensuring that the prediction results conform to basic economic laws and improving the usability and reliability of the prediction results in practical decision-making scenarios.

[0162] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A dynamic model design method for predicting long-term changes in industry market share, characterized in that, Includes the following steps: S1. Collect historical output data and multi-source feature data from various industries, calculate the industry share for each period after preprocessing, and construct three types of feature datasets: macro-level variables, industry-level variables, and industrial policy text features. S2. Calculate the static baseline share based on the historical sequence of industry share using a time decay weighted method; S3. Based on the static baseline share, introduce three types of feature datasets to construct a multi-source feature dynamic adjustment factor, and calculate the adjusted unnormalized industry share. S4. Identify the industry's life cycle stage based on multi-dimensional indicators, adaptively select a matching prediction model according to the life cycle stage, and predict the adjusted industry share to obtain the initial predicted share. S5. Normalize the initial forecast share of each industry to obtain the final forecast share that meets the total share constraint. S6. Perform rolling updates according to preset trigger conditions, iteratively optimize model parameters, and output the updated prediction results.

2. The dynamic model design method for predicting long-term changes in industry market share according to claim 1, characterized in that: In S1, the industry share is the ratio of a single industry's output to the total output of all statistical industries; the macroeconomic variables include the growth rate of the overall macroeconomy, interest rates, inflation rates, and the industrial policy intensity index; the industry-level variables include the industry's own output growth rate, industry activity index, capacity utilization rate and its rate of change, market concentration index, and cyclical fluctuation index; the industrial policy text features are structured feature vectors obtained by semantically encoding, time-weighted aggregation, and dimensionality reduction of the industrial policy texts within the cycle.

3. The dynamic model design method for predicting long-term changes in industry share as described in claim 2, characterized in that: In S1, the process of constructing the features of industrial policy texts is as follows: collect industrial policy texts periodically and complete text preprocessing; use a pre-trained language model to extract the semantic vector of a single text; perform exponential decay weighted averaging according to the release time to obtain a periodic aggregate vector; and obtain the policy text feature vector corresponding to the industry through dimensionality reduction processing; further, map it to an interpretable policy strength score through supervised learning, with positive values ​​representing encouragement strength and negative values ​​representing restriction strength.

4. The dynamic model design method for predicting long-term changes in industry market share according to claim 3, characterized in that: In S1, the industry activity index is obtained by standardizing four sub-indicators: capital investment growth rate, new orders index, technological innovation index, and market transaction activity index, and then extracting the first principal component through principal component analysis; the cyclical fluctuation index is the rolling standard deviation of the industry share in the fourth quarter; and the market concentration index is the sum of the market share of the top four companies in the industry.

5. The dynamic model design method for predicting long-term changes in industry market share according to claim 1, characterized in that: In S2, the static baseline share is calculated using an exponential weighted average method. The historical sequence of industry shares over the past N time periods is taken, and weights that decrease with increasing time intervals are assigned. The weighted sum is then used to obtain the static baseline share for the current period. The time decay coefficient of the weights ranges from 0 to 1.

6. The dynamic model design method for predicting long-term changes in industry market share according to claim 1, characterized in that: In S3, the multi-source feature dynamic adjustment factor is constructed using an exponential adjustment equation. The adjusted unnormalized industry share is the product of the static baseline share and the index term. The index term is a linear combination of industry-level variables, macro-level variables, and industrial policy text features with the corresponding industry adjustment coefficients. The adjustment coefficients are calibrated by regression of historical data using the maximum likelihood estimation method or the Bayesian estimation method.

7. The dynamic model design method for predicting long-term changes in industry market share according to claim 2, characterized in that: In S4, a life cycle quantitative index LCI-Score is constructed by weighting six types of indicators: industry output growth rate, industry activity index, capacity utilization rate and its change rate, market concentration index, cyclical fluctuation index, and policy intensity score. After standardization, it is mapped to the range of [-1,1]. The weight coefficients are determined by logistic regression calibration of historical samples or by the analytic hierarchy process.

8. The dynamic model design method for predicting long-term changes in industry share as described in claim 7, characterized in that: In S4, based on the LCI-Score value and its continuous trend, the industry is divided into three life cycle stages: emerging, mature, and declining. When making the determination, the corresponding threshold conditions are met and the preset number of cycles are maintained continuously. Correspondingly, mature industries are predicted using the ARIMA model, emerging industries are predicted using the Long Short-Term Memory (LSTM) network model, and declining industries are predicted using the state-space model combined with Kalman filtering.

9. The dynamic model design method for predicting long-term changes in industry market share according to claim 1, characterized in that: In S5, the normalization process adopts proportional normalization or Softmax normalization with temperature parameters to ensure that the sum of the final predicted shares of each industry in the same period is always 1. At the same time, a minimum share lower limit is set. When the normalized share of an industry is lower than the lower limit, it is corrected to the lower limit value, and the difference is deducted from the shares of the remaining industries proportionally before being normalized again.

10. The dynamic model design method for predicting long-term changes in industry market share according to claim 1, characterized in that: In S6, the triggering conditions for rolling updates include at least one of fixed-time-period timed triggering and event-driven triggering when the industry lifecycle label changes; after triggering, all feature data are updated, the adjustment coefficient is recalibrated, and the share calculation, prediction, and normalization process is repeated.