Economic indicator prediction method and system based on markov transition chain, electronic device and medium

By using a three-zone dynamic factor model based on Markov transition chains and employing multi-source data screening and machine learning algorithms to construct an economic climate prediction model, a fine-grained classification and real-time early warning of economic climate status are achieved. This solves the problems of insufficient accuracy in state differentiation and arbitrary indicator selection in existing technologies, and improves the robustness and reliability of prediction.

CN122434348APending Publication Date: 2026-07-21STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
Filing Date
2026-04-27
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing economic forecasting techniques based on the two-zone system assumption suffer from insufficient precision in distinguishing states, making it difficult to accurately identify severe downturns. Furthermore, the indicator selection process lacks a systematic approach, affecting the model's generalization ability and forecast reliability.

Method used

A three-zone dynamic factor model based on Markov transfer chains is adopted. Predictive variables are screened through multi-source economic time series data, and dynamic factor models for upward, general downward and severe downward states are constructed. Machine learning algorithms are used to solve parameters and make real-time predictions, and the predicted economic indicators are output.

Benefits of technology

It enables precise classification and real-time early warning of economic conditions, improves the timeliness and accuracy of early warnings for severe downturns, solves the problems of insufficient precision in state differentiation and arbitrary indicator selection in traditional methods, and enhances the robustness and predictive reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434348A_ABST
    Figure CN122434348A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of big data analysis, and particularly relates to an economic indicator prediction method and system based on Markov transition chain, an electronic device and a medium; the method comprises: acquiring multi-source economic time series data; filtering a predetermined number of prediction variables from the multi-source economic time series data; based on the prediction variables, constructing a dynamic factor model configured with a three-zone Markov transition chain; performing parameter solving on historical time series data to generate a trained dynamic factor model including zone transition probabilities; inputting real-time economic data into the trained dynamic factor model, performing state recognition and indicator calculation, and outputting economic indicator prediction results. In this way, the technical problem of insufficient state classification accuracy of existing economic prosperity prediction technology based on a two-zone assumption is solved, and the real-time performance, accuracy and decision support capability of macroeconomic prediction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data analytics, and in particular to a method, system, electronic device, and medium for predicting economic indicators based on Markov transition chains. Background Technology

[0002] In the field of economic forecasting, existing technologies have established a certain foundation. The Markov-switching dynamic factor model (abbreviated as Markov regime-switching dynamic factor model) is one such model. It is widely used in the identification and prediction of economic conditions, and the current mainstream method adopts the two-zone system model ( This paper simplifies economic activity by dividing it into two states: upward and downward, to reflect economic cycle fluctuations. Existing technology (application publication number CN113052430A) discloses a comprehensive power economic activity index analysis system and its analysis method. This system utilizes multiple prediction models (such as least squares method and support vector machine) and a signal light mechanism, employing power data as an indicator for principal component dimensionality reduction and prediction, aiming to monitor macroeconomic activity. This method, with its multi-model integration and dynamic optimization features, improves the robustness of predictions to some extent. However, this technology still relies on the two-zone assumption and fails to fully consider the multi-state characteristics of economic fluctuations, particularly its inability to distinguish between general downturns and severe downturns, resulting in insufficient adaptability when predicting major economic recessions.

[0003] Specifically, existing technologies suffer from two prominent problems: First, the two-zone model oversimplifies the characterization of economic conditions, lacking the fine-grained identification of severe downturns and failing to accurately capture economic turning points. Second, the indicator selection process often relies on empirical methods, lacking a systematic screening mechanism, and cannot efficiently extract predictive indicators from a large number of latent variables, thus affecting the model's generalization ability. These problems are intertwined, making it difficult for existing economic forecasting technologies to balance state differentiation accuracy and indicator reliability in the context of intensified economic volatility, and failing to meet the need for timely early warning of severe economic downturns. Therefore, existing economic forecasting technologies based on the two-zone assumption suffer from insufficient state differentiation accuracy. Summary of the Invention

[0004] To address the aforementioned shortcomings or drawbacks, this invention provides an economic indicator forecasting method, system, electronic device, and medium based on Markov transition chains, which can solve the technical problem of insufficient state differentiation accuracy in existing economic prosperity forecasting technologies based on the two-zone assumption.

[0005] This invention provides a method for predicting economic indicators based on Markov transition chains, comprising: Obtain multi-source economic time series data.

[0006] Based on a pre-defined machine learning screening algorithm, a predetermined number of predictive variables are selected from multi-source economic time series data.

[0007] Based on multiple predictor variables, a dynamic factor model with a three-zone Markov transfer chain is constructed. The three-zone Markov transfer chain is used to divide the economic prosperity state into three discrete zones: an upward state, a general downward state, and a severe downward state.

[0008] Based on the constructed dynamic factor model, parameters are solved on historical time series data to generate a trained dynamic factor model including regional transition probabilities. The regional transition probabilities are defined in the transition matrix of the three-region Markov transition chain and are used to characterize the probability of transition between different economic states.

[0009] Real-time economic data is input into a trained dynamic factor model, which is then used to identify the execution status and calculate indicators based on regional transition probabilities, and outputs economic indicator prediction results.

[0010] According to a second aspect, this invention provides an economic indicator prediction system based on Markov transfer chains, comprising: The time series data acquisition module is used to acquire multi-source economic time series data.

[0011] The economic forecast variable screening module is used to select a predetermined number of forecast variables from multi-source economic time series data based on a preset machine learning screening algorithm.

[0012] The dynamic factor model building module is used to construct a dynamic factor model with a three-zone Markov transition chain based on multiple predictor variables. The three-zone Markov transition chain is used to divide the economic prosperity state into three discrete zones: an upward state, a general downward state, and a severe downward state.

[0013] The dynamic factor model training module is used to solve parameters on historical time series data based on the constructed dynamic factor model, and generate a trained dynamic factor model including regional transition probabilities. The regional transition probabilities are defined in the transition matrix of the three-region Markov transition chain and are used to characterize the transition probability between different economic prosperity states.

[0014] The real-time economic indicator forecasting module is used to input real-time economic data into a trained dynamic factor model, perform state identification and indicator extrapolation based on regional transition probabilities, and output economic indicator forecasting results.

[0015] According to a third aspect, the present invention provides an electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to execute any of the economic indicator prediction methods based on Markov transition chains in the embodiments of the present invention.

[0016] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute any of the economic indicator prediction methods based on Markov transition chains in the embodiments of the present invention.

[0017] The present invention provides an economic indicator forecasting method based on Markov transition chains. This method is achieved through four core steps: multi-source data acquisition and screening, construction of a three-zone dynamic factor model, model parameter solving, and real-time forecasting. Specifically, multi-source economic time-series data is acquired to integrate multi-dimensional information reflecting economic activities; a predetermined number of predictive variables are selected from the data using a pre-set machine learning screening algorithm to address the arbitrariness of indicator selection in traditional methods and improve the quality and predictive ability of input information; based on the selected predictive variables, a dynamic factor model configured with a three-zone Markov transition chain is constructed to finely divide the economic climate into three discrete zones: upward, moderate downward, and severe downward; based on the constructed model, parameters are solved on historical data to generate a trained model containing zone transition probabilities, which quantify the likelihood of transitions between different economic states; real-time economic data is input into the trained model, and state identification and indicator calculation are performed based on the zone transition probabilities, ultimately outputting the economic indicator forecast results.

[0018] In this technical solution, the present invention addresses the problem described in the background section of the existing two-zone model, which oversimplifies the characterization of economic conditions and struggles to effectively identify severe downturns. By introducing a three-zone Markov transition chain, the present invention provides a more refined classification of economic conditions, enabling the differentiation and modeling of three states: upward, moderate downward, and severe downward. This overcomes the shortcomings of the traditional two-zone model, which lacks sufficient precision in state differentiation and cannot accurately predict major economic recessions. Furthermore, addressing the issue of existing technologies relying on experience and lacking a systematic screening mechanism for indicator selection, the present invention employs a pre-defined machine learning screening algorithm to automatically select predictive variables from multi-source time-series data, constructing a data-driven indicator optimization process. This overcomes the drawbacks of traditional methods, such as weak model generalization ability and low predictive reliability due to the arbitrariness of indicator selection. Therefore, the technical solution of this invention solves the technical problem of insufficient precision in state differentiation in existing economic forecasting technologies based on the two-zone assumption, improving the timeliness and accuracy of early warnings of economic turning points, especially severe downturns. Attached Figure Description

[0019] Figure 1 This is a flowchart of an economic indicator prediction method based on Markov transition chains according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an economic indicator prediction system based on Markov transition chains according to an embodiment of the present invention; Figure 3 This is a block diagram of an electronic device used to implement embodiments of the present invention. Detailed Implementation

[0020] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] During the development of this invention, researchers, through extensive experiments and data analysis, revealed the intrinsic relationship between the volatility of economic indicators and the probability of state transitions: the state transitions of the economic system are not simple binary leaps, but rather exhibit a gradual transition pattern of "upward movement - general downward movement - severe downward movement." Based on this relationship, this invention innovatively proposes this technical solution, utilizing the dynamic characteristics of a three-zone Markov transition chain to couple micro-indicator volatility with macro-state transition probabilities through a state-space model. This achieves a precise characterization of economic conditions and prediction of turning points, embodying the core concept of "multi-state division - probabilistic transition - dynamic prediction."

[0022] Specifically, through comparative experiments, the invention team discovered that the traditional two-zone dynamic factor model method suffers from a technical bottleneck in its coarse state characterization: it simply divides economic prosperity into two states, "upward" and "downward," failing to distinguish the characteristics of regular fluctuations from major recessions. These technical deficiencies result in insufficient sensitivity to early warning of major economic events, failing to meet the needs of precise macroeconomic regulation.

[0023] Therefore, this invention provides an economic indicator forecasting method based on Markov transition chains, which can be applied to a macroeconomic monitoring and early warning system (hereinafter referred to as the "system"). This system can run on a server cluster via cloud services or local deployment to automate the processing and intelligent forecasting of multi-source economic data.

[0024] Specifically, this system can be deployed in various hardware environments, including but not limited to: cloud computing platforms, physical servers in specific data centers, and private deployment environments of other commercial organizations. This flexible deployment architecture allows the system to meet both the elastic computing needs of large-scale data processing in the cloud and the local deployment needs of data-sensitive scenarios. In terms of operation, the system achieves full automation from data acquisition and model training to prediction output through modular design.

[0025] Specifically, the system can be configured with a data acquisition and preprocessing unit, an intelligent variable selection unit, a model training unit, and a prediction execution unit. The data acquisition and preprocessing unit is responsible for periodically collecting economic time-series data from multiple heterogeneous data sources, such as specific statistical bureau databases and trading market data interfaces, and performing missing value imputation and standardization. The intelligent variable selection unit automatically completes the selection and dimensionality reduction of predictive variables through an elastic network algorithm combined with a rolling time window verification mechanism. The model training unit adopts a distributed computing architecture, supporting parallel parameter solving and periodic model updates for the three-zone dynamic factor model. The prediction execution unit provides online services for real-time economic status identification and indicator prediction based on the trained model. All units collaborate through standardized data interfaces to form a complete economic monitoring and analysis closed loop. In practice, the system can adopt a microservice architecture design, with core algorithm modules supporting independent upgrades and expansions, ensuring system stability and providing technical support for subsequent algorithm optimization. Through a visual interactive interface, users can view the probability distribution of economic conditions, trend analysis reports, and early warning signals in real time, providing intuitive data support for macroeconomic decision-making.

[0026] like Figure 1 As shown, the method may include: Step S110: Obtain multi-source economic time series data.

[0027] Multi-source economic time series data refers to a set of economic indicator observations collected from different statistical channels and arranged in chronological order. These include economic statistical indicators (such as industrial production index, fixed asset investment, money supply, and consumer price index) and survey indicators (such as consumer confidence index and purchasing managers' index). These data are usually recorded monthly or quarterly to comprehensively reflect the overall trend of economic activity.

[0028] Specifically, the system can automatically download raw data from specific financial data platforms and commercial research institutions' servers at regular intervals through an Application Programming Interface (API). The data then undergoes data cleaning (including imputation of missing values ​​and removal of outliers) and standardization (such as Z-score standardization, a data standardization method based on sample mean and standard deviation, used to transform raw data into standardized data with a mean of 0 and a standard deviation of 1 to eliminate differences in units and orders of magnitude between different indicators), forming a well-organized time series dataset.

[0029] For example, the system retrieves 35 alternative variables from historical data from January 2000 to December 2023, including the growth rate of industrial added value (monthly year-on-year), total retail sales of consumer goods (monthly month-on-month), and broad money supply M2 (quarterly year-on-year). The data is stored in CSV (Comma-Separated Values) format, with a file size of approximately 50MB.

[0030] Step S120: Based on the preset machine learning screening algorithm, a predetermined number of predictive variables are selected from the multi-source economic time series data.

[0031] Among them, machine learning screening algorithm refers to variable selection method based on regularized regression, such as the Elastic Net screening algorithm; predictor variable refers to the subset of indicators selected by the algorithm that have significant predictive ability for economic fluctuations; predetermined number refers to the optimal number of variables determined by cross-validation, used to balance model complexity and prediction accuracy.

[0032] Specifically, the system can determine the optimal combination of hyperparameters (such as L1 regularization coefficients) for the elastic net algorithm using a rolling window verification method (with a window length of 20 quarters). L2 regularization coefficient The regression coefficients of each candidate variable were calculated using the Least Angle Regression (LARS) method, and the variable with a non-zero regression coefficient was selected as the predictor variable.

[0033] For example, the system selects 6 core predictive variables from 35 candidate variables, including the Purchasing Managers' Index (PMI), the month-on-month change rate of total stock market capitalization, and the year-on-year monthly export value. The selection process is completed in a Python (a programming language) environment and takes about 3 minutes.

[0034] In another embodiment, the system can filter variables in the candidate variable set using the elastic net algorithm shown in formula (1): (1) Formula (1) is the objective function of the elastic net algorithm, which is used to select variables with significant predictive power for the dependent variable from the candidate variable set by minimizing the sum of the loss function and the regularization penalty term. It is a vector of dependent variables, with dimensions of . , The sample size, in this embodiment, refers to the time series observations of the economic indicator to be predicted (such as GDP growth rate); It is an independent variable matrix with dimensions of . , The total number of candidate variables is represented by the time series data of one candidate economic indicator in each column of the matrix. It is the vector of regression coefficients to be estimated, with dimension . The variables corresponding to non-zero coefficients are selected as predictors. and These are regularization adjustment parameters, all of which are non-negative real numbers. Controlling the strength of L1 regularization (Least Absolute Shrinkage and Selection Operator, or LASSO for short, is a regression analysis method used for variable selection and regularization in linear regression models. It achieves automatic feature selection by adding an L1 regularization penalty term to the loss function, causing some regression coefficients to shrink to zero) is used to generate sparse solutions for variable selection. The strength of L2 regularization (ridge regression) is controlled to handle multicollinearity among variables and improve model stability. (Symbols) The vector norm is represented by the subscripts 1 and 2, which represent the L1 norm and L2 norm, respectively.

[0035] Next, in this embodiment, formula (1) is equivalent to the following constrained optimization problem, denoted as formulas (2) and (3): (2) (3) in, and Is with , The corresponding constraint boundary constants. This equivalent form clarifies that the algorithm achieves the dual objectives of variable selection and stable estimation by imposing constraints on the L1 and L2 norms of the regression coefficient vector.

[0036] Furthermore, in this embodiment, in order to select the optimal regularization parameter and (i.e., determine) and The system employs a forward rolling window validation method. This method is designed specifically for the characteristics of time series data to avoid future information leakage. Specifically, an initial training set time window is set (e.g., from Q1 2006 to Q4 2017), and the first validation set is the immediately following time point (e.g., Q1 2018). In each iteration, a set of candidate data is... Substitute the variables into formula (1), select the variables and build the model, make predictions on the validation set and calculate the prediction error. Then, scroll the training window forward by one time unit (e.g., one quarter), add one sample, and repeat the above training and validation process until the entire sample period has been covered. Finally, calculate the prediction error for each group. The optimal hyperparameter combination is determined by the combination of parameters that minimizes the average prediction error (root mean square error, or RMSE) across all rolling windows.

[0037] Then, in this embodiment, after determining the optimal adjustment parameters... and Then, the system uses the least angle regression method to evaluate the parameters in the elastic network model. Estimate the value. Least angle regression is an efficient computational algorithm that can accurately solve the optimization problem defined by formula (1) and obtain the regression coefficient vector. The estimated value .according to The system assigns values ​​to each component, with the absolute value of the coefficient greater than a very small positive threshold (e.g., ...). The corresponding variables of the candidate variables are determined as valid predictors, thus completing the screening from the candidate variable set to the predetermined number of predictors.

[0038] Therefore, by using the optimization framework defined by formulas (1) to (3) above, combined with the forward rolling window verification method and the minimum angle regression estimation algorithm, the system can automatically and data-drivenly select a subset of core variables with strong predictive power and robustness from a large number of candidate economic indicators, laying a reliable data foundation for the subsequent construction of a high-precision dynamic factor model.

[0039] Step S130: Based on multiple predictor variables, construct a dynamic factor model configured with a three-zone Markov transition chain.

[0040] Among them, the three-zone Markov transition chain is used to divide the economic prosperity state into three discrete zones: upward state, general downward state, and severe downward state. The upward state indicates that the economic growth rate is higher than the long-term trend level, the general downward state indicates that the economic growth rate is moderately declining, and the severe downward state indicates that the economic growth rate is sharply declining (such as a drop of more than 2 percentage points). The dynamic factor model refers to a statistical model based on state space representation that can extract common fluctuation components from multiple variables.

[0041] Specifically, the system can construct a state-space model, including measurement equations (establishing a linear mapping between observed variables and state vectors) and transition equations (describing the autoregressive evolution process of state vectors), and embed a three-zone Markov transition chain into the transition equations, so that the model intercept term changes dynamically with the switching of economic states.

[0042] For example, the system uses the MSBVAR package in R (a programming language) (a package in R specifically designed for estimating and analyzing Markov-based transition vector autoregressive models). The model was constructed using a statistical software package. The measurement equation of the state-space model has a dimension of 6 (corresponding to 6 predictors). The transition equation sets the state vector to include one common factor and its second-order lag term. The model parameters were estimated using the Bayesian method.

[0043] In another embodiment, the system can describe the linear mapping relationship between hard indicators and potential common factors and trait components using the following formula (4), thereby constructing the measurement equation part of the model: (4) Formula (4) is the standard factor model structure, used to characterize the first factor. One hard indicator Monthly growth rate at any time How to be a scalar dynamic common factor and its own special components A joint explanation. The total number of hard indicators; It is the first Each hard indicator corresponds to a common factor. The factor loading coefficients reflect the sensitivity of this indicator to common economic fluctuations; It is the first The lag order of a hard indicator represents the time offset of the indicator's growth rate data relative to the current economic cycle state, and is used to handle the leading or lagging relationship between different indicators and the economic cycle.

[0044] Next, in this embodiment, the system can also describe the relationship between the survey indicators and the common factors and trait components with multiple lags using formula (5): (5) Formula (5) is the structure of a factor model with lag terms, used to characterize the first... Monthly growth rate of each survey indicator . The total number of survey indicators; These are the corresponding factor loading coefficients; It is the lag order of the survey indicator; It is a pre-set lag period, indicating that the survey indicator is affected by the past continuous period. The setting of the cumulative impact of common factors over a period reflects the smooth response of survey indicators (usually based on expectations) to economic trends.

[0045] Furthermore, in this embodiment, the system assumes that the idiomatic error term components of all indicators in formulas (4) and (5) It follows a vector autoregression process of order q, as shown in equation (6): (6) in, yes There are several autoregressive coefficient matrices, and their covariance matrices with the disturbance term are... All are constrained to be diagonal matrices, which means that the eigenvalues ​​of each indicator only follow their own independent autoregressive processes that do not intersect with other indicators, thus simplifying the model and highlighting the role of common factors. It has a mean of zero and a covariance matrix of The independent and identically distributed Gaussian perturbation terms.

[0046] Furthermore, in this embodiment, the scalar common factor in formulas (4) and (5) It is set to conform to an intercept term with a region dependency. The autoregressive process is shown in equation (7): (7) in, Common factors The autoregressive coefficients describe its own dynamic persistence; It is a random perturbation that follows a standard normal distribution. Core parameters It is a time-varying intercept term whose value depends on an unobservable discrete state variable. This state variable corresponds to three economic prosperity zones: "upward trend," "moderate downward trend," and "severe downward trend." Specifically, ,in If and only if Otherwise, it is 0. These are the intercept level parameters corresponding to the three zones.

[0047] State variables Assuming to follow a first-order traversal three-zone Markov chain process, its state transition law is determined by a... The transition probability matrix Definition, matrix element Indicates from the district system Transfer to district system The probability of.

[0048] Wherein, the transition probability matrix It can be as follows: ; To make the model estimate more stable and economically meaningful, the system imposes the following constraints on the transition matrix, namely Equations (8) and (9): (8) (9) Formula (8) prohibits direct transitions between "upward states" (block 1) and "severe downward states" (block 3), which reflects the reasonable assumption that extreme states usually need to go through intermediate states to transition in the evolution of the economic cycle. Formula (9) is the standard normative constraint of the Markov chain, ensuring that the sum of the probabilities of transitioning from any state to all possible states (including itself) is 1.

[0049] Finally, in this embodiment, the system defines the complete set of equations (4) to (9). The model is represented in a compact state-space form, namely, equations (10) and (11): (10) (11) in, It is made by all The observed values ​​of the hard indicators and survey indicators over time Superimposed Dimensional column vector. It is the expanded state vector, which includes common factors. and its lag terms, as well as all trait components And its lag terms. B is the observation matrix (or factor loading matrix), which establishes the state vector. With observation vector The linear mapping between them (corresponding to formulas (4) and (5)). It is the state transition matrix, which describes the state vector. The dynamic evolution law (the AR process that integrates formulas (6) and (7), the AR process is a time series analysis method, also known as the autoregressive process). It depends on the region system The intercept term vector has a structure such that only the part corresponding to the common factor equation (i.e., the intercept of formula (7)) varies with the regime. R is the selection matrix of the perturbation term. It is the joint Gaussian perturbation term vector, and its covariance matrix is ​​Q.

[0050] Therefore, by combining and defining the hierarchy of the above formulas (4) to (11), the system can construct a complete three-zone Markov regional transition dynamic factor model. The model first extracts a leading common factor from multi-source mixed-frequency economic indicators through measurement equations (formulas (4), (5) and (10)); then, through the transition equations (formulas (6), (7) and (11)) and embedding a three-zone Markov chain with economic constraints (formulas (8) and (9)), it jointly characterizes the dynamic evolution path of the factor and the underlying economic state transition mechanism; finally, the system can use algorithms such as Kalman filtering and expectation maximization to perform joint parameter estimation and state inference based on the state space model, thereby realizing real-time identification of economic prosperity zones, quantification of state transition probabilities, and accurate prediction of key economic indicators (such as GDP growth rate).

[0051] Step S140: Based on the constructed dynamic factor model, perform parameter solving on the historical time series data to generate a trained dynamic factor model including regional transition probabilities.

[0052] The zone transition probability is defined in the transition matrix of a three-zone Markov transition chain and is used to characterize the probability of transitions between different economic states. For example, the probability of transitioning from an upward state to a general downward state is denoted as... The probability of recovering from a severe downtrend to an uptrend is denoted as . Parameter solving refers to estimating unknown parameters (such as factor loadings and autoregressive coefficients) in a model using numerical optimization algorithms.

[0053] Specifically, the system can combine the Expectation-Maximization Algorithm (EM algorithm, an iterative optimization algorithm for parameter estimation in statistical models with latent variables) with Kalman filtering. An optimal recursive estimation algorithm for linear Gaussian state-space models is proposed, which iteratively solves the problem as follows: In the expectation step, the conditional expectation of the state vector is calculated based on the current parameters; in the maximization step, the parameter estimates are updated; and the iteration continues until the change in the log-likelihood function is less than 1 / 3. The event will end at that time.

[0054] For example, the system uses data from the first quarter of 1995 to the fourth quarter of 2020 to train the model, and the final transition matrix contains... The probability of maintaining the upward state is 0.85. The probability of maintaining a downward trend is generally 0.70. The probability of maintaining a severe downlink state is 0.60, and the training process takes about 15 minutes on a server (configured with an 8-core CPU and 32GB of memory).

[0055] Step S150: Input real-time economic data into the trained dynamic factor model, perform state identification and indicator calculation based on regional transition probability, and output the economic indicator prediction results.

[0056] Among them, state identification refers to calculating the marginal probability of the economy being in each state based on the regional transition probability; indicator extrapolation refers to extrapolating future economic indicator values ​​(such as GDP growth rate) using the common factor sequence output by the model.

[0057] Specifically, the system can use a recursive filtering algorithm to input the latest data into the trained model, calculate the state probability distribution and indicator prediction values ​​for the next 1 to 4 quarters, and generate a visual report.

[0058] For example, the system takes real-time data from the first quarter of 2024 as input (e.g., PMI of 52.5 and M2 growth rate of 8.2%) and outputs results showing that the probability of the economy being in an upward trend is 65%, the probability of a general downward trend is 30%, and the probability of a severe downward trend is 5%. It also predicts that the GDP growth rate for the second quarter of 2024 will be 5.2%, and the prediction results are output in JSON (JavaScript Object Notation, a data exchange format).

[0059] Therefore, according to the above implementation method, the system acquires multi-dimensional information and supplementary information to collect electricity consumption data and related auxiliary information characterizing economic activities; performs standardized preprocessing on this information to achieve normalization and comparability of multi-source heterogeneous data; inputs the preprocessed data into a preset mixing dynamic factor model to perform feature extraction and fusion processing; generates common factor sequences and specific component sequences through the model to characterize the common driving force of macroeconomics and sectoral specific fluctuations; uses Kalman filtering and expectation-maximization algorithms to perform parameter estimation and update model parameters to achieve adaptive optimization of the model; and performs state prediction and filtering updates on real-time updated data to generate real-time prediction results of gross domestic product.

[0060] Specifically, in this implementation, the technical solution addresses the problems of traditional economic forecasting methods, such as strong lag and difficulty in capturing high-frequency real-time signals, as described in the background. By comprehensively utilizing high-frequency power big data and multi-dimensional auxiliary information, a forecasting foundation capable of reflecting real-time changes in economic activities is established, overcoming the shortcomings of traditional forecasting methods relying on quarterly or monthly statistical data in terms of poor timeliness. Addressing the data fusion challenges caused by heterogeneous multi-source information and frequency mixing, a combination of standardized preprocessing and a mixed-frequency dynamic factor model achieves effective alignment and collaborative analysis of information at different frequencies and scales. Addressing the challenges of the time-varying nature of economic systems and the need for dynamic adjustment of model parameters, a closed-loop process integrating parameter estimation and model updates enables continuous optimization and state tracking of model parameters. To ensure that forecast results meet the needs of real-time business decision-making, a state prediction and filtering update mechanism ensures that forecast results can be quickly corrected and output based on the latest data stream. Therefore, this implementation's technical solution solves the technical problem of insufficient state differentiation accuracy in existing economic prosperity forecasting techniques based on the two-zone assumption, improving the real-time performance, accuracy, and decision support capabilities of macroeconomic forecasting.

[0061] In some embodiments, multi-source economic time series data includes statistical indicators and survey indicators; based on a preset machine learning screening algorithm, a predetermined number of predictive variables are selected from the multi-source economic time series data, including: The target hyperparameter combination for machine learning screening algorithms is determined by using the rolling window validation method.

[0062] Among them, the rolling window validation method is a cross-validation strategy unique to time series data. It uses a fixed-length time window to roll over the time series, using the data within the window as the training set and the data immediately following it as the validation set, in order to evaluate the performance of hyperparameters and avoid future information leakage. The target hyperparameter combination refers to the combination of regularization parameter values ​​that minimizes the prediction error of the model on the validation set.

[0063] Specifically, the system divides the total duration into multiple consecutive fixed-length windows. Each time, data from the current window is used to train the model, and performance is tested on a validation set consisting of the next time unit. Hyperparameter performance is comprehensively evaluated through multiple rolling calculations. For example, the system sets the rolling window length to 20 quarters (i.e., 5 years), starting from the first quarter of 2006. Data from the first quarter of 2006 to the fourth quarter of 2017 (a total of 48 quarters) is used as the first training window, and the first quarter of 2018 is used as the first validation set for prediction and error calculation. The window then rolls forward one quarter, repeating training and validation until the entire sample period is covered. Finally, the hyperparameter combination that minimizes the average root mean square error over all rolling validation cycles is selected as the target hyperparameter combination, such as the L1 regularization coefficient. The L2 regularization coefficient is 0.1. It is 0.05.

[0064] The application of target hyperparameter combinations is used to control the sparsity and regularization strength in the variable selection process, and the regression coefficients of each candidate variable are calculated using the least angle regression method. Each candidate variable is a potential predictor variable pre-selected from statistical and survey indicators based on economic prosperity analysis theory.

[0065] Here, sparsity refers to the fact that regularization penalties make the regression coefficients of most variables zero, thereby achieving automatic variable selection; regularization strength refers to the degree of constraint of the penalty term on the magnitude of the regression coefficients; least angle regression is an efficient algorithm for solving linear models, which can generate piecewise linear solution paths.

[0066] Specifically, the system substitutes the target hyperparameter combination into the objective function of the elastic net algorithm and solves it using the least angle regression method. This algorithm continuously adjusts the coefficients of the predictor variables with the highest correlation to the residuals to traverse all possible sparse solutions in the fewest steps.

[0067] For example, the system uses Python. In the library The class (which is a concrete, operational, pre-built machine learning model class for implementing the Elastic Net filtering algorithm described above) sets... (correspond ), (correspond The algorithm fits the monthly growth rate data of 35 candidate variables (with a dimension of T×35, where T is the sample period length). The algorithm converges within 15 iterations and outputs 35 regression coefficient estimates.

[0068] Based on the non-zero coefficients in the regression coefficients, a predetermined number of predictive variables are selected from the candidate variables.

[0069] Here, a non-zero coefficient refers to a coefficient whose absolute value, after estimation using the least angle regression method, is greater than the numerical precision threshold (e.g., ...). The regression coefficients; screening refers to selecting the corresponding variables to enter the final model based on whether the coefficients are non-zero.

[0070] Specifically, the system sets a very small positive number as the threshold (e.g. The system compares the absolute value of the regression coefficients with a threshold. Variables with absolute values ​​greater than the threshold are considered valid variables that contribute to the prediction and are retained. For example, after fitting the system to 35 candidate variables, 6 variables had regression coefficients with absolute values ​​greater than the threshold. These six variables (including the Manufacturing Purchasing Managers' Index (PMI) and the growth rate of industrial added value) are selected as a predetermined number of predictive variables, which constitute the input dataset for the subsequent dynamic factor model.

[0071] Therefore, according to the above implementation method, the system can automatically and data-drivenly complete the process of screening key predictive variables from a large number of candidate indicators, effectively avoiding model bias caused by subjective variable selection in traditional methods, and improving the robustness and interpretability of the economic prosperity prediction model.

[0072] In some embodiments, a dynamic factor model configured with a three-zone Markov transition chain is constructed based on multiple predictor variables, including: Construct a state-space model, which includes measurement equations and transition equations.

[0073] Among them, the state-space model refers to a mathematical model that uses state variables to describe the evolution of a dynamic system. It consists of measurement equations and transition equations. The measurement equations are used to establish the relationship between observable variables and unobservable state variables, while the transition equations are used to describe the dynamic process of state variables changing over time.

[0074] Specifically, the system can construct a state-space model by defining the dimension of the state vector (such as including common factors and their lags) and the form of the equation (such as a linear Gaussian model), where the measurement equation linearly relates the observed variables to the state vector, and the transition equation sets the autoregressive evolution of the state vector.

[0075] For example, the system constructs a state-space model with a state vector dimension of 3 (containing one common factor and its second-order lag term), and the measurement equation is expressed as follows: (in It is a 6-dimensional observation vector. It is a 3-dimensional state vector, Z is Factor loading matrix, (This is the observation error), the transfer equation is expressed as: (where T is) The transition matrix, where R is the coefficient matrix of the perturbation term. It is a state error.

[0076] Multiple predictor variables are incorporated as observed variables into the measurement equation, which is used to establish the mapping relationship between the observed variables and the state vector.

[0077] Among them, the observed variables refer to directly measurable economic indicators (such as GDP growth rate and inflation rate); the state vector refers to the non-observable variables that represent the potential common fluctuation components of the economic system; and the mapping relationship refers to the mathematical relationship that transforms the state vector into the observed variables through linear transformation.

[0078] Specifically, the system can define the measurement equation by estimating the factor loading matrix, such that each observed variable is a linear combination of state vectors, and assumes that the observation errors follow independent and identically distributed rules. For example, for six predictor variables (such as the industrial production index and the consumer confidence index), the measurement equation sets the factor loading matrix Z as follows: The matrix, whose parameters are obtained through maximum likelihood estimation, explains more than 80% of the variance of the observed variables.

[0079] The transition equations are configured as autoregressive processes with time-varying intercept terms to describe the dynamic evolution of the state vector.

[0080] Among them, the time-varying intercept term refers to the constant term in the transfer equation that changes with time, and its value depends on the economic state; the autoregressive process refers to a linear model in which the current value of the state vector depends on its past values.

[0081] Specifically, the system can extend the transfer equation to (in It is the time-varying intercept term. Indicates economic status. and These are autoregressive coefficients), and the intercept is adjusted based on state dependence. For example, setting the autoregressive order to 2, the time-varying intercept term takes different values ​​depending on the economic state: in an upward trend... In general downtrend In a severe downtrend .

[0082] A three-zone Markov transition chain is embedded in the state-space model, so that the value of the time-varying intercept term is determined by the current state of the Markov transition chain, which can be an up-going state, a normal down-going state, or a severe down-going state.

[0083] Here, a Markov transition chain refers to a stochastic process in which the state transition probability depends only on the current state; the current state refers to the discrete state of the economic situation at a specific point in time (encoded by 1, 2, and 3 to represent upward, moderate downward, and severe downward states, respectively).

[0084] Specifically, the system can define state variables. and the transition probability matrix P ( (matrix), making the time-varying intercept term Follow The system switches states and estimates the state sequence using filtering algorithms. For example, the system uses the Hamiltonian filtering algorithm (a core inference engine, alongside the Kalman filtering algorithm, specifically designed for inferring states in a discrete Markov system) to estimate state probabilities based on historical data. The transition matrix P is initially set to a uniform distribution and updated iteratively.

[0085] Specific constraints are imposed on the transition probability matrix of the three-zone Markov transition chain. These constraints include prohibiting the transition between the uplink state and the severe downlink state.

[0086] The transition probability matrix refers to a square matrix (elements) that describes the transition probabilities between states. Indicates from state Transition to state The probability); specific constraints refer to mathematical restrictions introduced to enhance the economic rationality of the model.

[0087] Specifically, the system can set constraints. and (Direct conversion between upward and severe downward states is prohibited), while ensuring that the sum of probabilities for each row is 1 (i.e.) ).

[0088] For example, the transition matrix constraint is: ; in, The parameters are then solved using maximum likelihood estimation.

[0089] Therefore, based on the above implementation method, the system can construct a three-zone dynamic factor model with a rigorous structure and clear economic significance. Through the state-space framework and Markov chain constraints, it can accurately characterize the state transition law of economic prosperity and provide a reliable basis for subsequent forecasting.

[0090] In some embodiments, based on the constructed dynamic factor model, parameter solving is performed on historical time series data to generate a trained dynamic factor model including regional transition probabilities, including: The state-space model is solved using the Kalman filter algorithm, which includes a state prediction step and a state update step.

[0091] Among them, the Kalman filter algorithm is an optimal recursive estimation algorithm for linear Gaussian state-space models. It optimizes the estimation of the state vector by alternating between two steps: prediction and update. The state prediction step refers to calculating the prior estimate of the state at the current time based on the dynamic equation of the model. The state update step refers to correcting the prior estimate using new observation data to obtain the posterior estimate.

[0092] Specifically, the system can initialize the state vector and covariance matrix, and then recursively execute prediction and update loops: in the prediction step, the state is deduced based on the transition equation; in the update step, the estimate is adjusted by combining the measurement equation and the observations. For example, the system uses Python (a programming language). In the library The class (a Python programming interface and algorithm implementation class for building, configuring and solving linear Gaussian state-space models) sets the state vector dimension to 3, the observation vector dimension to 6, and the initial state covariance matrix to be a diagonal matrix (diagonal elements are 1.0). The algorithm calculates the state estimate in each iteration.

[0093] In the state prediction step, the prior prediction value of the state vector at the current time point is derived based on the transition equation and the posterior state estimate of the previous time point.

[0094] Among them, the posterior state estimate refers to the optimal estimate of the state vector after correction based on the observed data at the previous time step; the prior prediction refers to the uncorrected prediction of the state at the current time step based on the model's dynamic equations; and the derivation refers to the calculation of the temporal evolution of the state vector through mathematical equations.

[0095] Specifically, the system can be described by the recurrence relation. (in It is the current state vector, and T is the transition matrix. It is the posterior estimate of the previous time step. (where R is the process noise and R is the disturbance term coefficient matrix) Calculate the prior prediction value and update the prediction error covariance.

[0096] For example, for a point in time The system is based on posterior estimate and transition matrix T ( Matrix, elements such as ), calculated Prior prediction value The prediction error covariance matrix is ​​updated simultaneously.

[0097] In the state update step, the prior prediction value is corrected based on the measurement equation and the observed variable value at the current time point to obtain the posterior estimate of the state vector at the current time point.

[0098] Among them, the observed variable value refers to the economic indicator data actually measured at the current moment; the correction refers to minimizing the estimation error by fusing prior predictions and observed values ​​through Kalman gain weighted fusion; the posterior estimate refers to the optimal estimate of the state vector after correction.

[0099] Specifically, the system can measure the equations (in Z is the observation vector, and Z is the factor loading matrix. (Observation noise) Calculate the observation prediction value, and then calculate the Kalman gain. and update the state estimate. Covariance matrix.

[0100] For example, for Observed values Factor loading matrix Z ( (The matrix is ​​estimated using historical data), and the system calculates the Kalman gain. ( (matrix), after correction, the posterior estimate is obtained. The estimation error covariance decreases.

[0101] The state prediction and state update steps are executed iteratively. The model parameters of the dynamic factor model and the zone transition probabilities of the three-zone Markov transition chain are solved by combining the expectation-maximization algorithm to obtain the trained dynamic factor model.

[0102] Among them, the expectation-maximization algorithm is a statistical method for solving model parameters by iterating through expectation steps and maximization steps; model parameters include factor loadings, autoregressive coefficients, etc.; the zone transition probability refers to the elements in the Markov transition matrix, representing the probability of transition between states.

[0103] Specifically, the system can be executed iteratively: in the expectation step, the conditional expectation and likelihood function of the state vector are calculated using Kalman filtering and smoothing algorithms; in the maximization step, the model parameters and transition probabilities are updated to maximize the likelihood function; iteration continues until the parameter changes are less than a threshold (e.g., ...). Stop when ).

[0104] For example, the system is set to a maximum of 100 iterations, with initial parameters randomly initialized. After 20 iterations, it converges, ultimately yielding the factor loading matrix Z (e.g., elements). ), autoregression coefficient And the transition probability matrix P (such as The training process takes about 10 minutes on the server.

[0105] Therefore, according to the above implementation method, the system can accurately estimate the parameters and state transition probabilities of the dynamic factor model through efficient recursive algorithms and iterative optimization, generate a reliable trained model, and provide a solid foundation for economic forecasting.

[0106] In some embodiments, the step of performing state identification and indicator calculation based on regional transition probability and outputting economic indicator forecast results includes: Based on the regional transfer probability of the three-region Markov transfer chain, the probability values ​​of the economic climate being in an upward, moderate downward, and severe downward state are calculated.

[0107] Among them, the zone transition probability refers to the conditional probability that the economic prosperity will shift to a certain state in the next moment, derived from the transition matrix of the three-zone Markov transition chain; the probability value refers to the marginal probability that the economy is in each state at the current or future moment, calculated by a filtering algorithm (such as Hamiltonian filtering), and the value ranges from 0 to 1.

[0108] Specifically, the system can calculate the probability distribution of each state at each time point based on the transition probability matrix and observation data using a recursive Bayesian update formula, as follows: ; in, express Current state This represents the observation data up to time t.

[0109] For example, for a point in time In the first quarter of the year, the system calculated that the probability of the economy being in an upward trend was 0.65 (65 percent), the probability of a moderate downward trend was 0.30 (30 percent), and the probability of a severe downward trend was 0.05 (5 percent), with the sum of the probability values ​​being 1.

[0110] The current and future economic conditions are determined by the probability values ​​to determine the regional category to which they belong.

[0111] Among them, the zone category refers to the specific classification of the economic prosperity status (upward, moderate downward or severe downward); determination means to compare the probability values ​​of each state and take the state with the highest probability as the zone category at the current or predicted time.

[0112] Specifically, the system can select the state index with the highest probability value as the zone category by taking the maximum value operation. ,in Values ​​1, 2, and 3 correspond to upward, moderate downward, and severe downward states, respectively. For example, based on the above probability values ​​(0.65 for upward, 0.30 for moderate downward, and 0.05 for severe downward), the system determines the current economic climate status as upward, because the upward state has the highest probability.

[0113] A common factor sequence is extracted from the posterior estimate of the state vector. The common factor sequence represents the synthetic leading index sequence of economic prosperity.

[0114] Among them, the posterior estimate of the state vector refers to the optimal estimate of the state vector obtained after correction by the Kalman filter algorithm; the common factor sequence refers to the time series data extracted from the state vector that represents the common fluctuation components of the economy; and the synthetic leading index sequence refers to the standardized indicator composed of the common factor sequence that can predict economic trends in advance.

[0115] Specifically, the system can inversely deduce common factor components from the posterior estimates of the state vectors using the measurement equations of the state-space model, and then standardize these components (e.g., subtracting the mean and dividing by the standard deviation) to generate a synthetic leading index sequence. For example, the system extracts a common factor sequence from the posterior estimates of the state vectors, with the sequence value being... (Length corresponds to time point), and then scale it to a synthetic leading index series with an average of 100 and a standard deviation of 10 for the base year (e.g., 2015), such as an index value of 105.3 for the first quarter of 2024.

[0116] By combining regional classifications and common factor sequences, economic indicator forecasts are generated.

[0117] Here, "combination" refers to using regional category information (such as status labels) and common factor sequences (such as leading index values) as inputs to calculate the target economic indicator value through a prediction model; the economic indicator prediction result refers to the numerical prediction of future economic indicators (such as GDP growth rate).

[0118] Specifically, the system can use a linear regression model or state-space extrapolation to input the regional category as a dummy variable (e.g., uplink state code is 1, others are 0) along with the common factor sequence into the prediction equation, as shown in the formula: ,in It is a predicted value. These are model parameters.

[0119] For example, the system uses a pre-trained regression model (parameters) Input the common factor value of 0.55 for the first quarter of 2024 and the upward state dummy variable 1 to calculate the GDP growth rate forecast for the second quarter of 2024 as 5.2%.

[0120] Therefore, according to the above implementation method, the system can automatically complete the identification of economic status and the prediction of indicators, and output intuitive and quantitative economic indicator prediction results to support decision-making.

[0121] In some embodiments, before embedding a three-zone Markov transition chain in the state-space model, the method further includes: Define multiple state transition probability elements in the transition probability matrix, where each element represents the probability of transitioning from an upward state, a normal downward state, and a severe downward state to any other state.

[0122] Among them, the state transition probability element refers to the value at each position in the transition probability matrix, which is used to quantify the objective possibility of transitioning from one state to another; the transition probability matrix is ​​a 3-row, 3-column mathematical matrix, with its rows and columns corresponding to three economic states (upward, moderate downward, and severe downward).

[0123] Specifically, the system can create A matrix data structure, where each element ( and The value (1, 2, or 3) indicates the state. Transition to state The probability is calculated, and these probability values ​​are estimated using historical data. For example, the system initializes the transition probability matrix as a uniform distribution, i.e., each... Setting it to 1 / 3 ≈ 0.333, and then adjusting the probability value based on economic time series data using maximum likelihood estimation, the final result is as follows: (The probability of maintaining an upward trend from an upward trend) (The probability of transitioning from an upward state to a normal downward state) (The probability of transitioning from an upward trend to a severe downward trend).

[0124] Set the transition probability elements representing the direct transition from the uptrend state to the severe downtrend state, and the direct transition from the severe downtrend state to the uptrend state, to zero.

[0125] Direct transfer refers to a one-step transition of economic state between adjacent time points without going through intermediate states; setting it to zero means that these transfer paths are forcibly prohibited through mathematical constraints in order to conform to the continuous law of economic cycle evolution.

[0126] Specifically, the system can explicitly set the transition probability matrix. and This ensures that the economic situation does not jump directly from an upward trend to a severe downturn or vice versa, thereby avoiding unreasonable predictions of sudden changes.

[0127] For example, in the transition matrix, the system is fixed. and Other elements such as The final matrix, estimated using the EM (Expectation-Maximization) algorithm, is in the following form: First row The second line The third line .

[0128] Set the sum of all transition probability elements in each row of the transition probability matrix to one.

[0129] The fact that the sum of probabilities in each row is one refers to the basic canonical condition of the Markov chain, which ensures that the sum of all possible transition probabilities from any state is 100% (i.e., 1), representing the completeness of state transitions.

[0130] Specifically, the system can use the normalization step in the iterative algorithm to scale the elements of each row after each parameter update, so that for each state... ,satisfy For example, if the estimated probability value for the second row is... The sum is 1.05, which the system will normalize to. (To make the sum equal to 1), ensuring mathematical rigor.

[0131] Therefore, according to the above implementation method, the system can construct a transition probability matrix that conforms to economic laws and mathematical norms, providing a reliable foundation for the embedding of the three-zone Markov transition chain.

[0132] In some embodiments, the state prediction step and the state update step are executed iteratively, and the model parameters of the dynamic factor model and the zone transition probabilities of the three-zone Markov transition chain are solved by combining the expectation-maximization algorithm to obtain the trained dynamic factor model, including: Initialize the model parameters and initial values ​​of the regime transition probability of the dynamic factor model.

[0133] Among them, model parameters refer to the unknowns that need to be estimated in the dynamic factor model, including the factor loading matrix, autoregressive coefficients and noise covariance matrix; the initial value of the zone transition probability refers to the initial setting value of the elements of the transition probability matrix of the three-zone Markov transition chain, which is usually based on prior knowledge or uniform distribution.

[0134] Specifically, the system can use random initialization methods to assign small random numbers to the model parameters (e.g., sampling from a normal distribution with a mean of 0 and a standard deviation of 0.1), and set uniform initial values ​​for the transition probability matrix (e.g., each element is 0.1). For example, the system initialization factor loading matrix is: The identity matrix has autoregressive coefficients. The transition probability matrix is ​​initially set to have all elements equal to 0.333, and the initialized parameters are stored as follows: (Numerical Python, a Python library) array format.

[0135] The expected step is executed, which calculates the conditional expectation and covariance matrix of the state vector based on the current parameter values ​​using the Kalman filter algorithm.

[0136] The expectation step refers to the first step in the expectation-maximization algorithm, which aims to calculate the conditional expectation (average estimate given the observed data) and conditional covariance matrix (a measure of the uncertainty of the estimate) of the state vector. The conditional expectation is the optimal estimate of the state vector given the observed data; the covariance matrix is ​​the variance-covariance matrix of the state estimation error.

[0137] Specifically, the system can use a smoothing algorithm based on Kalman filtering (such as...) The smoother (an optimal fixed-interval smoothing algorithm for linear Gaussian state-space models) recursively calculates the conditional expectation and covariance matrix of the state vector at each time point based on the current parameters and all observed data, using the following formula: ,in This represents all time series data.

[0138] For example, for a point in time The conditional expectation value of the system's state vector is calculated as follows: The conditional covariance matrix is Diagonal matrix (diagonal elements are) The calculation process is performed in a Python environment. The library (an open-source software library focused on statistical modeling, econometric analysis, and hypothesis testing) is implemented.

[0139] Perform a maximization step to update the model parameters and estimates of regime transition probabilities based on conditional expectations and covariance matrices.

[0140] The maximization step refers to the second step in the expectation-maximization algorithm, which updates the model parameters and transition probabilities by maximizing the likelihood function of the complete data; updating the estimate refers to adjusting the parameter values ​​based on the output of the expectation step to improve the model fit.

[0141] Specifically, the system can use numerical optimization methods (such as gradient ascent) to calculate the gradient of the likelihood function using conditional expectation and covariance matrix, and update parameters. For example, the factor loading matrix is ​​adjusted using least squares estimation, and the transition probability is updated using counting estimation. The formula is as follows: , where Q is the expectation function.

[0142] For example, the updated factor loading matrix elements become: The transition probability matrix is ​​updated to The update process takes about 2 seconds.

[0143] Iteratively execute the expectation step and the maximization step until the estimated values ​​of the model parameters and the regime transition probability satisfy the convergence condition.

[0144] Here, iteration refers to the cyclical process of repeatedly executing the expected step and the maximization step; the convergence condition is the criterion for stopping iteration when the parameter change is less than a preset threshold, which is used to ensure the stability of the algorithm.

[0145] Specifically, the system can set a maximum number of iterations (e.g., 100 times) and a convergence threshold (e.g., ...). The algorithm calculates the norm of the parameter change after each iteration, and terminates if the change is less than a threshold or the maximum number of iterations is reached. For example, after 15 iterations, the L2 norm of the parameter change decreases from the initial 0.1 to... less than the threshold The algorithm converged, with a total computation time of approximately 5 minutes.

[0146] The model parameters and regional transition probabilities after iterative convergence are used as the objective solution parameters to complete the construction of the trained dynamic factor model.

[0147] Here, the target solution parameters refer to the optimized model parameters and transition probability values ​​output by the algorithm; construction refers to using these parameters to instantiate a usable dynamic factor model for prediction tasks.

[0148] Specifically, the system can save the converged parameters as a model file and integrate it into the prediction pipeline, for example, by using Python's pickle module to serialize the model object. For instance, the final model parameters include a factor loading matrix (such as elements...). ), autoregressive coefficients (e.g.) ), transition probability matrix (such as The model file size is 2MB (megabytes), and it can be directly loaded for real-time prediction.

[0149] Therefore, according to the above implementation method, the system can efficiently solve complex model parameters through automated iterative optimization, generate a high-precision trained dynamic factor model, and provide a reliable basis for economic forecasting.

[0150] Figure 2 This is a structural block diagram of an economic indicator prediction system based on Markov transition chains according to an embodiment of the present invention.

[0151] like Figure 2 As shown, this economic indicator forecasting system based on Markov transfer chains includes: The time series data acquisition module 210 is used to acquire multi-source economic time series data.

[0152] The economic forecast variable screening module 220 is used to select a predetermined number of forecast variables from multi-source economic time series data based on a preset machine learning screening algorithm.

[0153] The dynamic factor model construction module 230 is used to construct a dynamic factor model with a three-zone Markov transition chain based on multiple predictor variables. The three-zone Markov transition chain is used to divide the economic prosperity state into three discrete zones: an upward state, a general downward state, and a severe downward state.

[0154] The dynamic factor model training module 240 is used to perform parameter solving on historical time series data based on the constructed dynamic factor model, and generate a trained dynamic factor model including regional transition probabilities. The regional transition probabilities are defined in the transition matrix of the three-region Markov transition chain and are used to characterize the transition probability between different economic prosperity states.

[0155] The real-time economic indicator prediction module 250 is used to input real-time economic data into a trained dynamic factor model, perform state identification and indicator extrapolation based on regional transition probabilities, and output economic indicator prediction results.

[0156] The specific functions and examples of each module and submodule of the device in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0157] According to embodiments of the present invention, the above-described method of the present invention can be applied to an electronic device and a readable storage medium.

[0158] Figure 3 A schematic block diagram of an electronic device 600 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0159] like Figure 3 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0160] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0161] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a Markov transition chain-based economic indicator forecasting method. For example, in some embodiments, a Markov transition chain-based economic indicator forecasting method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the Markov transition chain-based economic indicator forecasting method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured, by any other suitable means (e.g., by means of firmware), to perform an economic indicator forecasting method based on a Markov transition chain.

[0162] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0163] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0164] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0165] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0166] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0167] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0168] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0169] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for predicting economic indicators based on Markov transition chains, characterized in that, include: Acquire multi-source economic time series data; Based on a preset machine learning screening algorithm, a predetermined number of predictive variables are selected from the multi-source economic time series data; Based on the aforementioned multiple predictor variables, a dynamic factor model configured with a three-zone Markov transfer chain is constructed. The three-zone Markov transfer chain is used to divide the economic prosperity state into three discrete zones: an upward state, a general downward state, and a severe downward state. Based on the constructed dynamic factor model, parameters are solved on historical time series data to generate a trained dynamic factor model including regional transition probabilities. The regional transition probabilities are defined in the transition matrix of the three-region Markov transition chain and are used to characterize the transition probability between different economic prosperity states. Real-time economic data is input into the trained dynamic factor model, and state identification and indicator calculation are performed based on the regional transition probability to output the economic indicator prediction results.

2. The method according to claim 1, characterized in that, The multi-source economic time series data includes statistical indicators and survey indicators; the predetermined number of predictive variables are selected from the multi-source economic time series data based on a preset machine learning screening algorithm, including: The target hyperparameter combination of the machine learning screening algorithm is determined by the rolling window validation method. The target hyperparameter combination is applied to control the sparsity and regularization strength in the variable selection process, and the regression coefficient of each candidate variable is calculated by the least angle regression method. Each candidate variable is a potential predictor variable pre-selected from statistical indicators and survey indicators based on economic prosperity analysis theory. Based on the non-zero coefficients in the regression coefficients, a predetermined number of predictive variables are selected from the candidate variables.

3. The method according to claim 1, characterized in that, The construction of a dynamic factor model with a three-zone Markov transition chain based on the multiple predictor variables includes: Construct a state-space model, which includes measurement equations and transition equations; The multiple predictor variables are incorporated as observation variables into the measurement equation, which is used to establish the mapping relationship between the observation variables and the state vector; The transition equation is configured as an autoregressive process with a time-varying intercept term to describe the dynamic evolution of the state vector; The three-zone Markov transition chain is embedded in the state-space model, such that the value of the time-varying intercept term is determined by the current state of the Markov transition chain, where the current state is the uplink state, the general downlink state, or the severe downlink state. Specific constraints are imposed on the transition probability matrix of the three-zone Markov transition chain, including prohibiting the mutual transition between the uplink state and the severe downlink state.

4. The method according to claim 3, characterized in that, The aforementioned dynamic factor model, based on the constructed model, performs parameter solving on historical time series data to generate a trained dynamic factor model including regional transition probabilities, comprising: The state-space model is solved using the Kalman filter algorithm, which includes a state prediction step and a state update step. In the state prediction step, based on the transition equation and the posterior state estimate of the previous time point, the prior prediction value of the state vector at the current time point is deduced. In the state update step, the prior prediction value is corrected based on the measurement equation and the observed variable value at the current time point to obtain the posterior estimate of the state vector at the current time point; The state prediction step and the state update step are executed iteratively, and the model parameters of the dynamic factor model and the zone transition probabilities of the three-zone Markov transition chain are solved by combining the expectation-maximization algorithm to obtain the trained dynamic factor model.

5. The method according to claim 4, characterized in that, The step of performing state identification and index calculation based on the regional transition probability and outputting economic indicator prediction results includes: Based on the regional transfer probability of the three-region Markov transfer chain, calculate the probability values ​​of the economic climate being in the upward state, the general downward state, and the severe downward state. The current and future economic conditions are classified into different regional categories based on the probability values. A common factor sequence is extracted from the posterior estimate of the state vector, and the common factor sequence represents a synthetic leading index sequence of economic prosperity. The economic indicator prediction results are generated by combining the regional category and the common factor sequence.

6. The method according to claim 3, characterized in that, Before embedding the three-region Markov transition chain into the state-space model, the method further includes: Define multiple state transition probability elements in the transition probability matrix, wherein the multiple state transition probability elements respectively represent the probability of transitioning from the uplink state, the normal downlink state, and the severe downlink state to any other state; Set the transition probability elements representing the direct transition from the uplink state to the severe downlink state and the direct transition from the severe downlink state to the uplink state to zero; Set the sum of all transition probability elements in each row of the transition probability matrix to one.

7. The method according to claim 4, characterized in that, The iterative execution of the state prediction and state update steps, combined with the expectation-maximization algorithm to solve for the model parameters of the dynamic factor model and the region transition probabilities of the three-region Markov transition chain, yields the trained dynamic factor model, including: Initialize the model parameters of the dynamic factor model and the initial values ​​of the regional transition probability; The expected step is to calculate the conditional expectation and covariance matrix of the state vector based on the current parameter values ​​using the Kalman filter algorithm. Perform a maximization step to update the model parameters and the estimated values ​​of the regime transition probability based on the conditional expectation and covariance matrix; The expected step and the maximization step are executed iteratively until the model parameters and the estimated values ​​of the regime transition probability satisfy the convergence condition; The model parameters and regional transition probabilities after iterative convergence are used as the objective solution parameters to complete the construction of the trained dynamic factor model.

8. An economic indicator forecasting system based on Markov transition chains, characterized in that, include: The time series data acquisition module is used to acquire multi-source economic time series data; The economic forecast variable screening module is used to select a predetermined number of forecast variables from the multi-source economic time series data based on a preset machine learning screening algorithm. The dynamic factor model construction module is used to construct a dynamic factor model configured with a three-zone Markov transition chain based on the multiple predictor variables. The three-zone Markov transition chain is used to divide the economic prosperity state into three discrete zones: an upward state, a general downward state, and a severe downward state. The dynamic factor model training module is used to perform parameter solving on historical time series data based on the constructed dynamic factor model, and generate a trained dynamic factor model including regional transition probabilities. The regional transition probabilities are defined in the transition matrix of the three-region Markov transition chain and are used to characterize the transition probability between different economic prosperity states. The real-time economic indicator prediction module is used to input real-time economic data into the trained dynamic factor model, perform state identification and indicator calculation based on the regional transition probability, and output the economic indicator prediction results.

9. An electronic device, characterized in that, include: At least one processor; and a memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-7.