Stock timing system based on hidden Markov model

Through the stock timing system based on the Hidden Markov model, the hidden state of the market is speculated and trading signals are generated, which solves the problem that traditional technical analysis methods are difficult to adapt to market dynamic changes, and improves the accuracy and return of investment decisions.

CN120013677APending Publication Date: 2025-05-16BEIYIN FINANCIAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510178069.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional technical analysis methods rely on fixed technical indicators and empirical rules, are difficult to adapt to market dynamic changes, and are easily affected by market noise, which cannot effectively reduce risks and improve returns.

Method used

A stock timing system based on the hidden Markov model is used to model the stock price sequence data statistical probability, infer the hidden state of the market, and generate trading signals based on these states.

Benefits of technology

The system can dynamically adapt to market changes, adjust its status and decision-making by continuously observing and learning market behavior, improving the accuracy and rate of return of investment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013677A_ABST
    Figure CN120013677A_ABST
Patent Text Reader

Abstract

The invention discloses a stock timing system based on a hidden Markov model, and the system comprises a data obtaining and preprocessing module which is used for obtaining and preprocessing data; the model design and training module is used for dividing a training set and an observation set and training the model; the time selection strategy module is used for state recognition, signal generation, income calculation and model performance judgment; and the strategy optimization module is used for optimizing the strategy. The method pays attention to the dependency relationship between the dynamic change of the market state and the time sequence, speculates the hidden state of the market by adopting the statistical probability model, makes a decision on the basis, and has great advantages in the aspect of dealing with the market uncertainty and complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of financial technology, and in particular is a stock timing system based on a hidden Markov model. Background Art

[0002] In recent years, the financial market has become highly competitive, driving institutional investors and hedge funds to rely on advanced forecasting models to gain competitive advantages. Investors and institutions hope to reduce risks and increase returns through more accurate forecasting models. In order to gain a leading position in the market, it is necessary to analyze massive amounts of data, dig out potential investment opportunities and risks, and thus improve the accuracy of investment decisions.

[0003] The system of stock prediction based on machine learning has emerged, mainly to cope with the complexity and uncertainty of the financial market, overcome the limitations of traditional analysis methods, and use big data and powerful computing power, combined with advanced algorithms, to improve the accuracy of predictions and the scientific nature of investment decisions. With the continuous development of data technology and machine learning technology, stock prediction based on machine learning will become more and more popular and play an important role in the financial market.

[0004] Traditional technical analysis, taking the stock price of a certain company as an example, uses three commonly used technical indicators for analysis and prediction:

[0005] 1. Calculate the Moving Average (MA)

[0006] Moving averages are commonly used tools in technical analysis that can help smooth price data and show price trends. For example, a simple moving average is calculated for the short term (5 days) and long term (10 days).

[0007] 5-day moving average = (216.86 + 232.07 + 222.62 + 232.10 + 219.80) / 5 = 224.29

[0008] 10-day moving average = (216.86+232.07+222.62+232.10+219.80+220.25+215.99+246.38+251.51+239.20) / 10 = 229.78

[0009] From the above calculation, we can see that the stock’s short term moving average (224.29) is lower than its long term moving average (229.78), which suggests that the current trend could be down.

[0010] 2. Calculate the Relative Strength Index (RSI)

[0011] RSI is a momentum indicator used to assess overbought or oversold conditions in stock prices. RSI values ​​range from 0 to 100, with RSI values ​​above 70 indicating overbought conditions and below 30 indicating oversold conditions. Calculating RSI requires data from a longer period of time, and the simplified calculation is as follows:

[0012] Assume that there are 5 rising days and 5 falling days in the past 10 days. The total increase is (232.07-227.69+232.25-232.41+224.90-221.19+225.42-216.80+244.21-253.60) = 22.85

[0015] The total decline is (227.69-216.86+232.41-222.62+232.10-224.70+221.19-219.80+226.00-225.42) = 30.99

[0016] RSI=100-(100 / (1+(22.85 / 30.99)))=100-(100 / 1.36)=26.47

[0017] The RSI is around 26.47, which suggests that the stock is oversold and could see a rebound.

[0018] 3. Volume Analysis

[0019] Looking at the volume over the past 10 days, you can see that there are days with particularly high volume, such as 07 / 24 / 2024 (167942900) and 07 / 29 / 2024 (129201800). These days of high volume may be related to major news or events, and usually indicate increased market attention to the stock.

[0020] Based on the above analysis, we can conclude that: from the analysis of moving average and RSI, the stock is currently in a downward trend and may be oversold, with the possibility of a rebound. The fluctuation of trading volume also indicates that the market has a greater interest in the stock.

[0021] The existing technical analysis solutions mainly have the following problems:

[0022] 1. Traditional technical analysis relies on specific technical indicators (such as moving averages, RSI, MACD, etc.) to analyze price and volume data. These indicators are calculated using predefined mathematical formulas and are used to identify market trends, momentum, or overbought / oversold conditions. Technical analysis is usually rule-driven, that is, buying and selling signals are generated based on fixed trading rules (such as moving average crossovers). These rules are usually empirical and based on patterns of historical market behavior.

[0023] 2. Traditional technical analysis is often based on current or recent data. For example, if the current short-term moving average breaks through the long-term moving average, a buy signal may be generated. The decision-making process usually relies on pre-set fixed rules, which may not be applicable under different market conditions.

[0024] 3. Traditional technical analysis is easily affected by market noise, and its rules are too rigid to adapt to different market conditions. Summary of the invention

[0025] In view of the above problems, the present invention is proposed to provide a stock timing system based on a hidden Markov model that overcomes the above problems or at least partially solves the above problems.

[0026] To achieve the above object, the present invention adopts the following technical solutions:

[0027] A stock timing system based on a hidden Markov model, the system comprising:

[0028] Data acquisition and preprocessing module, used to acquire data and preprocess the data;

[0029] Model design and training module, used to divide the training set and observation set, and train the model;

[0030] Timing strategy module, used for state identification and signal generation, calculating returns and judging model performance;

[0031] Strategy optimization module, used to optimize the strategy.

[0032] Optionally, the data acquisition and preprocessing module includes:

[0033] The data crawling unit obtains Nasdaq historical stock data through crawlers and selects the daily level data of the target stock for the past 10 years;

[0034] The data preprocessing unit generates the maximum difference between the daily highest and lowest prices, the increase or decrease between the closing price and the previous day, the increase or decrease between the closing price and the previous day, the price change rate, technical indicators and the transaction volume change rate, and normalizes the data.

[0035] Optionally, the daily level data includes opening price, closing price, highest price, lowest price and trading volume.

[0036] Optional, model design and training modules include:

[0037] Divide the training set and observation set units, use random sampling method to divide the data set into training set and test set. If the number of samples of different categories in the data set is unbalanced, use stratified sampling to ensure that the proportion of samples of each category in the training set and test set is the same as that in the original data set;

[0038] The model parameter initialization unit is initialized according to the state distribution in the historical data, initially set to a uniform distribution, and adjusted during the training process;

[0039] The model training unit uses the forward algorithm and the backward algorithm to calculate the forward probability and the backward probability, calculate the expected value, and update the model parameters;

[0040] The model evaluation unit uses the forward algorithm to calculate the fit of the model to the test data, calculates the accuracy, precision, and recall rate to evaluate the model, and adjusts the number of states and model complexity based on the model evaluation results;

[0041] The model result analysis unit analyzes the model results.

[0042] Optional, timing strategy modules include:

[0043] The state recognition unit decodes based on the final generated transition probability matrix and infers the multi-head strategy state sequence of the hidden state;

[0044] The signal generation unit generates a buy signal if the system predicts an increase; generates a sell signal if the system predicts a decrease; and generates a hold or wait-and-see signal if the system predicts a fluctuation.

[0045] Calculate the profit unit. In the conversion probability matrix generated by the model, if the signal of the second day is rising, buy on the same day; if the signal of the second day is falling, liquidate on the same day; if the signal of the second day is oscillating, hold the position and observe; finally, make daily predictions for the prediction set, accumulate daily profits, and calculate the final total profit.

[0046] Optionally, the timing strategy module includes generating a long strategy state sequence through the model, accumulating the final benefits and total assets calculated every day, judging the model performance through the final total asset profitability, and then assisting in investment decision-making.

[0047] Optionally, the strategy optimization module includes:

[0048] The backtest analysis unit backtests the generated trading signals based on historical data to evaluate the performance of the strategy;

[0049] Parameter optimization unit, which adjusts the parameters and trading strategies of the hidden Markov model to improve the stability and yield of the strategy;

[0050] The transaction cost consideration unit takes into account the costs in actual transactions during backtesting and adjusts the strategy to increase net returns.

[0051] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0052] 1. The present invention pays more attention to the dynamic changes of market status and the dependency of time series, adopts statistical probability model to infer the hidden status of the market, and makes decisions based on this, which has great advantages in dealing with market uncertainty and complexity.

[0053] 2. The present invention does not directly use technical indicators, but infers the hidden market status through the sequence data of stock prices, and generates trading signals based on these inferred states; the present invention focuses on the characteristics of time series data, especially the dependencies between these data at different time points, and predicts future market trends by modeling these dependencies.

[0054] 3. The present invention is based on the inferred hidden states and their transition probabilities, not only taking into account the current market behavior, but also taking into account the possible market states and their evolution; it is able to dynamically adapt to market changes and adjust its states and decisions by continuously observing and learning market behavior. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 A schematic diagram of a stock timing system flow based on a hidden Markov model is provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0057] See also Figure 1 This embodiment provides a stock timing system based on a hidden Markov model, the system comprising:

[0058] The data acquisition and preprocessing module is used to acquire and preprocess data.

[0059] The data acquisition and preprocessing modules include:

[0060] The data crawling unit obtains Nasdaq historical stock data through a crawler and selects the daily level data of the target stock for the past 10 years, wherein the daily level data includes the opening price, closing price, highest price, lowest price and trading volume.

[0061] Main technical implementation:

[0062] Get web page information: get_html(address, stock abbreviation)

[0063] Importing the requests library allows for conveniently sending HTTP requests to websites and obtaining response results. The requests module is more concise than the urllib module. To send an HTTP request using requests, you need to import the requests module first:

[0064] From Agent import Agent. Import the proxy library to generate a proxy pool and avoid having the IP blocked by being monitored.

[0065] The address for scraping stock data: https: / / www.nasdaq.com / symbol / Stock Abbreviation / historical.

[0066] Simulate a browser: Construct the header.

[0067] Get web page data: requests.post(url, data, headers, proxies)

[0068] Get detailed stock information: get_stock_detail(html_data)

[0069] From lxml import etree. Import the lxml library, a Python library for parsing XML and HTML. It provides efficient and easy-to-use APIs for parsing and manipulating XML / HTML documents. Among them, ETree is an important module in the lxml library, which provides a simple and easy-to-use interface for traversing, querying, and modifying XML documents. XPath expressions are used to query specific elements or attributes.

[0070] Based on elements.xpath, obtain element data such as date, opening price, highest price, lowest price, closing price, trading volume, etc.

[0071] Data preprocessing unit, which generates the maximum difference between the daily highest price and the lowest price, the increase or decrease of the closing price compared to the previous day, the percentage change of the closing price compared to the previous day, the price change rate, technical indicators, and the change rate of trading volume, and normalizes the data to ensure that each feature has a similar dimension in the model as the model input.

[0072] Model design and training module, used to divide the training set and the observation set and train the model.

[0073] The model design and training module includes:

[0074] Divide the data set into training set and observation set units, and use random sampling method to divide the data set into training set and test set. If the number of samples of different categories in the data set is unbalanced, use stratified sampling to ensure that the proportion of samples of each category in the training set and test set is the same as that in the original data set.

[0075] The following steps can usually be used to construct training and test sets from the dataset:

[0076] Use random sampling to split the dataset into training and test sets. Common ratios are 70% for training and 30% for testing, or 80% / 20%.

[0077] #Assume X is the feature and y is the label

[0078] Python core code:

[0079] X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=0.2,random_state=42)

[0080] If the number of samples of different categories in the dataset is unbalanced, stratified sampling can be used to ensure that the proportion of samples of each category in the training set and test set is the same as in the original dataset.

[0081] Python core code:

[0082] X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=0.2,stratify=y,random_state=42)

[0083] The model parameter initialization unit is initialized according to the state distribution in the historical data. It is initially set to a uniform distribution and adjusted during the training process.

[0084] Model parameter initialization specifically includes:

[0085] (1) Initial state distribution: It is usually set to a uniform distribution, but can also be initialized based on the state distribution in historical data.

[0086] (2) State transfer matrix: It can be initially set to a uniform distribution and then adjusted through the training process.

[0087] (3) Observation probability matrix: initialized according to the distribution of observed data, usually assuming that the observed data obeys a certain probability distribution (such as Gaussian distribution).

[0088] (4) n_components = 3: indicates that three hidden layer states are used

[0089] (5) covariance_type = 'diag': defines the covariance matrix as a diagonal matrix, that is, the Gaussian distribution of each feature has its own variance parameter and is independent of each other.

[0090] (6) n_iter = 10000: defines the maximum number of iterations for Baum-Welch unsupervised learning

[0091] (7) fit(X): After the function completes training, model has become an HMM model that fits the stock's rise and fall.

[0092] The model training unit uses the forward algorithm and the backward algorithm to calculate the forward probability and the backward probability, calculate the expected value, and update the model parameters until the parameters converge.

[0093] The model evaluation unit uses the forward algorithm to calculate the model's fit to the test data, calculates the accuracy, precision, and recall rate to evaluate the model, and adjusts the number of states and model complexity based on the model evaluation results.

[0094] The model result analysis unit analyzes the model results.

[0095] The analysis of the model results includes:

[0096] (1) Mean matrix: Since the Baum-Welch unsupervised training method is used, no information about the implicit state is given during training.

[0097] Therefore, it is necessary to observe the specific values ​​of the mean matrix after training to understand the meaning of each hidden state.

[0098] #The matrix has three rows, each row represents a state of the node in the hidden layer (flat, up, down), and the two elements in each row represent: the mean of the ups and downs and the mean of the transaction volume

[0099] #Since the goal of the system is to predict the rise and fall, we only look at the data in the first column: the mean of state 0 is -0.0075 (close to zero), which means that the state is flat;

[0100] #The mean value of state 1 is 0.15 (an increase of more than 1 cent), indicating that the state is rising;

[0101] #The mean value of state 3 is -0.049 (a drop of more than 4 points), indicating that the state is falling;

[0102] #Then state 2 is oscillation;

[0103] (2) Covariance matrix

[0104] The covariance matrix of the eigenvalues ​​in each hidden layer state

[0105] Here we only consider the covariance of the rise and fall features (the upper left corner of each two-dimensional matrix)

[0106] The variance of the state "flat" is 0.111, and the credibility is medium

[0107] The variance of the state "Up" is 1.317, which is the maximum variance among the three states. That is, the range of change of this state is large and it is the least credible.

[0108] The variance of the state "fall" is 0.074, which is the smallest variance among the three states, which means that the prediction of this state is very reliable.

[0109] With u and sigma, we can use the normal distribution table to estimate the confidence interval of the increase or decrease.

[0110] (3) Transition probability matrix

[0111] The three lines still represent the states of the three hidden layers: shock, rise, and fall

[0112] #The maximum value in the first row: 0.7415, so the "oscillation" tends to maintain its own state, that is, the probability of still oscillating on the second day is high

[0113] #The maximum value of the second line: 0.5756. After learning that it is "rising", the next day tends to change to "rising", but there is also a high probability that it will "oscillate" (0.4234), and the possibility of falling after rising is relatively small.

[0114] #The maximum value in the third row is 0.7855. After learning that it has fallen, it is more likely that it will fall again the next day.

[0115] The timing strategy module is used for state identification and signal generation, calculating returns and judging model performance.

[0116] The timing strategy module includes:

[0117] The state recognition unit decodes based on the final generated transition probability matrix and infers the multi-head strategy state sequence of the hidden state.

[0118] The signal generating unit generates a buy signal if the system predicts an increase; generates a sell signal if the system predicts a decrease; and generates a hold or wait-and-see signal if the system predicts a shock.

[0119] Calculate the profit unit. In the conversion probability matrix generated by the model, if the signal of the second day is rising, buy on the same day; if the signal of the second day is falling, liquidate on the same day; if the signal of the second day is oscillating, hold the position and observe; finally, make daily predictions for the prediction set, accumulate daily returns, and calculate the final total return. If the market value increases by more than 30%, the model is considered to be effective.

[0120] The specific calculation logic is as follows:

[0121] When the buy signal is up on the next day and there is no position held, hold the full position:

[0122] Purchase share = balance / yesterday's closing price

[0123] Balance = Total assets – Yesterday’s closing price * Purchase shares

[0124] Market value = today's closing price * purchase shares

[0125] Average holding price = yesterday's closing price

[0126] Floating profit and loss = (today's closing price – yesterday's closing price) * purchase shares

[0127] Total assets = balance + floating profit and loss

[0128] On the second day, when the buy signal is up and the position is already held, if there is spare money, hold the full position:

[0129] Purchase share = balance / yesterday's closing price

[0130] Cost = yesterday's shares * average holding price + purchased shares * yesterday's closing price

[0131] Total shares = yesterday's shares + purchased shares

[0132] Average holding price = cost / total shares

[0133] Market value = total shares * today's closing price

[0134] Total assets = market value + balance

[0135] Floating profit and loss = (today's closing price – average holding price) * total shares

[0136] When the buy signal is down on the next day, sell all:

[0137] Remaining amount = yesterday's balance + total shares * yesterday's closing price

[0138] Floating profit and loss = (yesterday's closing price – average holding price) * total shares

[0139] When the buy signal on the next day is hold, hold the position and observe without taking any action.

[0140] The timing strategy module also includes generating a long strategy state sequence through the model, accumulating the final benefits and total assets calculated every day, judging the model performance through the final total asset profitability, and then assisting in investment decisions.

[0141] Strategy optimization module, used to optimize the strategy.

[0142] The strategy optimization module includes:

[0143] The backtest analysis unit backtests the generated trading signals based on historical data to evaluate the performance of the strategy (return, risk, winning rate, etc.).

[0144] The parameter optimization unit adjusts the parameters of the hidden Markov model (such as the number of states, the selection of observed variables) and the trading strategy (such as the setting of stop-loss and take-profit lines) to improve the stability and yield of the strategy.

[0145] The transaction cost consideration unit takes into account the costs of actual transactions (such as commissions and slippage) in the backtest and adjusts the strategy to increase net returns.

[0146] Through the above detailed technical implementation scheme, a stock timing system based on the hidden Markov model can be constructed. The system uses the Markov model to identify the market state and generates trading signals based on the identification results to assist investment decisions. The system needs to be fully backtested and optimized before actual deployment to ensure its performance in the real market environment.

[0147] This embodiment pays more attention to the dynamic changes of market status and the dependencies of time series, adopts statistical probability models to infer the hidden status of the market, and makes decisions based on this, which has great advantages in dealing with market uncertainty and complexity.

[0148] This embodiment does not directly use technical indicators, but instead infers hidden market states through stock price sequence data, and generates trading signals based on these inferred states; the present invention focuses on the characteristics of time series data, especially the dependencies between these data at different time points, and predicts future market trends by modeling these dependencies.

[0149] This embodiment is based on the inferred hidden states and their transition probabilities, and takes into account not only the current market behavior but also the possible market states and their evolutions; it is able to dynamically adapt to market changes and adjust its states and decisions by continuously observing and learning market behavior.

[0150] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A stock timing system based on hidden Markov model, characterized in that: The system comprises: Data acquisition and preprocessing module, used to acquire data and preprocess the data; Model design and training module, used to divide the training set and observation set, and train the model; Timing strategy module, used for state identification and signal generation, calculating returns and judging model performance; Strategy optimization module, used to optimize the strategy.

2. A stock timing system based on a hidden Markov model as claimed in claim 1, characterized in that: The data acquisition and preprocessing modules include: The data crawling unit obtains Nasdaq historical stock data through crawlers and selects the daily level data of the target stock for the past 10 years; The data preprocessing unit generates the maximum difference between the daily highest and lowest prices, the increase or decrease between the closing price and the previous day, the increase or decrease between the closing price and the previous day, the price change rate, technical indicators and the transaction volume change rate, and normalizes the data.

3. A stock timing system based on a hidden Markov model as claimed in claim 2, characterized in that: The daily level data includes opening price, closing price, highest price, lowest price and trading volume.

4. A stock timing system based on a hidden Markov model as claimed in claim 1, characterized in that: The model design and training modules include: Divide the training set and observation set units, use random sampling method to divide the data set into training set and test set. If the number of samples of different categories in the data set is unbalanced, use stratified sampling to ensure that the proportion of samples of each category in the training set and test set is the same as that in the original data set; The model parameter initialization unit is initialized according to the state distribution in the historical data, initially set to a uniform distribution, and adjusted during the training process; The model training unit uses the forward algorithm and the backward algorithm to calculate the forward probability and the backward probability, calculate the expected value, and update the model parameters; The model evaluation unit uses the forward algorithm to calculate the fit of the model to the test data, calculates the accuracy, precision, and recall rate to evaluate the model, and adjusts the number of states and model complexity based on the model evaluation results; The model result analysis unit analyzes the model results.

5. A stock timing system based on a hidden Markov model as claimed in claim 1, characterized in that: The timing strategy module includes: The state recognition unit decodes based on the final generated transition probability matrix and infers the multi-head strategy state sequence of the hidden state; The signal generation unit generates a buy signal if the system predicts an increase; generates a sell signal if the system predicts a decrease; and generates a hold or wait-and-see signal if the system predicts a fluctuation. Calculate the profit unit. In the conversion probability matrix generated by the model, if the signal of the second day is rising, buy on the same day; if the signal of the second day is falling, liquidate on the same day; if the signal of the second day is oscillating, hold the position and observe; finally, make daily predictions for the prediction set, accumulate daily profits, and calculate the final total profit.

6. A stock timing system based on a hidden Markov model as claimed in claim 1, characterized in that: The timing strategy module includes generating a long strategy state sequence through the model, accumulating the final benefits and total assets calculated every day, judging the model performance through the final total asset profitability, and then assisting in investment decisions.

7. A stock timing system based on a hidden Markov model as claimed in claim 1, characterized in that: The strategy optimization module includes: The backtest analysis unit backtests the generated trading signals based on historical data to evaluate the performance of the strategy; Parameter optimization unit, which adjusts the parameters and trading strategies of the hidden Markov model to improve the stability and yield of the strategy; The transaction cost consideration unit takes into account the costs in actual transactions during backtesting and adjusts the strategy to increase net returns.