Highly-frequent transaction-adaptive intelligence system
The integrated system for high-frequency trading addresses label imbalance, domain shift, and noise by processing data at multiple time scales, dynamically adjusting class weights, and reducing noise, enhancing prediction accuracy and adaptability in high-frequency trading.
Patent Information
- Application Number
- JP2025042975
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-01
AI Technical Summary
High-frequency trading faces challenges such as label imbalance, temporal domain shift, high noise levels, and computational efficiency, with existing solutions failing to provide an integrated approach that adapts to the dynamic nature of the market.
A comprehensive system integrating a multi-resolution time processing engine, adaptive label balancing framework, domain adaptation module, ensemble model architecture, and noise reduction component to process market data at multiple time scales, dynamically adjust class weights, adapt to market shifts, and reduce noise, while maintaining computational efficiency.
The system enhances prediction accuracy, adaptability, and computational efficiency, improving profitability and reducing unnecessary transactions and model drift, by effectively handling label imbalance, domain shift, and noise in high-frequency trading environments.
Abstract
Description
Technical Field
[0001] The present invention relates to a prediction model for high-frequency trading (HFT) in the financial market, and more particularly to an adaptive intelligence system for addressing the label imbalance problem. More specifically, the present invention relates to a comprehensive system that integrates market data processing at multiple time scales, dynamic label balancing, domain adaptation, model ensemble, and noise reduction.
Background Art
[0002] High-frequency trading (HFT) is a trading strategy in the financial market that uses computer algorithms to process a large number of orders in a very short time frame (from milliseconds to microseconds). This trading method aims to identify market inefficiencies and temporary price distortions and profit from them. High-frequency trading plays an important role in modern financial markets and has reached the point of accounting for more than 50% of trading volume in some markets. Characteristics of high-frequency trading include an extremely short holding period, high turnover rate, a large number of trades, and an extremely small profit margin. These characteristics present unique challenges and opportunities different from traditional trading methods. The success of high-frequency trading depends largely on the ability to accurately predict market trends. This prediction is usually performed by a machine learning model trained based on past market data. These models use various indicators such as price fluctuations, trading volume, order book imbalance, bid-ask spread, etc. as inputs to predict the direction and magnitude of future price fluctuations. In the context of high-frequency trading, the prediction time frame is very short and generally ranges from a few milliseconds to a few minutes. Predictions within such a short time frame present many technical challenges such as market noise, high dimensionality of data, and non-stationarity. However, there are several important challenges in developing a prediction model for high-frequency trading. Among them, the problem of label imbalance is particularly important. Label imbalance refers to a situation in which the number of samples in different classes (e.g., price increase, price decrease, no price change) in a prediction task is significantly different. In the context of high-frequency trading, especially in a short time frame (e.g., 1 minute), the profit opportunities (positive or negative labels) that exceed transaction costs are significantly fewer than those that do not (neutral labels). This imbalance causes problems such that standard machine learning algorithms are biased towards the majority class and the prediction accuracy of the minority class decreases. Specifically, in a typical high-frequency trading dataset, there is an extreme imbalance where neutral labels (small price changes that cannot cover transaction costs) account for approximately 80% of all samples, while positive labels (significant price increases) and negative labels (significant price decreases) each account for approximately 10%. The causes of this imbalance include market efficiency (prices quickly reflect new information), transaction costs (spread, fees, slippage, etc.), and market noise (temporary price fluctuations masking true price movements). In such a situation, a model that simply predicts the majority class can achieve high accuracy, but it cannot generate profits in actual trading. Traditional approaches to address label imbalance include data preprocessing techniques (such as oversampling, undersampling, hybrid methods, etc.), cost-sensitive learning (considering misclassification costs for different classes), ensemble methods (combining multiple classifiers), etc. Oversampling techniques include simple random oversampling (ROS) that randomly replicates samples of the minority class, and synthetic minority oversampling technique (SMOTE) that generates synthetic samples of the minority class. Undersampling techniques include random undersampling (RUS) that randomly deletes samples of the majority class, and clustering-based undersampling that deletes samples of the majority class based on clustering. In cost-sensitive learning, different costs are assigned to misclassifications of different classes, and the model is trained to minimize the total cost. In ensemble methods, the robustness to imbalanced data is improved by combining multiple models trained on different subsamples. However, these methods often cannot adequately address the complex issues of high-frequency trading on their own. For example, oversampling increases the risk of overfitting, and undersampling may lose important information. Cost-sensitive learning is difficult to set an appropriate cost matrix, and ensemble methods have high computational costs. Furthermore, these methods are usually based on the static characteristics of the data and do not consider the dynamic nature of the market. In a high-frequency trading environment, the market situation is constantly changing, and the label distribution may also change over time. Therefore, static label balancing techniques may become less effective over time. Furthermore, prediction models in high-frequency trading face additional challenges such as temporal domain shift (the phenomenon where the market distribution changes over time) and high noise-to-signal ratios (where market microstructure noise complicates the extraction of predictive signals), in addition to label imbalance. Temporal domain shift is caused by changes in market regimes, the introduction of new regulations, changes in macroeconomic factors, etc. This leads to the problem that the distribution of training data and test data is different. High noise-to-signal ratios are caused by market microstructure noise such as bid-ask bounce, market impact, order splitting, etc. This makes it difficult to extract the true price signal. In the prior art, approaches to address these challenges individually have been proposed, but an integrated system that simultaneously considers label imbalance, domain shift, and high noise levels has not been well studied. For example, many methods have been proposed to address label imbalance, but most of them assume a static environment and do not consider the dynamic nature of the market. Similarly, methods for domain adaptation and noise reduction have been proposed, but they are often not considered in combination with the problem of label imbalance. These challenges are interrelated, and a comprehensive solution is needed. In particular, the understanding of how these challenges interact in the specific context of high-frequency trading is limited. In high-frequency trading, predictions are required within an extremely short time frame, the market noise is very high, and the market situation changes rapidly. In such an environment, the problems of label imbalance, domain shift, and high noise levels are intricately intertwined, which may limit the effectiveness of conventional methods. Therefore, a new approach to comprehensively address these challenges is needed. Furthermore, the requirement for computational efficiency in a high-frequency trading environment is also an important issue. In high-frequency trading, latency in milliseconds can significantly impact the profitability of trades. Therefore, the prediction model needs to generate predictions in an extremely short time while maintaining high accuracy. This may limit the implementation of complex algorithms. For example, deep learning models may offer high prediction accuracy, but their computational cost may not be acceptable within the time constraints of high-frequency trading. Therefore, a new approach is needed to balance accuracy and computational efficiency. As described above, there are many interrelated issues in the development of prediction models in high-frequency trading, such as label imbalance, domain shift, high noise levels, and computational efficiency. A new approach is needed to comprehensively address these issues.
Summary of the Invention
Problems to be Solved by the Invention
[0003] The object of the present invention is to solve the following interrelated problems in high-frequency trading: The first problem is the serious label imbalance problem. In high-frequency trading, especially in short-term profit prediction (e.g., within 1 minute), profitable trading opportunities (labels +1 and -1) are significantly fewer than unprofitable opportunities (label 0). Typically, there is an extreme imbalance where label 0 accounts for about 80% of all samples, and labels +1 and -1 each account for about 10%. This imbalance causes problems where standard machine learning algorithms are biased towards the majority class and the prediction accuracy of the minority class decreases. Specifically, the model may be biased to always predict "no trade" (label 0), potentially missing profitable opportunities. Also, this imbalance is not static and may change dynamically according to market conditions. For example, during periods of high market volatility, the proportion of profitable opportunities (labels +1 and -1) may increase. Therefore, static label balancing techniques are insufficient, and techniques that can adapt to the dynamic nature of the market are needed. The second challenge is the problem of temporal domain shift. The distribution of financial data changes over time, and due to changes in market regimes, introduction of new regulations, changes in macroeconomic factors, etc., the underlying statistical characteristics of the data change. As a result, a "model drift" phenomenon occurs where the performance of a model trained on past data deteriorates when applied to new market situations. In a dynamic environment such as high-frequency trading, the ability to continuously adapt to this domain shift is essential. Specifically, the problem arises that the distribution of training data and test data is different, which reduces the prediction accuracy. For example, when a model trained during normal times is applied during a market stress period, its prediction ability may significantly decline. Also, this domain shift can occur not only at the feature level (e.g., changes in the distribution of price, trading volume, etc.) but also at the label level (e.g., changes in the distribution of profit opportunities). Therefore, a method that can handle domain shift at both the feature and label levels is needed. The third challenge is the problem of high noise-to-signal ratio. Financial market data, especially high-frequency data, contains market microstructure noise (bid-ask bounce, market impact, order splitting, etc.), which significantly complicates the extraction of prediction signals. This noise can increase the risk of overfitting of the model and reduce its generalization ability. Effective noise reduction techniques are essential for improving prediction accuracy. Specifically, the problem occurs that noise obscures the signal and makes it difficult to identify the true price trend. For example, bid-ask bounce (the phenomenon where price bounces between the bid and ask prices) can generate false signals regarding the price direction. Also, market impact (the phenomenon where a large order temporarily affects the price) can also cause temporary price fluctuations and distort the true price trend. Therefore, a method that can extract the true price signal and effectively reduce noise is needed. The fourth problem is the lack of an integrated approach to address these issues. In the prior art, approaches to individually address these problems have been proposed, but an integrated system that simultaneously considers label imbalance, domain shift, and high noise levels has not been sufficiently studied. These problems are interrelated and a comprehensive solution is needed. For example, if a method for addressing label imbalance is not robust to domain shift, its effectiveness may decline over time. Similarly, if noise reduction techniques do not consider label imbalance, they may erroneously remove important features of the minority class as noise. Therefore, an integrated approach that comprehensively considers these problems and their interactions is required. The fifth problem is the requirement of computational efficiency. In a high-frequency trading environment, the prediction model needs to operate within strict time constraints in milliseconds, which may limit the implementation of complex algorithms. Specifically, there is a problem that a complex model (e.g., a deep learning model) that may provide high prediction accuracy cannot be used within the time constraints of high-frequency trading due to its computational cost. For example, a complex ensemble model may provide high prediction accuracy, but if its inference time is too long, it is not suitable for high-frequency trading. Also, real-time feature engineering and preprocessing may have high computational costs, which may reduce the efficiency of the overall prediction pipeline. Therefore, a new approach is needed to balance high prediction accuracy and computational efficiency. To solve the above problems, the present invention proposes an integrated system that simultaneously considers label imbalance, domain shift, high noise levels, and the requirement of computational efficiency. This system aims to adapt to the dynamic nature of the market and provide robust predictions for market conditions that change over time.
Means for Solving the Problems
[0004] To address the above problems, the present invention proposes a High-Frequency Trading Adaptive Intelligence System (HFT-AIS) that integrates multiple innovative components. This system provides a comprehensive solution to address the interrelated problems of label imbalance, domain shift, and high noise levels. Hereinafter, the main components of the present invention and their functions will be described in detail. The first component is a multi-resolution time processing engine. This component simultaneously processes market data at multiple time scales (e.g., 5 seconds, 15 seconds, 30 seconds, 60 seconds). At each time scale, a dedicated feature extractor extracts the relevant temporal features. This enables capturing both the micro market structure (short-term patterns) and the macro market trends (long-term patterns). The intuition behind this approach is that different time scales can capture different aspects of the market. For example, a short time scale (5 seconds) can capture the micro structure of the market (order arrivals, cancellations, executions, etc.), while a long time scale (60 seconds) can capture more persistent market trends. By integrating information from these different time scales, a more complete understanding of the complex dynamics of the market can be achieved. The multi-resolution time processing engine includes a time scale setting unit, a feature extractor generation unit, and a feature integration unit. The time scale setting unit sets the multiple time scales to be used for analysis. By default, four time scales of 5 seconds, 15 seconds, 30 seconds, and 60 seconds are used, but they can be customized according to the characteristics of the market and trading strategies. For example, for a more high-frequency trading strategy, short time scales such as 1 second, 3 seconds, 5 seconds, 10 seconds, etc. may be appropriate. On the other hand, for a more long-term trading strategy, long time scales such as 1 minute, 5 minutes, 15 minutes, 30 minutes, etc. may be appropriate. The time scale setting unit can also dynamically adjust these time scales. For example, during periods of high market volatility, shorter time scales can be emphasized, and during periods of low volatility, longer time scales can be emphasized. The feature extractor generation unit generates feature extractors corresponding to each time scale. Each feature extractor extracts relevant features from market data at the corresponding time scale. The features to be extracted include the following: 1. Price features: average price, standard deviation, skewness, kurtosis, maximum value, minimum value, range, etc. These features capture the distribution and volatility of prices. For example, the standard deviation of price measures the market volatility, and the skewness of price measures the asymmetry of the price distribution. 2. Trading volume features: total trading volume, average, standard deviation, ratio to past periods, etc. These features capture the activity level and liquidity of the market. For example, an increase in trading volume may indicate an increase in market interest or the release of important news. 3. Order book features: bid-ask spread, order book imbalance, order book depth, order arrival rate, etc. These features capture the balance between market supply and demand. For example, the order book imbalance (the difference in the quantity of buy and sell orders) may indicate the future direction of the price. 4. Technical indicators: moving average, relative strength index (RSI), Bollinger Bands, MACD (Moving Average Convergence Divergence), etc. These indicators capture price trends, momentum, overbought / oversold conditions, etc. 5. Market microstructure features: effective spread, price impact, liquidity indicator, volatility indicator, etc. These features capture the microstructure of the market and transaction costs. For example, the effective spread measures the actual cost of a trade, and the price impact measures the impact of a large order on the price. Each feature extractor is implemented using a Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM) network, or attention mechanism. These architectures can effectively extract temporal patterns from time series data. CNNs are suitable for detecting local temporal patterns (e.g., price spikes, sudden increases in trading volume, etc.). LSTMs are suitable for capturing long-term dependencies (e.g., the relationship between past price patterns and future price fluctuations). The attention mechanism can dynamically evaluate the importance of different time points and focus on important time points. By combining these architectures, complex patterns in time series data can be effectively captured. Specifically, the CNN-based feature extractor consists of a 1D convolutional layer, a pooling layer, a batch normalization layer, and an activation layer (such as ReLU, LeakyReLU, ELU, etc.). The convolutional layer uses multiple filters with different kernel sizes (e.g., 3, 5, 7, etc.) to detect temporal patterns at different scales. The pooling layer reduces the size of the feature map and improves computational efficiency. The batch normalization layer reduces internal covariate shift and improves the stability of training. The activation layer introduces non-linearity and improves the model's representation ability. The LSTM-based feature extractor consists of an LSTM layer, a dropout layer, and a fully connected layer. The LSTM layer uses memory cells with input gates, forget gates, and output gates to learn long-term dependencies. The dropout layer prevents overfitting and improves the model's generalization ability. The fully connected layer converts the output of the LSTM layer into the final feature representation. By using bidirectional LSTM (BiLSTM), feature extraction considering both past and future contexts becomes possible. The attention mechanism-based feature extractor is composed of a self-attention layer (self-attention layer), a feed-forward layer, a residual connection, and layer normalization. The self-attention layer calculates how each position in the input sequence is related to all other positions. This enables effective capture of long-range dependencies within the sequence. The feed-forward layer further transforms the output of the attention mechanism. The residual connection alleviates the vanishing gradient problem and enables training of deeper networks. Layer normalization improves the stability of training. By using an attention mechanism based on the Transformer architecture, complex temporal relationships can be effectively captured. The feature integration part hierarchically integrates the features extracted by each feature extractor. The integration process is implemented using a cross-scale attention mechanism. This mechanism dynamically evaluates the importance of features at each time scale and weights and integrates the features based on that. Specifically, the integrated feature is calculated using the following formula: F_integrated = Σ(α_i * F_i) Here, F_integrated is the integrated feature, F_i is the feature at time scale i, and α_i is the importance (attention score) of time scale i. The attention score is calculated based on the current market situation and can capture the importance of different time scales in different market regimes. For example, during periods of high market volatility, the importance of short time scales may increase. Conversely, during periods when the market is following a trend, the importance of long time scales may increase. For the calculation of the attention score, an additive attention mechanism or a multiplicative attention mechanism can be used. In the additive attention mechanism, the attention score is calculated using the following formula: e_i = v^T * tanh(W_1 * F_i + W_2 * q) α_i = softmax(e_i) Here, \(F_i\) is the feature at time scale \(i\), \(q\) is the query vector (representing the current market situation), \(W_1\), \(W_2\), \(v^T\) are learnable parameters, \(\tanh\) is the hyperbolic tangent function, and \(\text{softmax}\) is the softmax function. In the multiplicative attention mechanism, the attention score is calculated using the following formula: \(e_i=(F_i * W_1)*(q * W_2)^T / \sqrt{d}\) \(\alpha_i=\text{softmax}(e_i)\) Here, \(d\) is the dimensionality of the feature, and \(\sqrt{}\) is the square root function. Furthermore, the feature integration part can also use the multi - head attention mechanism to capture the interaction between features of different time scales. In the multi - head attention mechanism, multiple attention mechanisms (heads) are applied in parallel, and their outputs are concatenated to generate the final feature representation. This enables capturing different types of temporal relationships simultaneously. In addition to the attention mechanism, the feature integration part also uses conditional feature modulation techniques such as gated linear units (GLU) and fast weight adaptation (FiLM). These techniques allow dynamically adjusting the importance of features according to the market situation. For example, GLU is defined by the following formula: \(\text{GLU}(F_i,q)=\sigma(W_1 * q)\odot(W_2 * F_i)\) Here, \(\sigma\) is the sigmoid function, \(\odot\) is the Hadamard product (element - wise product), and \(W_1\) and \(W_2\) are learnable parameters. FiLM is defined by the following formula: \(\text{FiLM}(F_i,q)=\gamma(q)\odot F_i+\beta(q)\) Here, \(\gamma(q)\) and \(\beta(q)\) are the scaling parameter and shift parameter calculated from the query vector \(q\). Finally, the feature integration part normalizes the integrated features and applies dimensionality reduction techniques (such as principal component analysis (PCA), t-SNE, UMAP, etc.) to reduce the dimensionality of the features and improve computational efficiency. This enables an efficient mapping from a high-dimensional feature space to a low-dimensional latent space. The second component is an adaptive label balancing framework. This framework dynamically adjusts the importance of each class (-1, 0, +1). Different from the conventional static weighting approach, this framework takes into account both the current distribution of each class and the recent performance of the model for each class. This enables adaptation to changes in market conditions and improvement in the prediction accuracy of minority classes (profitable trading opportunities). The adaptive label balancing framework includes a class distribution recording part, a weight holding part, a performance tracking part, and a weight calculation part. The class distribution recording part records the changes in class distribution over time. Specifically, it calculates the class distribution in each time window (e.g., 1 hour, 1 day, 1 week, etc.) and saves it as historical data. This enables tracking of the temporal variations in class distribution and identification of seasonality and changes in market regimes. For example, the class distribution may change before and after a specific time period (e.g., at the start or end of the market) or a specific market event (e.g., release of economic indicators, central bank policy announcements, etc.). By tracking such variations, the label balancing strategy can be adjusted more effectively. The class distribution recording part records the following information: 1. The absolute number and relative frequency of each class in each time window 2. Temporal variations (trends, seasonality, periodicity, etc.) of the class distribution 3. The relationship between the class distribution and market conditions (volatility, trading volume, spread, etc.) 4. Transition probabilities between classes (e.g., the probability that label +1 follows label 0) This information is stored in a time - series database and used for subsequent analysis and weight calculation. The time - series database supports high - speed data insertion and efficient querying, and can efficiently manage a large amount of historical data.
[0005] The weight holding part holds the current weights of each class. The initial weights are set based on the inverse frequency of the classes (for example, if the distribution of classes 0, - 1, + 1 is 8:1:1, the initial weights are 1:8:8). These weights are dynamically updated during the training process. The weight holding part holds the following information: 1. The current weights of each class 2. The weight update history 3. The parameters used for weight update (such as learning rate, regularization parameter, etc.) 4. The weight constraints (such as minimum value, maximum value, the constraint that the sum is 1, etc.) This information is stored in an in - memory data structure (such as a hash map, priority queue, etc.) and can be accessed and updated quickly. The weight holding part also applies smoothing techniques (such as exponential moving average, Kalman filter, etc.) to prevent sudden fluctuations in weights. The performance tracking part tracks the performance metrics of the model for each class. The metrics to be tracked include accuracy, recall, F1 - score, confusion matrix, area under the ROC curve (AUC), etc. These metrics are calculated at the end of each evaluation period (such as at the end of each training epoch, each trading day, etc.) and saved as historical data. The performance tracking part tracks the following information: 1. Basic classification metrics such as accuracy, recall, F1 - score, AUC for each class 2. The confusion matrix (number of true positives, false positives, true negatives, false negatives) for each class 3. The predicted probability distribution for each class (such as a histogram of predicted probabilities) 4. The misclassification cost for each class (such as the cost of false positives and false negatives) 5. The temporal variation of the performance for each class These information are stored in a time - series database and used for subsequent analysis and weight calculation. The performance tracking part can also update the performance indicators in real - time using an online learning algorithm. The weight calculation part calculates the optimal weights based on the distribution factor and the performance factor. Specifically, the weights for each class are calculated using the following formula: w_c = α * (N / N_c) + (1 - α) * (1 / P_c) Here, w_c is the weight of class c, N is the total number of samples, N_c is the number of samples in class c, P_c is the most recent performance of the model for class c (normalized to the range from 0 to 1), and α is a parameter (in the range from 0 to 1) that controls the balance between the distribution factor and the performance factor. The first term (N / N_c) of this formula is the weight based on the inverse frequency of the class, which assigns a high weight to the minority class. The second term (1 / P_c) is the weight based on the reciprocal of the performance, which assigns a high weight to the class with low performance. The α parameter controls the balance between these two factors. The closer the value of α is to 1, the greater the influence of the distribution factor, and the closer the value of α is to 0, the greater the influence of the performance factor. The value of α can be adjusted dynamically according to the market situation and the performance of the model. For example, if the model consistently shows low performance for a specific class, the value of α can be decreased to increase the influence of the performance factor. Conversely, if the class distribution is extremely unbalanced, the value of α can be increased to increase the influence of the distribution factor. The following formula can be used for the dynamic adjustment of α: α = sigmoid(β * (I_distribution - I_performance)) Here, sigmoid(x) = 1 / (1 + exp(-x)) is the sigmoid function, β is the parameter that controls the sensitivity of adjustment, I_distribution is the index that measures the degree of imbalance in class distribution (e.g., Gini coefficient, entropy, etc.), and I_performance is the index that measures the degree of imbalance in performance (e.g., standard deviation of performance between classes, etc.). Furthermore, the weight calculation unit calculates the weight using the class distribution and performance index within an exponentially decaying time window: N_c(t) = Σ(e^(-λ * (t - s)) * n_c(s)) P_c(t) = Σ(e^(-λ * (t - s)) * p_c(s)) / Σ(e^(-λ * (t - s))) Here, N_c(t) is the weighted number of samples of class c at time t, n_c(s) is the number of samples of class c at time s, P_c(t) is the weighted performance of class c at time t, p_c(s) is the performance of class c at time s, and λ is the decay rate. Due to this exponential decay, the weight is continuously updated based on the recent data and performance, and it can adapt to changes in the market situation. The decay rate λ can be adjusted according to the dynamics of the market. For example, during a period when the market is changing rapidly, the value of λ can be increased to assign a higher weight to the recent data. Conversely, during a period when the market is stable, the value of λ can be decreased to consider data over a longer period. For the dynamic adjustment of λ, the following formula can be used: λ = λ_base + γ * V Here, λ_base is the basic decay rate, γ is the parameter that controls the sensitivity of adjustment, and V is the index that measures the volatility of the market (e.g., standard deviation of price, mean absolute deviation, etc.). The weight calculation unit normalizes the calculated weights to ensure numerical stability and maintain the ratio of weights between different classes. The following formula can be used for normalization: w_c_normalized = w_c / Σ(w_c') Here, w_c_normalized is the normalized weight of class c, and Σ(w_c') is the sum of the weights of all classes. Furthermore, the weight calculation unit smooths the weights using the following formula to prevent rapid fluctuations in the weights: w_c_smoothed = (1 - β) * w_c_previous + β * w_c_current Here, w_c_smoothed is the smoothed weight of class c, w_c_previous is the weight of class c at the previous time point, w_c_current is the currently calculated weight of class c, and β is the smoothing parameter (in the range from 0 to 1). Finally, the weight calculation unit integrates the calculated weights into the loss function. By applying the weights to a general classification loss function (e.g., cross-entropy loss, focal loss, etc.), the importance of the minority classes can be enhanced. For example, the weighted cross-entropy loss is defined by the following formula: L = -Σ(w_c * y_c * log(p_c)) Here, L is the loss, y_c is the true label (0 or 1) of class c, and p_c is the predicted probability of class c. The third component is the domain adaptation module. This module continuously monitors the statistical characteristics of the training data (source domain) and the recent market data (target domain). Using these statistics, it transforms the input features to mitigate the impact of domain shift. This approach enables the model to remain effective even as market conditions evolve. The domain adaptation module includes a source statistics calculation unit, a target statistics calculation unit, and a feature adaptation unit. The source statistics calculation unit calculates and stores the statistical information of the training data. The calculated statistical information includes the mean, standard deviation, minimum value, maximum value, quartiles, skewness, kurtosis, etc. of each feature. These statistics can be calculated from the entire training dataset or calculated for each time window. With the latter approach, the temporal variations in the training data can be captured. The source statistics calculation unit calculates the following statistical information: 1. First-order statistics: mean, median, mode, minimum value, maximum value, range, quartiles, etc. 2. Second-order statistics: variance, standard deviation, mean absolute deviation, interquartile range, etc. 3. Higher-order statistics: skewness, kurtosis, entropy, etc. 4. Correlation statistics: correlation matrix, covariance matrix, principal components, etc. 5. Temporal statistics: autocorrelation, partial autocorrelation, spectral density, etc. These statistics are stored in an efficient data structure (e.g., hash map, B-tree, etc.) for fast access. The source statistics calculation unit uses robust estimation methods (e.g., median absolute deviation, winsorized mean, etc.) for calculating the statistics to minimize the impact of outliers. The target statistics calculation unit calculates and stores the statistical information of the recent market data. The calculated statistical information is the same as that of the source statistics calculation unit. However, these statistics are calculated from the latest market data (e.g., the past 1 hour, the past 1 day, etc.). The calculation frequency of the statistics is set according to the dynamics of the market and the constraints of the computing resources. For example, in a high-frequency trading environment, it may be appropriate to update the statistics every few seconds or minutes. On the other hand, for a more long-term trading strategy, it may be appropriate to update the statistics every few hours or days. The target statistics calculation unit provides the following functions: 1. Real-time statistics calculation: incrementally update the statistics every time a new data point arrives 2. Time Window Statistical Calculation: Calculate statistics based on data within a specified time window 3. Weighted Statistical Calculation: Calculate statistics by assigning higher weights to more recent data points 4. Outlier Detection and Handling: Detect outliers before statistical calculation and handle them appropriately 5. Confidence Interval Calculation for Statistics: Calculate confidence intervals to quantify the uncertainty of statistics The target statistics calculation unit also provides a function to track the temporal variation of the calculated statistics and detect abrupt changes in the statistics (which may indicate a change in the market regime, for example). The feature adaptation unit adapts features using the statistical information of the source domain and the target domain. The basic approaches are feature normalization and inverse normalization: 1. Normalize features using the mean and standard deviation of the source domain: x_norm = (x - μ_source) / σ_source 2. Inverse normalize the normalized features using the mean and standard deviation of the target domain: x_adapted = x_norm * σ_target + μ_target Here, x is the original feature, x_norm is the normalized feature, x_adapted is the adapted feature, μ_source and σ_source are the mean and standard deviation of the source domain, and μ_target and σ_target are the mean and standard deviation of the target domain.
[0006] This approach aligns the first moment (mean) and second moment (variance) of the features but does not consider differences in higher moments. As a more advanced approach, the feature adaptation unit also implements the following method: 1. Minimization of Maximum Mean Discrepancy (MMD): Minimize the distance between the feature distributions of the source domain and the target domain. MMD measures the distance between two distributions in a Reproducing Kernel Hilbert Space (RKHS). The minimization of MMD is formulated as the following optimization problem: min_θ MMD^2(X_source, X_target_θ) Here, X_source is the feature of the source domain, X_target_θ is the feature of the target domain transformed by parameter θ, and MMD^2 is the squared MMD distance. 2. Correlation Alignment (CORAL): Align the covariance matrices of the features of the source domain and the target domain. CORAL applies the following transformation: X_adapted = X * W Here, W = C_source^(-1 / 2) * C_target^(1 / 2), C_source is the covariance matrix of the source domain, and C_target is the covariance matrix of the target domain. 3. Adversarial Domain Adaptation: Use adversarial training to learn domain-invariant feature representations. In this method, a feature extractor and two classifiers (a label classifier and a domain classifier) are trained. The feature extractor is trained to maximize the performance of the label classifier and minimize the performance of the domain classifier. As a result, feature representations that are useful for label prediction but independent of the domain are learned. These methods can handle more complex domain shifts (e.g., non-linear transformations, changes in conditional distributions, etc.). The appropriate method is selected according to the characteristics of the market data and the nature of the domain shift. The feature adaptation part also provides the following functions: 1. Automatic adjustment of adaptation parameters: Adjust the adaptation parameters according to the magnitude of the domain shift 2. Gradual Adaptation: To avoid abrupt conversions, apply the adaptation gradually 3. Conditional Adaptation: Apply the adaptation only under specific conditions (e.g., when the domain shift exceeds a threshold) 4. Feature-Selective Adaptation: Apply the adaptation only to features that are susceptible to the influence of domain shift 5. Multi-Source Adaptation: Integrate information from multiple source domains to perform adaptation The feature adaptation component also implements a cache mechanism and parallel computing to improve the efficiency of the adaptation process. The cache mechanism stores frequently used statistics and transformation parameters in memory to reduce the computational overhead. Parallel computing uses multiple CPU cores or GPUs to parallelize the adaptation calculations and improve the processing speed. The fourth component is the ensemble model architecture. This architecture combines multiple prediction models such as multi-layer perceptrons (MLPs), long short-term memory (LSTM) networks, and Mamba models. The contribution of each model is dynamically adjusted based on recent performance. This approach enables leveraging the strengths of different models in different market situations. The ensemble model architecture includes a model holding part, a weight holding part, a performance recording part, a prediction aggregation part, and a weight update part. The model holding part holds multiple models with different architectures and learning algorithms. By default, the following three models are used: 1. Multi-Layer Perceptron (MLP): A feed-forward neural network with multiple hidden layers. Each layer is fully connected and uses a non-linear activation function (e.g., ReLU, LeakyReLU, SELU). The MLP can capture complex non-linear relationships between features. The structure of the MLP is as follows: - Input layer: The number of neurons corresponding to the dimension of the features - Hidden layer 1: 256 neurons, with the activation function LeakyReLU - Hidden layer 2: 128 neurons, with the activation function being LeakyReLU - Hidden layer 3: 64 neurons, with the activation function being LeakyReLU - Output layer: The number of neurons corresponding to the number of classes, with the activation function being softmax The MLP also includes a dropout layer (dropout rate 0.2) and a batch normalization layer. These layers prevent overfitting and improve the stability of training. 2. Long Short-Term Memory (LSTM) network: A type of recurrent neural network that can learn long-term dependencies. LSTM uses memory cells with input gates, forget gates, and output gates to perform selective retention and update of information. LSTM is suitable for capturing the temporal patterns of time series data. The structure of LSTM is as follows: - Input layer: The input size corresponding to the dimension of the features - LSTM layer 1: 128 units, bidirectional - Dropout layer: Dropout rate 0.3 - LSTM layer 2: 64 units, bidirectional - Dropout layer: Dropout rate 0.3 - Fully connected layer: 32 neurons, with the activation function being ReLU - Output layer: The number of neurons corresponding to the number of classes, with the activation function being softmax Using bidirectional LSTM enables prediction considering both past and future contexts. 3. Mamba model: A recent architecture based on the structured state space sequence model (SSM). Mamba uses a selection mechanism to dynamically adjust the parameters of the state space model according to the input. This approach is similar to the attention mechanism of transformers but allows for more efficient inference. Mamba can capture both long-distance dependencies and local patterns. The structure of Mamba is as follows: - Input layer: The input size corresponding to the dimension of the features - Embedding layer: Embedding of 128 dimensions - Mamba block 1: State size 64, expansion factor 4 - Mamba block 2: State size 64, expansion factor 4 - Mamba block 3: State size 64, expansion factor 4 - Mamba block 4: State size 64, expansion factor 4 - Fully connected layer: 32 neurons, activation function is ReLU - Output layer: Number of neurons corresponding to the number of classes, activation function is softmax Each mamba block is composed of layer normalization, SSM layer, feed-forward layer, and residual connection. In addition to these models, the model storage unit can also include conventional machine learning models such as support vector machines (SVM), gradient boosting trees (e.g., XGBoost, LightGBM), and random forests. This enables the realization of an ensemble of diverse models that capture different types of patterns. The model storage unit also saves the checkpoints of each model and can restore the model to its previous state as needed. This allows the model to be restored to a previous good state in case of performance degradation. The checkpoints include the model's parameters, hyperparameters, training state, and performance metrics. The weight storage unit holds the current weights of each model. The initial weights are set evenly (e.g., for three models, the weight of each model is 1 / 3). These weights are updated periodically based on the recent performance of each model. The weight storage unit holds the following information: 1. The current weights of each model 2. The weight update history 3. The parameters used for weight update (e.g., learning rate, regularization parameter, etc.) 4. Weight constraints (e.g., constraints such as minimum value, maximum value, and the sum being 1) These pieces of information are stored in in-memory data structures (such as hash maps, priority queues, etc.) and can be accessed and updated quickly. To prevent sudden fluctuations in weights, the weight holding unit also applies smoothing techniques (such as exponential moving average, Kalman filter, etc.). The performance recording unit records the past performance of each model. The performance metrics to be recorded include prediction accuracy, F1 score, profit, Sharpe ratio, etc. These metrics are calculated at the end of each evaluation period (such as each trading day, each week, etc.) and stored as historical data. The performance recording unit records the following information: 1. Basic classification metrics such as accuracy, recall, F1 score, AUC, etc. for each model 2. Confusion matrix (number of true positives, false positives, true negatives, false negatives) for each model 3. Predicted probability distribution for each model (such as a histogram of predicted probabilities) 4. Profitability metrics for each model (such as profit, Sharpe ratio, maximum drawdown, etc.) 5. Temporal variation of the performance of each model These pieces of information are stored in a time-series database and used for subsequent analysis and weight updates. The performance recording unit also provides a function to detect outliers in performance metrics and handle them appropriately. The prediction aggregation unit weights and aggregates the predictions of each model. The basic approach is a weighted average: y_final = Σ(w_i * y_i) / Σ(w_i) Here, y_final is the final prediction, y_i is the prediction of model i, and w_i is the weight of model i. When the prediction is in the form of a probability distribution (probability of each class), the prediction aggregation unit calculates the final probability distribution using the following formula: p_final(c) = Σ(w_i * p_i(c)) / Σ(w_i) Here, p_final(c) is the final probability of class c, and p_i(c) is the probability of class c by model i. As a more advanced approach, the prediction aggregation part can also implement stacking. In stacking, the predictions of each model are used as new features, and a meta-model (e.g., logistic regression, neural network, etc.) is trained to generate the final prediction. This approach can capture more complex relationships between models than simple weighted averaging. The implementation of stacking includes the following steps: 1. Train each base model with the training data 2. Use each base model to generate predictions for the validation data 3. Use these predictions as new features and train the meta-model 4. For the test data, generate predictions for each base model and use them as inputs to the meta-model to generate the final prediction As the meta-model, logistic regression, random forest, gradient boosting, neural network, etc. are used. The selection of the meta-model depends on the complexity of the problem and the constraints of computational resources. The prediction aggregation part also provides the following functions: 1. Conditional aggregation: Use different aggregation methods according to market conditions 2. Confidence weighting: Adjust the weights based on the confidence of the predictions of each model 3. Diversity promotion: Assign high weights to models that provide diverse predictions 4. Detection and handling of abnormal predictions: Detect abnormal predictions and handle them appropriately 5. Optimization of aggregation parameters: Optimize the aggregation parameters regularly
[0007] The weight update unit updates the weights based on the recent performance of each model. The basic approach is to calculate the recent performance of each model using exponential weighted averaging and update the weights based on it: w_i = exp(λ * p_i) / Σ(exp(λ * p_j)) Here, p_i is the recent performance of model i (e.g., accuracy, F1 score, profit, etc.), and λ is a parameter that emphasizes the difference in performance. The larger the value of λ, the higher the sensitivity to the difference in performance, and a higher weight is assigned to the model showing the best performance. For the calculation of p_i, performance metrics within an exponentially decaying time window are used: p_i = Σ(e^(-μ * (t - s)) * perf_i(s)) / Σ(e^(-μ * (t - s))) Here, perf_i(s) is the performance of model i at time s, and μ is the decay rate. Due to this exponential decay, a higher weight is assigned to the recent performance. As a more advanced approach, the weight update unit can also implement online learning algorithms (e.g., exponential weighted average predictor, hedge algorithm, etc.). These algorithms can adapt to the temporal variations in the performance of each model and emphasize different models in different market regimes. For example, the exponential weighted average predictor is defined by the following formula: w_i(t+1) = w_i(t) * exp(η * r_i(t)) w_i(t+1) = w_i(t+1) / Σ(w_j(t+1)) Here, \(w_i(t)\) is the weight of model \(i\) at time \(t\), \(r_i(t)\) is the reward of model \(i\) at time \(t\) (e.g., 1 for a correct prediction and 0 for an incorrect prediction), and \(\eta\) is the learning rate. The hedge algorithm is defined by the following formula: \(w_i(t + 1)=w_i(t)\times(1-\varepsilon+\varepsilon\times r_i(t) / r_{avg}(t))\) Here, \(\varepsilon\) is the hedge parameter and \(r_{avg}(t)\) is the average reward at time \(t\). The weight update part also provides the following functions: 1. Adaptive learning rate: Adjust the learning rate according to the performance variation 2. Regularization: Apply regularization to prevent excessive concentration of weights 3. Balance between exploration and exploitation: Balance the exploration of new models and the exploitation of known good models 4. Weight constraints: Apply constraints so that the weights fall within a specific range 5. Weight smoothing: Apply smoothing to prevent sudden fluctuations in weights The fifth component is a noise reduction component using wavelet decomposition. This component removes noise from the input data and improves the signal-to-noise ratio. Wavelet decomposition can analyze signals in both the time and frequency domains, so it is suitable for dealing with noises of various scales contained in financial time series data. The noise reduction component includes a wavelet setting part, a wavelet decomposition part, a threshold calculation part, a threshold processing part, and a signal reconstruction part. The wavelet setting part sets the wavelet type (e.g., Daubechies, Symlet, Coiflet, etc.) and the decomposition level to be used. The selection of the wavelet type depends on the characteristics of the signal. The following wavelet types are generally used for financial time series data: 1. Daubechies wavelet: There are different orders (such as db1, db2, db4, db8, etc.), and the balance between smoothness and locality is different. db1 is the most local but not the smoothest, and the higher the order, the smoother it becomes but the locality decreases. For financial time series data, db4 or db8 is generally used. 2. Symlet wavelet: A symmetric version of the Daubechies wavelet with less phase distortion. Symlets also have different orders (such as sym2, sym4, sym8, etc.), and the higher the order, the smoother it becomes. For financial time series data, sym4 or sym8 is generally used. 3. Coiflet wavelet: Has high vanishing moments and is suitable for approximating smooth signals. Coiflets also have different orders (such as coif1, coif2, coif3, etc.), and the higher the order, the smoother it becomes. For financial time series data, coif2 or coif3 is generally used. 4. Discrete Meyer wavelet: Defined in the frequency domain and has good time - frequency localization. It is suitable for the analysis of financial time series data. The selection of the wavelet type depends on the characteristics of the signal (such as smoothness, discontinuity, etc.) and the purpose of the analysis (such as noise reduction, feature extraction, etc.). A common approach is to try multiple wavelet types and select the one that provides the best results. The selection of the decomposition level depends on the length of the signal and the characteristics of the noise. Generally, levels from 3 to 5 are used. The higher the level, the more low - frequency noise can be removed, but the distortion of the signal may also increase. The upper limit of the decomposition level is restricted by the length of the signal. Specifically, the maximum decomposition level should not exceed log2(N) (where N is the length of the signal). The wavelet setting part provides the following functions: 1. Automatic Wavelet Selection: Select the optimal wavelet type based on the characteristics of the signal 2. Automatic Level Selection: Select the optimal decomposition level based on the length of the signal and the characteristics of the noise 3. Parameter Optimization: Optimize the wavelet parameters using cross-validation or information criteria 4. Adaptive Setting: Adapt the wavelet settings according to the temporal changes of the signal 5. Multi-Wavelet Analysis: Analyze the signal using multiple wavelet types and integrate the results The wavelet decomposition unit performs wavelet decomposition on the time-series data. Wavelet decomposition decomposes the signal into components of different scales. Specifically, the signal is decomposed into an approximation coefficient (low-frequency component) and detail coefficients (high-frequency components). In an n-level decomposition, one approximation coefficient and multiple detail coefficients are generated: [a_n, d_n, d_{n - 1}, ..., d_1] = DWT(x, n) Here, a_n is the approximation coefficient at level n, d_i is the detail coefficient at level i, DWT is the discrete wavelet transform, x is the input signal, and n is the decomposition level The approximation coefficient represents the low-frequency component of the signal and captures the overall trend. The detail coefficients represent the high-frequency components of the signal and capture the fine fluctuations and noise. Generally, since noise mainly appears in the detail coefficients, noise can be reduced by applying threshold processing to the detail coefficients The wavelet decomposition unit provides the following functions: 1. Efficient Implementation: Improve the computational efficiency using the fast wavelet transform (FWT) algorithm 2. Boundary Processing: Provide boundary processing methods (such as periodic extension, symmetric extension, etc.) to minimize distortion at the boundaries of the signal 3. Parallel Computation: Parallelize the decomposition calculation using multiple CPU cores or GPUs 4. Incremental decomposition: Update the decomposition incrementally every time a new data point arrives. 5. Multilevel analysis: Analyze the detail coefficients at different levels to identify noise at different scales. The threshold calculation unit calculates an adaptive threshold to be applied to the detail coefficients. The selection of the threshold affects the balance between the effect of noise reduction and the distortion of the signal. The following are common threshold selection methods: 1. Universal threshold: T = σ * sqrt(2 * log(N)) Here, σ is the noise level (usually estimated using the median absolute deviation (MAD) of the detail coefficients), and N is the data length. MAD is calculated by the following formula: σ = median(|d_i - median(d_i)|) / 0.6745 Here, d_i is the detail coefficient, and 0.6745 is the normalization coefficient of the median absolute deviation of the normal distribution. 2. SURE threshold (Stein's Unbiased Risk Estimator): Select the threshold that minimizes the risk function. The SURE threshold is calculated as the value that minimizes the risk function defined by the following formula: SURE(T) = N - 2 * #{i: |d_i| ≦ T} + Σ(min(|d_i|, T)^2) Here, #{i: |d_i| ≦ T} is the number of detail coefficients with absolute values smaller than T. 3. Minimax threshold: Select the threshold that minimizes the risk in the worst case. The minimax threshold is approximated by the following formula: T = σ * (0.3936 + 0.1829 * log2(N)) 4. Adaptive Threshold: Adjust the threshold based on the local characteristics of the detail coefficients. For example, use a low threshold in areas where the local variance of the detail coefficients is large, and use a high threshold in areas where the variance is small. The adaptive threshold is calculated by the following formula: T_i = σ_i * sqrt(2 * log(N)) Here, σ_i is the local standard deviation of the detail coefficients. The threshold calculation unit also provides the following functions: 1. Level-Dependent Threshold: Calculate different thresholds for each decomposition level 2. Signal-Dependent Threshold: Adjust the threshold based on the characteristics of the signal 3. Noise Estimation: Provide multiple methods to estimate the noise level from the detail coefficients 4. Threshold Optimization: Optimize the threshold using cross-validation or information criteria 5. Threshold Based on Confidence Interval: Set the threshold based on the confidence interval of the detail coefficients The threshold processing unit applies threshold processing to the detail coefficients to remove noise. Common threshold processing methods include the following: 1. Hard Thresholding: Set coefficients smaller than the threshold to zero and keep the other coefficients unchanged. d'_i = d_i * I(|d_i| > T) Here, d_i is the original detail coefficient, d'_i is the processed detail coefficient, T is the threshold, and I() is the indicator function. 2. Soft Thresholding: Set coefficients smaller than the threshold to zero and shrink the other coefficients by the amount of the threshold. d'_i = sign(d_i) * max(0, |d_i| - T) Here, sign() is the sign function. 3. Garrote Thresholding: An intermediate method between hard thresholding and soft thresholding. d'_i = d_i * max(0, 1 - (T / |d_i|)^2) Soft thresholding is commonly used in financial time series data to preserve the continuity of the signal. Hard thresholding preserves the amplitude of the signal but may introduce discontinuities. Galois thresholding has intermediate characteristics between these two methods. The thresholding unit also provides the following functions: 1. Adaptive thresholding: Adapt the thresholding method based on the local characteristics of the detail coefficients 2. Hybrid thresholding: Combine multiple thresholding methods 3. Hierarchical thresholding: Apply different thresholding methods at different levels 4. Block thresholding: Apply thresholding to blocks of detail coefficients 5. Probabilistic thresholding: Process detail coefficients based on a probability model
[0008] The signal reconstruction unit reconstructs the original signal from the processed coefficients. The reconstruction is performed using the inverse discrete wavelet transform (IDWT): x' = IDWT(a_n, d'_n, d'_{n - 1},..., d'_1) Here, x' is the reconstructed (denoised) signal, a_n is the approximation coefficient, d'_i is the processed detail coefficient, and IDWT is the inverse discrete wavelet transform. The reconstructed signal is a noise-reduced version of the original signal and is used as the input to the prediction model. The signal reconstruction unit also provides the following functions: 1. Efficient implementation: Improve computational efficiency using the fast inverse wavelet transform (IFWT) algorithm 2. Boundary processing: Provide a boundary processing method to minimize distortion at the boundaries of the reconstructed signal 3. Parallel computing: Parallelize the reconstruction calculation using multiple CPU cores or GPUs 4. Incremental Reconfiguration: Update the reconfiguration incrementally every time a new processing coefficient becomes available. 5. Selective Reconfiguration: Reconstruct the signal using only the coefficients of a specific level or frequency band. The noise reduction component also provides the following additional functions: 1. Noise Characteristic Analysis: Analyze the noise characteristics (distribution, spectrum, autocorrelation, etc.) of the input signal. 2. Noise Reduction Evaluation: Quantitatively evaluate the effect of noise reduction (e.g., signal-to-noise ratio, mean squared error, etc.). 3. Parameter Optimization: Optimize the noise reduction parameters using cross-validation or information criteria. 4. Adaptive Noise Reduction: Adapt the noise reduction parameters according to the temporal changes of the signal. 5. Multiscale Noise Reduction: Apply different reduction strategies for noises of different scales. In addition to these five main components, the system includes an integration module. This module integrates the information from each component and generates the final trading decision. The integration module includes an information integration unit, a decision generation unit, and an execution management unit. The information integration unit integrates the information (predictions, reliability, risk assessments, etc.) from each component. The decision generation unit generates trading decisions (buy, sell, hold, etc.) based on the integrated information. The execution management unit manages the execution of the trading decision and optimizes the execution strategy according to the market situation. The information integration unit integrates the information from each component. The information to be integrated includes the following: 1. Predictions from the ensemble model (probabilities of each class) 2. Reliability scores of the predictions 3. Importance (attention scores) at each time scale 4. Contribution degrees of each model 5. Market liquidity indicators 6. Risk assessment This information is passed to the decision-making generation unit. The information integration unit provides the following functions: 1. Information normalization: Normalize information on different scales to make it comparable. 2. Information weighting: Weight information based on the importance of each information source. 3. Information verification: Verify the consistency and reliability of information. 4. Information aggregation: Aggregate information from multiple information sources. 5. Information visualization: Visually represent the integrated information. The decision-making generation unit generates a trading decision based on the integrated information. The trading decision includes the direction of the trade (buy, sell, hold), the trading size, and the order type (market order, limit order, etc.). The direction of the trade is determined based on the predicted class and the confidence level of the prediction. For example, if the probability of class +1 exceeds a threshold (e.g., 0.6) and the confidence score is high, a buy decision is generated. Similarly, if the probability of class -1 exceeds the threshold and the confidence score is high, a sell decision is generated. Otherwise, a hold decision is generated. The trading size is determined based on the confidence level of the prediction, market liquidity, and risk assessment. A common approach is to calculate the trading size using a weighted sum of these factors: size = base_size * (w_1 * confidence + w_2 * liquidity + w_3 * (1 - risk)) Here, base_size is the basic trading size, confidence is the confidence level of the prediction (range from 0 to 1), liquidity is the market liquidity (range from 0 to 1), risk is the risk assessment (range from 0 to 1), and w_1, w_2, w_3 are the weights of each factor. The order type is determined based on market liquidity and volatility. For example, if liquidity is high and volatility is low, a limit order may be appropriate. Conversely, if liquidity is low and volatility is high, a market order may be appropriate. The decision-making generation unit also provides the following functions: 1. Risk management: Evaluate the risk of a position and apply risk limits 2. Portfolio optimization: Optimize trading decisions across multiple assets 3. Transaction cost analysis: Optimize trading decisions considering transaction costs (spread, commission, slippage, etc.) 4. Scenario analysis: Simulate the results of trading decisions under different market scenarios 5. Explanation of decisions: Provide information explaining the basis for trading decisions The execution management unit manages the execution of trading decisions. This component is responsible for order submission, order tracking, order modification or cancellation, and evaluation of execution quality. Order submission is carried out through the API of the exchange or trading platform. The execution management unit optimizes the order submission timing and implements strategies to minimize market impact. For example, strategies such as splitting a large order into multiple small orders or submitting orders at staggered times are used. Order tracking monitors the status of the submitted order (pending, partially filled, fully filled, canceled, etc.). The execution management unit modifies or cancels the order as necessary based on the order status. Evaluation of execution quality includes comparison of the actual execution price with the expected price, measurement of slippage, evaluation of market impact, etc. These metrics are used to improve and optimize trading strategies. The execution management unit also provides the following functions: 1. Optimal Execution Algorithm: Select the optimal execution algorithm according to market conditions 2. Execution Cost Minimization: Implement a strategy to minimize execution costs (such as spread, market impact, timing cost, etc.) 3. Execution Risk Management: Manage risks during execution (such as price fluctuations, liquidity depletion, etc.) 4. Execution Performance Analysis: Analyze execution performance and identify areas for improvement 5. Execution Report Generation: Generate a report recording the details of the execution Furthermore, this system includes the following additional modules: 1. Outlier Handling Module: Detect and handle outliers in market data. This module uses outlier detection algorithms (such as Z-score, interquartile range, isolation forest, etc.) to identify abnormal data points and process them appropriately (exclude, interpolate, robust transformation, etc.). 2. Reliability Evaluation Module: Evaluate the reliability of predictions and adjust the trading size based on the reliability. This module calculates a reliability score considering factors such as prediction probability, agreement between models, stability of predictions, etc., and determines the trading size based on it. 3. Liquidity Evaluation Module: Evaluate market liquidity and optimize the trading execution strategy based on liquidity. This module uses indicators such as bid-ask spread, market depth, trading volume, etc. to evaluate liquidity and determine the optimal order type and order size. 4. Seasonality Detection Module: Detect the seasonality and periodicity of market data and adjust predictions based on it. This module uses techniques such as Fourier transform, wavelet transform, autocorrelation analysis, etc. to identify seasonality such as intraday patterns, weekly patterns, monthly patterns, etc. 5. Structural Change Detection Module: Detects structural changes in the market and retrains the model based on them. This module uses change point detection algorithms (e.g., CUSUM, PELT, BICP, etc.) to identify changes in the market regime. 6. Correlation Analysis Module: Analyzes the correlations between multiple assets and enhances predictions based on them. This module uses metrics such as Pearson correlation, Spearman correlation, and mutual information to quantify the relationships between assets. 7. Macroeconomic Integration Module: Considers the macroeconomic factors of the market and adjusts predictions based on them. This module collects information such as economic indicators, central bank policies, and geopolitical events, and analyzes the impact they have on the market. 8. Transaction Analysis Module: Analyzes transaction histories and provides insights for optimizing transaction performance. This module calculates metrics such as win rate, profit-loss ratio, maximum drawdown, and Sharpe ratio to identify the strengths and weaknesses of trading strategies. The training process of this system consists of the following steps: 1. Extract features from historical market data: Use a multi-resolution time processing engine to extract features at multiple time scales. 2. Remove noise from features using a noise reduction component: Apply wavelet decomposition and thresholding to improve the signal-to-noise ratio. 3. Update source domain statistics: Use the source statistics calculation part of the domain adaptation module to calculate and save the statistical information of the training data. 4. Update the weights of the labels based on the class distribution: Use an adaptive label balancing framework to calculate the optimal weights for each class. 5. Train each model using weighted loss: Train each model of the ensemble model architecture using the calculated weights. The training process can be executed batch - wise or incrementally. In batch processing, the model is trained once using all the training data. In the incremental approach, the model is updated every time new data becomes available. The latter approach is suitable for continuously adapting to changes in market conditions. In the case of batch processing, the training is executed in the following steps: 1. Split the training data into mini - batches 2. For each mini - batch, perform the following operations: a. Feature extraction and noise reduction b. Domain adaptation c. Forward propagation (generation of predictions by the model) d. Calculation of weighted loss e. Backward propagation (calculation of gradients) f. Parameter update (using an optimization algorithm) 3. After processing all mini - batches, evaluate the model using validation data 4. Repeat steps 1 - 3 until the convergence condition is met In the case of the incremental approach, the training is executed in the following steps: 1. Train an initial model (similar to batch processing) 2. Every time new data becomes available, perform the following operations: a. Feature extraction and noise reduction b. Domain adaptation c. Generation and evaluation of predictions d. Update the model (using an online learning algorithm) 3. Periodically, perform a complete retraining of the model (e.g., when performance degrades, or after a certain period of time)
[0009] The following optimization techniques are also used in the training process: 1. Learning rate scheduling: Adjust the learning rate according to the progress of training (e.g., cyclic learning rate, cosine annealing, step decay, etc.) 2. Regularization: Techniques to prevent overfitting (e.g., L1 regularization, L2 regularization, dropout, early stopping, etc.) 3. Batch normalization: Reduce internal covariate shift and improve the stability of training 4. Gradient clipping: Limit the magnitude of the gradient to prevent the gradient explosion problem 5. Model ensemble: Train multiple models and combine their predictions The prediction process consists of the following steps: 1. Extract features from current market data: Use a multi-resolution time processing engine to extract features at multiple time scales. 2. Remove noise from the features: Use a noise reduction component to remove noise from the extracted features. 3. Update target domain statistics and adapt the features: Use a domain adaptation module to calculate the statistical information of recent market data and adapt the features. 4. Generate predictions using the ensemble: Use an ensemble model architecture to weight and aggregate the predictions of each model to generate the final prediction. 5. Generate trading decisions based on the predictions: Use an integration module to generate trading decisions considering factors such as predictions, confidence levels, and market conditions. 6. Execute the trading decisions: Use the execution management unit of the integration module to execute the trading decisions and evaluate the execution quality. The prediction process can be executed in real-time or periodically. In the real-time approach, predictions are generated whenever new market data becomes available. In the periodic approach, predictions are generated at fixed intervals (e.g., every 1 second, every 5 seconds, etc.). The appropriate approach depends on the trading strategy and the characteristics of the market. In the case of the real-time approach, the prediction is executed in the following steps: 1. Continuously monitor the market data stream 2. Whenever a new data point arrives, perform the following operations: a. Feature extraction and noise reduction b. Domain adaptation c. Generation of predictions d. Generation and execution of trading decisions 3. Continuously monitor the market situation and the performance of the predictions, and adjust the parameters as needed In the case of the periodic approach, the prediction is executed in the following steps: 1. Collect data at fixed intervals 2. At the end of each interval, perform the following operations: a. Feature extraction and noise reduction b. Domain adaptation c. Generation of predictions d. Generation and execution of trading decisions 3. Periodically evaluate the prediction performance and adjust the parameters as needed In the prediction process, the following optimization techniques are also used: 1. Caching: Cache frequently used calculation results to reduce the computational overhead 2. Parallel processing: Use multiple CPU cores or GPUs to parallelize the calculations 3. Batch processing: Process multiple data points as a batch to improve the throughput 4. Model quantization: Convert the model parameters into a low-precision representation to reduce memory usage and computational cost. 5. Model distillation: Transfer the knowledge of a large model to a small model to improve computational efficiency. The system is periodically updated based on recent data, labels, and model performance. This update process includes the following steps: 1. Update of the domain adaptation module: Calculate the statistical information of recent market data and update the target domain statistics. 2. Update of the label balancer: Update the weights of each class based on the recent class distribution and model performance. 3. Update of the ensemble weights: Update the weights of the models based on the recent performance of each model. 4. Retraining of the model: Retrain the model using the latest data if necessary. The update frequency is set based on market dynamics and computational resource constraints. A common approach is to update different components at different frequencies. For example, the domain adaptation module and the ensemble weights are updated frequently (e.g., every few minutes), while the retraining of the model is performed less frequently (e.g., every few hours or days). The update process is executed according to the following schedule: 1. Real-time update (from milliseconds to seconds): - Feature extraction and noise reduction - Generation of predictions - Generation and execution of trading decisions 2. Short-term update (in minutes): - Update of target domain statistics - Update of ensemble weights - Update of confidence scores - Update of liquidity evaluation 3. Medium-term updates (time unit): - Update of label weights - Incremental update of the model - Update of outlier detection parameters - Update of seasonal patterns 4. Long-term updates (daily or weekly): - Complete retraining of the model - Optimization of hyperparameters - Update of feature selection - Performance evaluation of the entire system The following techniques are also used in the update process: 1. Online learning: Incrementally update the model every time new data becomes available 2. Transfer learning: Transfer the knowledge of the previous model to the new model 3. Active learning: Select the most informative data points for labeling and training 4. Continuous learning: Continuously learn new patterns while maintaining the performance of the model 5. Prevention of catastrophic forgetting: Use techniques to retain previous knowledge when learning new data
Advantages of the Invention
[0010] According to the present invention, the following effects can be obtained: First, effective handling of label imbalance becomes possible. The adaptive label balancing framework improves the prediction accuracy of the minority class (profitable trading opportunities) by dynamically adjusting weights based on the distribution and performance of each class. This reduces the problem that standard machine learning algorithms tend to be biased towards the majority class, and achieves a more balanced prediction performance. Specifically, in the conventional static weighting approach, it is not possible to adapt to changes in market conditions, and the effect may decrease over time. In the adaptive approach of the present invention, it is possible to continuously adapt to the dynamic nature of the market and always maintain optimal weights. As a result, the detection rate of profit opportunities (labels +1 and -1) is improved, and the proportion of false positives (predicting a profit opportunity when it is not actually a profit opportunity) is reduced. Consequently, the profitability of transactions is improved, and unnecessary transaction costs are reduced. Second, the ability to adapt to market changes is improved. The domain adaptation module enables the model to maintain robustness against changes in the market regime by continuously monitoring and adapting to the distribution shift of market data. This reduces the "model drift" phenomenon and achieves consistent performance over a long period. In the conventional approach, it is vulnerable to changes in the market regime (for example, the transition from a low volatility market to a high volatility market), and frequent retraining of the model is required. In the domain adaptation approach of the present invention, by continuously adjusting the distribution of features, it can function effectively even in different market regimes. As a result, the prediction accuracy is maintained even during sudden market changes (for example, the announcement of economic indicators, changes in central bank policies, geopolitical events, etc.). Consequently, the lifespan of the model is extended, the frequency of retraining is reduced, and the operation cost is reduced. Thirdly, the improvement of signal extraction is realized. The noise reduction component uses wavelet decomposition and adaptive thresholding to reduce noise while retaining the underlying signal structure. As a result, the impact of market microstructure noise is reduced, and a more reliable prediction signal is extracted. In the conventional approach, basic noise reduction techniques such as simple moving averages and low-pass filters were used, but these techniques lost time locality and could smooth out important market events. The wavelet-based approach of the present invention can analyze signals in both the time and frequency domains and effectively reduce noise while retaining important market events. This improves the prediction accuracy, especially in high-noise environments (e.g., at the start or end of the market, during periods of low liquidity, etc.). As a result, false trading signals are reduced, and the profitability of trading is improved. Fourthly, the improvement of prediction accuracy is realized. The multi-resolution time processing engine enables understanding of complex market dynamics that cannot be captured on a single time scale by processing market data on multiple time scales. Also, the ensemble model architecture realizes consistently higher performance by combining the strengths of multiple models. In the conventional approach, due to dependence on a single time scale and a single model architecture, the adaptability in different market situations was limited. The multi-scale, multi-model approach of the present invention enables a more comprehensive understanding of the market by integrating information from different time scales and combining the predictions of different models. For example, it becomes possible to make predictions considering both short-term price fluctuations (micro market structure) and long-term trends (macro market dynamics). Also, since each model (MLP, LSTM, Mamba) can capture different types of patterns, combining them realizes a more robust prediction. As a result, the prediction accuracy is improved, and the profitability of trading is improved. Fifthly, the computational efficiency is improved. This system is designed considering the computational efficiency of each component and can operate within the strict time constraints of a high-frequency trading environment. In particular, computationally intensive operations such as wavelet decomposition and feature adaptation are optimized using efficient algorithms and data structures. In the conventional approach, using complex models to achieve high prediction accuracy would result in high computational costs and the possibility of not being able to operate within the time constraints of high-frequency trading. With the approach of the present invention, by balancing computational efficiency and model complexity, it becomes possible to generate predictions in milliseconds while maintaining high prediction accuracy. For example, the fast wavelet transform (FWT) algorithm is used for wavelet decomposition, and caching technology is used for feature adaptation to minimize computational overhead. Also, techniques such as parallel processing, batch processing, and model quantization are used to improve computational efficiency. As a result, low-latency trade execution becomes possible, enabling market opportunities to be captured without being missed. Sixthly, the interpretability of the system is improved. This system provides not only the prediction results but also the intermediate outputs of each component (for example, the importance of each time scale, the contribution of each model, the reliability of the prediction, etc.). Thereby, traders and analysts can understand the decision-making process of the system more deeply and adjust it as needed. In the conventional approach, especially when using deep learning models, it was difficult to understand the basis for the prediction (the "black box problem"). With the approach of the present invention, the transparency of the prediction process is enhanced by visualizing the operation and contribution of each component. For example, it provides information such as which time scale is most affecting the current prediction, which model is providing the most reliable prediction, and to what extent the reliability of the prediction is. Thereby, traders can evaluate the quality of the prediction and add human judgment as needed. As a result, the reliability of the system is improved, enabling more effective cooperation between humans and AI. Seventhly, the improvement of risk management is realized. The reliability evaluation module and the liquidity evaluation module enable trading decisions considering the uncertainty of prediction and the liquidity of the market. This prevents excessive risk-taking and realizes more stable trading performance. In the conventional approach, trading decisions are often made based only on the probability of prediction, and the uncertainty of prediction and the liquidity situation of the market are not sufficiently considered. In the approach of the present invention, the reliability of prediction (for example, the degree of agreement between models, the stability of prediction, etc.) and the liquidity of the market (for example, bid-ask spread, market depth, etc.) are explicitly evaluated, and the trading size and order type are adjusted based on this. For example, in the case of a prediction with low reliability or a market situation with low liquidity, the trading size can be reduced or the trading can be postponed. This reduces the risk of large losses and realizes a more stable profit profile. As a result, the Sharpe ratio (risk-adjusted return) is improved and the maximum drawdown (maximum loss) is reduced. Eighthly, it becomes possible to utilize the seasonality and periodicity of the market. The seasonality detection module detects the seasonality and periodicity of market data and adjusts the prediction based on this. This enables trading strategies that utilize repetitive market behaviors such as intraday patterns, weekly patterns, and monthly patterns. In the conventional approach, seasonality and periodicity are often not explicitly considered, and the potential predictive power obtained from these patterns is not utilized. In the approach of the present invention, techniques such as Fourier transform, wavelet transform, and autocorrelation analysis are used to identify seasonality at different time scales and integrate it into the prediction model. For example, patterns specific to a particular time period (for example, the start or end of the market) or a particular day of the week (for example, Monday or Friday) can be learned, and the prediction can be adjusted based on this. As a result, additional predictive power is obtained from these repetitive patterns, and the profitability of trading is improved. Ninth, the ability to respond to structural changes in the market is improved. The structural change detection module detects structural changes in the market and retrains the model based on them. This improves the adaptability to changes in the market regime (e.g., the transition from a low-volatility market to a high-volatility market). In the conventional approach, structural changes are not detected, and the model continues to make predictions based on old patterns, which may lead to a decline in performance. In the approach of the present invention, a change point detection algorithm (e.g., CUSUM, PELT, BICP, etc.) is used to automatically identify changes in the market regime and update the model accordingly. For example, when structural changes such as a sharp increase in volatility, a change in the correlation structure, or a change in trading patterns are detected, the model can be retrained or switched to a different model. As a result, consistent performance is maintained across different market regimes, and the lifespan of the model is extended. Tenth, the utilization of correlations among multiple assets becomes possible. The correlation analysis module analyzes the correlations among multiple assets and strengthens the predictions based on them. This enables predictions that utilize information from related assets. In the conventional approach, each asset is often analyzed independently, and the potential predictive power obtained from the interrelationships among assets is not utilized. In the approach of the present invention, metrics such as Pearson correlation, Spearman correlation, and mutual information are used to quantify the relationships among assets and integrate them into the prediction model. For example, price movements of highly correlated assets can be used as additional features for prediction, or common patterns across multiple assets can be identified. As a result, a more comprehensive understanding of the market becomes possible, and the prediction accuracy is improved. Eleventh, macroeconomic factors can be integrated. The macroeconomic integration module takes into account the macroeconomic factors of the market and adjusts the prediction based on them. This enables predictions that consider external factors such as economic indicators, central bank policies, and geopolitical events. In the conventional approach, macroeconomic factors are often not explicitly considered, and the impact of these factors on the market is not reflected in the prediction model. In the approach of the present invention, information is collected from external data sources such as economic calendars, news feeds, and central bank statements and integrated into the prediction model. For example, the reliability of the prediction can be adjusted before and after the release of important economic indicators, or the prediction can be adjusted by analyzing the market reaction after a central bank policy change. As a result, more comprehensive predictions that consider the impact of external factors are possible, and the prediction accuracy is improved. Twelfth, continuous optimization of trading performance becomes possible. The trade analysis module analyzes the trading history and provides insights for optimizing trading performance. This enables the identification of the strengths and weaknesses of trading strategies and continuous improvement. In the conventional approach, the analysis of trading performance is limited, and there is often a lack of a systematic feedback loop for performance improvement. In the approach of the present invention, indicators such as win rate, profit-loss ratio, maximum drawdown, and Sharpe ratio are calculated, and the performance in different market conditions, time periods, asset classes, etc. is analyzed in detail. For example, weaknesses in specific market conditions (such as high volatility, low liquidity, etc.) can be identified, and the model can be adjusted to address them. As a result, the trading strategy is continuously improved, and the long-term performance is enhanced. By combining these effects, the present invention significantly improves the prediction accuracy, adaptability, robustness, and efficiency in a high-frequency trading environment. As a result, it is expected that the profitability of trading will be improved, the risk will be reduced, and the market efficiency will be enhanced.
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described in detail.
[0012] The High-Frequency Trading Adaptive Intelligence System (HFT-AIS) of the present invention provides a comprehensive solution to address the interrelated issues of label imbalance, domain shift, and high noise levels. The system consists of six main components: a multi-resolution time processing engine, an adaptive label balancing framework, a domain adaptation module, an ensemble model architecture, a noise reduction component, and an integration module.
[0013] The multi-resolution time processing engine processes market data simultaneously at multiple time scales. This engine includes a time scale setting unit, a feature extractor generator, and a feature integration unit.
[0014] The time scale setting unit sets multiple time scales for analysis. By default, four time scales of 5 seconds, 15 seconds, 30 seconds, and 60 seconds are used, but they can be customized according to market characteristics and trading strategies. For example, in a higher-frequency trading strategy, short time scales such as 1 second, 3 seconds, 5 seconds, 10 seconds, etc. may be appropriate. On the other hand, in a longer-term trading strategy, long time scales such as 1 minute, 5 minutes, 15 minutes, 30 minutes, etc. may be appropriate. The time scale setting unit can also dynamically adjust these time scales. For example, during periods of high market volatility, shorter time scales can be emphasized, and during periods of low volatility, longer time scales can be emphasized.
[0015] The time scale setting unit sets the following parameters: 1. Set of time scales: A list of time scales used for analysis (e.g., [5, 15, 30, 60] seconds) 2. Weights of each time scale: The initial importance of each time scale (e.g., [0.25, 0.25, 0.25, 0.25]) 3. Adjustment parameter for time scales: A parameter for adjusting the time scales according to market conditions 4. Minimum and Maximum Values of the Time Scale: The Effective Range of the Time Scale 5. Adjustment Frequency of the Time Scale: The Frequency at Which the Time Scale is Re-evaluated
[0016] For the dynamic adjustment of the time scale, the following formula can be used: w_i = softmax(β * f_i(M)) Here, w_i is the weight of time scale i, β is a parameter controlling the sensitivity of the adjustment, f_i(M) is a function measuring the fitness of time scale i in market situation M, and softmax is the softmax function. The fitness function f_i(M) is defined based on factors such as market volatility, trading volume, and strength of the trend.
[0017] The feature extractor generation unit generates a feature extractor corresponding to each time scale. Each feature extractor extracts relevant features from the market data at the corresponding time scale. The features extracted include the following: 1. Price Features: Average price, standard deviation, skewness, kurtosis, maximum value, minimum value, range, etc. These features capture the distribution and volatility of prices. For example, the standard deviation of price measures the volatility of the market, and the skewness of price measures the asymmetry of the price distribution. 2. Trading Volume Features: Total trading volume, average, standard deviation, ratio to the past period, etc. These features capture the activity level and liquidity of the market. For example, an increase in trading volume may indicate an increase in market interest or the release of important news. 3. Order Book Features: Bid-ask spread, order book imbalance, order book depth, order arrival rate, etc. These features capture the balance between market supply and demand. For example, the order book imbalance (the difference in the quantity of buy orders and sell orders) may indicate the future direction of the price. 4. Technical indicators: Moving average, Relative Strength Index (RSI), Bollinger Bands, MACD (Moving Average Convergence Divergence), etc. These indicators capture price trends, momentum, overbought / oversold conditions, etc. 5. Market microstructure characteristics: Effective spread, price impact, liquidity indicators, volatility indicators, etc. These characteristics capture the microstructure of the market and transaction costs. For example, the effective spread measures the actual cost of a trade, and the price impact measures the effect of a large order on the price.
[0018] Each feature extractor is implemented using a Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM) network, or attention mechanism. These architectures can effectively extract temporal patterns from time series data. CNNs are suitable for detecting local temporal patterns (e.g., price spikes, sudden increases in trading volume, etc.). LSTMs are suitable for capturing long-term dependencies (e.g., the relationship between past price patterns and future price movements). The attention mechanism can dynamically evaluate the importance of different time points and focus on important time points. By combining these architectures, complex patterns in time series data can be effectively captured.
[0019] Specifically, the CNN-based feature extractor has the following structure: 1. Input layer: Time series data of a length corresponding to the time scale 2. 1D convolutional layer 1: 64 filters, kernel size 3, stride 1, padding "same" 3. Batch normalization layer 4. Activation layer: LeakyReLU (α = 0.2) 5. 1D convolutional layer 2: 128 filters, kernel size 3, stride 1, padding "same" 6. Batch normalization layer 7. Activation layer: LeakyReLU(α = 0.2) 8. 1D Convolutional Layer 3: 256 filters, kernel size 3, stride 1, padding "same" 9. Batch Normalization Layer 10. Activation layer: LeakyReLU(α = 0.2) 11. Global Average Pooling Layer 12. Fully Connected Layer: 128 units 13. Activation layer: LeakyReLU(α = 0.2) 14. Dropout Layer: Dropout rate 0.2 15. Fully Connected Layer: 64 units 16. Activation layer: LeakyReLU(α = 0.2)
[0020] Specifically, the CNN-based feature extractor has the following structure: 1. Input Layer: Time series data of length corresponding to the time scale 2. 1D Convolutional Layer 1: 64 filters, kernel size 3, stride 1, padding "same" 3. Batch Normalization Layer 4. Activation layer: LeakyReLU(α = 0.2) 5. 1D Convolutional Layer 2: 128 filters, kernel size 3, stride 1, padding "same" 6. Batch Normalization Layer 7. Activation layer: LeakyReLU(α = 0.2) 8. 1D Convolutional Layer 3: 256 filters, kernel size 3, stride 1, padding "same" 9. Batch Normalization Layer 10. Activation layer: LeakyReLU(α = 0.2) 11. Global Average Pooling Layer 12. Fully Connected Layer: 128 units 13. Activation layer: LeakyReLU(α = 0.2) 14. Dropout Layer: Dropout rate 0.2 15. Fully Connected Layer: 64 units 16. Activation Layer: LeakyReLU(α = 0.2)
[0021] The attention mechanism-based feature extractor has the following structure: 1. Input Layer: Time series data with a length corresponding to the time scale 2. Position Encoding Layer 3. Self-Attention Layer 1: 8 heads, 64 dimensions 4. Residual Connection 5. Layer Normalization 6. Feed-Forward Layer: Internal dimension 256 7. Residual Connection 8. Layer Normalization 9. Self-Attention Layer 2: 8 heads, 64 dimensions 10. Residual Connection 11. Layer Normalization 12. Feed-Forward Layer: Internal dimension 256 13. Residual Connection 14. Layer Normalization 15. Global Average Pooling Layer 16. Fully Connected Layer: 32 units 17. Activation Layer: ReLU
[0022] The feature integration part hierarchically integrates the features extracted by each feature extractor. The integration process is implemented using a cross-scale attention mechanism. This mechanism dynamically evaluates the importance of the features at each time scale and weights and integrates the features based on that. Specifically, the integrated feature is calculated using the following formula: F_integrated = Σ(α_i * F_i) Here, F_integrated is the integrated feature, F_i is the feature at time scale i, and α_i is the importance (attention score) at time scale i. The attention score is calculated based on the current market situation and can capture the importance of different time scales in different market regimes. For example, during periods of high market volatility, the importance of short time scales may increase. Conversely, during periods when the market is following a trend, the importance of long time scales may increase.
[0023] For calculating the attention score, an additive attention mechanism or a multiplicative attention mechanism can be used. In the additive attention mechanism, the attention score is calculated using the following formula: e_i = v^T * tanh(W_1 * F_i + W_2 * q) α_i = softmax(e_i) Here, F_i is the feature at time scale i, q is the query vector (representing the current market situation), W_1, W_2, v^T are learnable parameters, tanh is the hyperbolic tangent function, and softmax is the softmax function.
[0024] In the multiplicative attention mechanism, the attention score is calculated using the following formula: e_i = (F_i * W_1) * (q * W_2)^T / sqrt(d) α_i = softmax(e_i) Here, d is the dimensionality of the feature, and sqrt is the square root function.
[0025] Furthermore, the feature integration part can also use a multi-head attention mechanism to capture the interaction between features of different time scales. In the multi-head attention mechanism, multiple attention mechanisms (heads) are applied in parallel, and their outputs are concatenated to generate the final feature representation. This enables capturing different types of temporal relationships simultaneously.
[0026] In addition to the attention mechanism, the feature integration unit also uses conditional feature modulation techniques such as the Gate Linear Unit (GLU) and Fast Weight Adaptation (FiLM). These techniques enable the dynamic adjustment of the importance of features according to market conditions. For example, GLU is defined by the following formula: GLU(F_i, q) = σ(W_1 * q) H (W_2 * F_i) Here, σ is the sigmoid function, H is the Hadamard product (element-wise product), and W_1 and W_2 are learnable parameters.
[0027] FiLM is defined by the following formula: FiLM(F_i, q) = γ(q) H F_i + β(q) Here, γ(q) and β(q) are scaling parameters and shift parameters calculated from the query vector q.
[0028] Finally, the feature integration unit normalizes the integrated features and applies dimensionality reduction techniques (such as principal component analysis (PCA), t-SNE, UMAP, etc.) to reduce the dimensionality of the features and improve computational efficiency. This enables an efficient mapping from a high-dimensional feature space to a low-dimensional latent space.
[0029] The adaptive label balancing framework dynamically adjusts the importance of each class (-1, 0, +1). Different from the conventional static weighting approach, this framework takes into account both the current distribution of each class and the recent performance of the model for each class. This enables adaptation to changes in market conditions and improvement in the prediction accuracy of minority classes (profitable trading opportunities).
[0030] The adaptive label balancing framework includes a class distribution recording unit, a weight holding unit, a performance tracking unit, and a weight calculation unit. The class distribution recording unit records changes in the class distribution over time. Specifically, it calculates the class distribution in each time window (e.g., 1 hour, 1 day, 1 week, etc.) and saves it as historical data. This enables tracking of the temporal variations in the class distribution and identification of seasonal and market regime changes. For example, the class distribution may change before and after a specific time period (e.g., at the start or end of the market) or a specific market event (e.g., release of economic indicators, central bank policy announcements, etc.). By tracking such variations, the label balancing strategy can be adjusted more effectively.
[0031] The adaptive label balancing framework includes a class distribution recording unit, a weight holding unit, a performance tracking unit, and a weight calculation unit. The class distribution recording unit records changes in the class distribution over time. Specifically, it calculates the class distribution in each time window (e.g., 1 hour, 1 day, 1 week, etc.) and saves it as historical data. This enables tracking of the temporal variations in the class distribution and identification of seasonal and market regime changes. For example, the class distribution may change before and after a specific time period (e.g., at the start or end of the market) or a specific market event (e.g., release of economic indicators, central bank policy announcements, etc.). By tracking such variations, the label balancing strategy can be adjusted more effectively.
[0032] This information is stored in a time - series database and used for subsequent analysis and weight calculation. The time - series database supports high - speed data insertion and efficient queries, and can efficiently manage large amounts of historical data.
[0033] This information is stored in a time - series database and used for subsequent analysis and weight calculation. The time - series database supports high - speed data insertion and efficient queries, and can efficiently manage large amounts of historical data.
[0034] These pieces of information are stored in an in-memory data structure (such as a hash map, priority queue, etc.) and can be accessed and updated quickly. To prevent sudden fluctuations in the weights, the weight holding unit also applies smoothing techniques (such as exponential moving average, Kalman filter, etc.).
[0035] The performance tracking unit tracks the performance metrics of the model for each class. The metrics to be tracked include accuracy, recall, F1-score, confusion matrix, area under the ROC curve (AUC), etc. These metrics are calculated at the end of each evaluation period (such as each training epoch, each trading day, etc.) and saved as historical data. The performance tracking unit tracks the following information: 1. Basic classification metrics such as accuracy, recall, F1-score, AUC for each class 2. Confusion matrix for each class (number of true positives, false positives, true negatives, false negatives) 3. Predicted probability distribution for each class (such as a histogram of predicted probabilities) 4. Misclassification cost for each class (such as the cost of false positives and false negatives) 5. Temporal variation of the performance for each class
[0036] These pieces of information are stored in a time-series database and used for subsequent analysis and weight calculation. The performance tracking unit can also use an online learning algorithm to update the performance metrics in real time.
[0037] The weight calculation unit calculates the optimal weights based on the distribution factor and the performance factor. Specifically, the weights for each class are calculated using the following formula: w_c = α * (N / N_c) + (1 - α) * (1 / P_c) Here, \(w_c\) is the weight of class \(c\), \(N\) is the total number of samples, \(N_c\) is the number of samples in class \(c\), \(P_c\) is the recent performance of the model for class \(c\) (normalized to the range from 0 to 1), and \(\alpha\) is a parameter (in the range from 0 to 1) that controls the balance between the distribution factor and the performance factor.
[0038] The first term \((N / N_c)\) of this formula is the weight based on the inverse frequency of the class, which assigns a high weight to the minority class. The second term \((1 / P_c)\) is the weight based on the reciprocal of the performance, which assigns a high weight to the class with low performance. The \(\alpha\) parameter controls the balance between these two factors. The closer the value of \(\alpha\) is to 1, the greater the influence of the distribution factor, and the closer the value of \(\alpha\) is to 0, the greater the influence of the performance factor.
[0039] The value of \(\alpha\) can be dynamically adjusted according to the market situation and the performance of the model. For example, when the model consistently shows low performance for a specific class, the value of \(\alpha\) can be decreased to increase the influence of the performance factor. Conversely, when the class distribution is extremely imbalanced, the value of \(\alpha\) can be increased to increase the influence of the distribution factor.
[0040] The following formula can be used for the dynamic adjustment of \(\alpha\): \(\alpha=\text{sigmoid}(\beta\times(I_{\text{distribution}} - I_{\text{performance}}))\) Here, \(\text{sigmoid}(x)=\frac{1}{1 + \exp(-x)}\) is the sigmoid function, \(\beta\) is a parameter that controls the sensitivity of the adjustment, \(I_{\text{distribution}}\) is an indicator that measures the degree of imbalance of the class distribution (e.g., Gini coefficient, entropy, etc.), and \(I_{\text{performance}}\) is an indicator that measures the degree of imbalance of the performance (e.g., standard deviation of performance between classes, etc.).
[0041] Furthermore, the weight calculation unit calculates the weight using the class distribution and the performance index within an exponentially decaying time window: N_c(t) = Σ(e^(-λ * (t - s)) * n_c(s)) P_c(t) = Σ(e^(-λ * (t - s)) * p_c(s)) / Σ(e^(-λ * (t - s))) Here, N_c(t) is the weighted number of samples of class c at time t, n_c(s) is the number of samples of class c at time s, P_c(t) is the weighted performance of class c at time t, p_c(s) is the performance of class c at time s, and λ is the decay rate. Due to this exponential decay, the weights are continuously updated based on recent data and performance, enabling adaptation to changes in market conditions.
[0042] Furthermore, the weight calculation unit calculates the weights using the class distribution and performance metrics within an exponentially decaying time window: N_c(t) = Σ(e^(-λ * (t - s)) * n_c(s)) P_c(t) = Σ(e^(-λ * (t - s)) * p_c(s)) / Σ(e^(-λ * (t - s))) Here, N_c(t) is the weighted number of samples of class c at time t, n_c(s) is the number of samples of class c at time s, P_c(t) is the weighted performance of class c at time t, p_c(s) is the performance of class c at time s, and λ is the decay rate. Due to this exponential decay, the weights are continuously updated based on recent data and performance, enabling adaptation to changes in market conditions.
[0043] For the dynamic adjustment of λ, the following formula can be used: λ = λ_base + γ * V Here, λ_base is the base decay rate, γ is a parameter controlling the sensitivity of adjustment, and V is an indicator measuring the volatility of the market (e.g., standard deviation of price, mean absolute deviation, etc.).
[0044] The weight calculation unit normalizes the calculated weights to ensure numerical stability and maintain the ratio of weights between different classes. The following formula can be used for normalization: w_c_normalized = w_c / Σ(w_c') Here, w_c_normalized is the normalized weight of class c, and Σ(w_c') is the sum of the weights of all classes.
[0045] Furthermore, the weight calculation unit smooths the weights using the following formula to prevent sudden fluctuations in the weights: w_c_smoothed = (1 - β) * w_c_previous + β * w_c_current Here, w_c_smoothed is the smoothed weight of class c, w_c_previous is the weight of class c at the previous time point, w_c_current is the currently calculated weight of class c, and β is the smoothing parameter (in the range from 0 to 1).
[0046] Finally, the weight calculation unit integrates the calculated weights into the loss function. By applying weights to common classification loss functions (such as cross-entropy loss, focal loss, etc.), the importance of minority classes can be enhanced. For example, weighted cross-entropy loss is defined by the following formula: L = -Σ(w_c * y_c * log(p_c)) Here, L is the loss, y_c is the true label (0 or 1) of class c, and p_c is the predicted probability of class c.
[0047] The domain adaptation module continuously monitors the statistical characteristics of the training data (source domain) and the recent market data (target domain). Using these statistics, it transforms the input features to mitigate the impact of domain shift. This approach enables the model to remain effective even as market conditions evolve.
[0048] The domain adaptation module includes a source statistics calculation unit, a target statistics calculation unit, and a feature adaptation unit. The source statistics calculation unit calculates and stores the statistical information of the training data. The calculated statistical information includes the mean, standard deviation, minimum value, maximum value, quartiles, skewness, kurtosis, etc. of each feature. These statistics can be calculated from the entire training dataset or for each time window. In the latter approach, the temporal variations of the training data can be captured. The source statistics calculation unit calculates the following statistical information: 1. First-order statistics: mean, median, mode, minimum value, maximum value, range, quartiles, etc. 2. Second-order statistics: variance, standard deviation, mean absolute deviation, interquartile range, etc. 3. Higher-order statistics: skewness, kurtosis, entropy, etc. 4. Correlation statistics: correlation matrix, covariance matrix, principal components, etc. 5. Temporal statistics: autocorrelation, partial autocorrelation, spectral density, etc.
[0049] These statistics are stored in an efficient data structure (e.g., hash map, B-tree, etc.) for fast access. The source statistics calculation unit uses robust estimation methods (e.g., median absolute deviation, winsorized mean, etc.) for calculating the statistics to minimize the influence of outliers.
[0050] The target statistical calculation unit calculates and stores the statistical information of recent market data. The statistical information to be calculated is the same as that of the source statistical calculation unit. However, these statistics are calculated from the latest market data (e.g., in the past 1 hour, in the past 1 day, etc.). The calculation frequency of the statistics is set according to the dynamics of the market and the constraints of the calculation resources. For example, in a high-frequency trading environment, it may be appropriate to update the statistics every few seconds or minutes. On the other hand, for a more long-term trading strategy, it may be appropriate to update the statistics every few hours or days.
[0051] The target statistical calculation unit provides the following functions: 1. Real-time statistical calculation: Update the statistics incrementally every time a new data point arrives 2. Time-window statistical calculation: Calculate the statistics based on the data within the specified time window 3. Weighted statistical calculation: Calculate the statistics by assigning higher weights to the recent data points 4. Outlier detection and handling: Detect outliers before statistical calculation and handle them appropriately 5. Confidence interval calculation of statistics: Calculate the confidence interval to quantify the uncertainty of the statistics
[0052] The target statistical calculation unit also provides a function to track the temporal variation of the calculated statistics and detect abrupt changes in the statistics (which may indicate a change in the market regime).
[0053] The feature adaptation unit adapts the features using the statistical information of the source domain and the target domain. The basic approaches are feature normalization and inverse normalization: 1. Normalize the features using the mean and standard deviation of the source domain: x_norm = (x - μ_source) / σ_source 2. Inverse-normalize the normalized features using the mean and standard deviation of the target domain: x_adapted = x_norm * σ_target + μ_target Here, x is the original feature, x_norm is the normalized feature, x_adapted is the adapted feature, μ_source and σ_source are the mean and standard deviation of the source domain, and μ_target and σ_target are the mean and standard deviation of the target domain.
[0054] This approach aligns the first moment (mean) and second moment (variance) of the features, but does not consider the differences in higher-order moments. As a more advanced approach, the feature adaptation part also implements the following methods: 1. Maximum Mean Discrepancy (MMD) minimization: Minimize the distance between the feature distributions of the source domain and the target domain. MMD measures the distance between two distributions in a Reproducing Kernel Hilbert Space (RKHS). The minimization of MMD is formulated as the following optimization problem: min_θ MMD^2(X_source, X_target_θ) Here, X_source is the feature of the source domain, X_target_θ is the feature of the target domain transformed by parameter θ, and MMD^2 is the squared MMD distance. 2. Correlation Alignment (CORAL): Align the covariance matrices of the features of the source domain and the target domain. CORAL applies the following transformation: X_adapted = X * W Here, W = C_source^(-1 / 2) * C_target^(1 / 2), C_source is the covariance matrix of the source domain, and C_target is the covariance matrix of the target domain. 3. Adversarial Domain Adaptation: Use adversarial training to learn domain-invariant feature representations. In this method, a feature extractor and two classifiers (a label classifier and a domain classifier) are trained. The feature extractor is trained to maximize the performance of the label classifier and minimize the performance of the domain classifier. As a result, feature representations that are useful for label prediction but independent of the domain are learned.
[0055] These methods can handle more complex domain shifts (e.g., non-linear transformations, changes in conditional distributions, etc.). The appropriate method is selected according to the characteristics of the market data and the nature of the domain shift.
[0056] The feature adaptation part also provides the following functions: 1. Automatic adjustment of adaptation parameters: Adjust the adaptation parameters according to the magnitude of the domain shift 2. Stepwise adaptation: Apply the adaptation step by step to avoid sudden transformations 3. Conditional adaptation: Apply the adaptation only under specific conditions (e.g., when the domain shift exceeds a threshold) 4. Feature-selective adaptation: Apply the adaptation only to features that are susceptible to the influence of the domain shift 5. Multi-source adaptation: Integrate information from multiple source domains to perform adaptation
[0057] The feature adaptation part also implements a cache mechanism and parallel computing to improve the efficiency of the adaptation process. The cache mechanism stores frequently used statistics and transformation parameters in memory to reduce the computational overhead. Parallel computing uses multiple CPU cores or GPUs to parallelize the adaptation calculations and improve the processing speed.
[0058] The ensemble model architecture combines multiple prediction models. This architecture includes a model holding part, a weight holding part, a performance recording part, a prediction aggregation part, and a weight update part. The model holding part holds multiple models with different architectures and learning algorithms. By default, the following three models are used: 1. Multilayer Perceptron (MLP): A feedforward neural network with multiple hidden layers. Each layer is fully connected and uses a non-linear activation function (e.g., ReLU, LeakyReLU, SELU). The MLP can capture complex non-linear relationships between features. The structure of the MLP is as follows: - Input layer: The number of neurons corresponding to the dimension of the features - Hidden layer 1: 256 neurons, activation function is LeakyReLU - Hidden layer 2: 128 neurons, activation function is LeakyReLU - Hidden layer 3: 64 neurons, activation function is LeakyReLU - Output layer: The number of neurons corresponding to the number of classes, activation function is softmax The MLP also includes a dropout layer (dropout rate 0.2) and a batch normalization layer. These layers prevent overfitting and improve the stability of training. 2. Long Short-Term Memory (LSTM) network: A type of recurrent neural network that can learn long-term dependencies. The LSTM uses memory cells with input gates, forget gates, and output gates to perform selective retention and update of information. The LSTM is suitable for capturing the temporal patterns of time series data. The structure of the LSTM is as follows: - Input layer: The input size corresponding to the dimension of the features - LSTM layer 1: 128 units, bidirectional - Dropout layer: Dropout rate 0.3 - LSTM layer 2: 64 units, bidirectional - Dropout layer: Dropout rate 0.3 - Fully connected layer: 32 neurons, ReLU activation function - Output layer: Number of neurons corresponding to the number of classes, softmax activation function By using bidirectional LSTM, predictions considering both past and future contexts become possible. 3. Mamba model: A state-of-the-art architecture based on the structured state space sequence model (SSM). Mamba dynamically adjusts the parameters of the state space model according to the input using a selection mechanism. This approach is similar to the attention mechanism of the Transformer but enables more efficient inference. Mamba can capture both long-range dependencies and local patterns. The structure of Mamba is as follows: - Input layer: Input size corresponding to the dimension of the features - Embedding layer: Embedding of 128 dimensions - Mamba block 1: State size 64, expansion factor 4 - Mamba block 2: State size 64, expansion factor 4 - Mamba block 3: State size 64, expansion factor 4 - Mamba block 4: State size 64, expansion factor 4 - Fully connected layer: 32 neurons, ReLU activation function - Output layer: Number of neurons corresponding to the number of classes, softmax activation function Each Mamba block is composed of layer normalization, an SSM layer, a feed-forward layer, and a residual connection.
[0059] In addition to these models, the model repository can also include conventional machine learning models such as support vector machines (SVM), gradient boosting trees (e.g., XGBoost, LightGBM), and random forests. This enables the realization of an ensemble of diverse models that capture different types of patterns.
[0060] The model retention part also stores the checkpoints of each model and can restore the model to its previous state if necessary. This enables it to return to a previous good state when the performance of the model deteriorates. The checkpoints include the model's parameters, hyperparameters, training state, and performance metrics.
[0061] The weight retention part holds the current weights of each model. The initial weights are set evenly (e.g., for three models, the weight of each model is 1 / 3). These weights are updated periodically based on the recent performance of each model. The weight retention part holds the following information: 1. The current weights of each model 2. The weight update history 3. The parameters used for weight update (e.g., learning rate, regularization parameter, etc.) 4. Weight constraints (e.g., constraints such as minimum value, maximum value, and the sum being 1)
[0062] This information is stored in an in-memory data structure (e.g., hash map, priority queue, etc.) to enable fast access and update. The weight retention part also applies smoothing techniques (e.g., exponential moving average, Kalman filter, etc.) to prevent sudden fluctuations in weights.
[0063] The performance recording part records the past performance of each model. The performance metrics to be recorded include prediction accuracy, F1 score, profit, Sharpe ratio, etc. These metrics are calculated at the end of each evaluation period (e.g., each trading day, each week, etc.) and saved as historical data. The performance recording part records the following information: 1. Basic classification metrics such as accuracy, recall, F1 score, AUC, etc. for each model 2. The confusion matrix (number of true positives, false positives, true negatives, false negatives) for each model 3. The predicted probability distribution for each model (e.g., histogram of predicted probabilities) 4. Profitability indicators of each model (e.g., profit, Sharpe ratio, maximum drawdown, etc.) 5. Temporal variations in the performance of each model
[0064] This information is stored in a time - series database and used for subsequent analysis and weight updates. The performance recording section also provides a function to detect outliers in performance indicators and handle them appropriately.
[0065] The prediction aggregation section weights and aggregates the predictions of each model. The basic approach is a weighted average: y_final = Σ(w_i * y_i) / Σ(w_i) Here, y_final is the final prediction, y_i is the prediction of model i, and w_i is the weight of model i.
[0066] When the predictions are in the form of a probability distribution (probabilities for each class), the prediction aggregation section calculates the final probability distribution using the following formula: p_final(c) = Σ(w_i * p_i(c)) / Σ(w_i) Here, p_final(c) is the final probability for class c, and p_i(c) is the probability for class c by model i.
[0067] As a more advanced approach, the prediction aggregation section can also implement stacking. In stacking, the predictions of each model are used as new features, and a meta - model (e.g., logistic regression, neural network, etc.) is trained to generate the final prediction. This approach can capture more complex relationships between models than a simple weighted average.
[0068] The implementation of stacking includes the following steps: 1. Train each base model with the training data 2. Use each base model to generate predictions for the validation data 3. Use these predictions as new features to train the meta-model 4. For the test data, generate predictions for each base model and use them as inputs to the meta-model to generate the final predictions
[0069] As the meta-model, logistic regression, random forest, gradient boosting, neural network, etc. are used. The selection of the meta-model depends on the complexity of the problem and the constraints of computational resources.
[0070] The prediction aggregation unit also provides the following functions: 1. Conditional aggregation: Use different aggregation methods according to the market situation 2. Confidence weighting: Adjust the weights based on the confidence of the predictions of each model 3. Diversity promotion: Assign high weights to models that provide diverse predictions 4. Detection and handling of abnormal predictions: Detect abnormal predictions and handle them appropriately 5. Optimization of aggregation parameters: Optimize the aggregation parameters regularly
[0071] The weight update unit updates the weights based on the recent performance of each model. The basic approach is to calculate the recent performance of each model using the exponentially weighted average and update the weights based on it: w_i = exp(λ * p_i) / Σ(exp(λ * p_j)) Here, p_i is the recent performance of model i (e.g., accuracy, F1 score, profit, etc.), and λ is a parameter that emphasizes the difference in performance. The larger the value of λ, the higher the sensitivity to the difference in performance, and the higher the weight is assigned to the model that shows the best performance.
[0072] For the calculation of \(p_i\), performance metrics within an exponentially decaying time window are used: \(p_i=\frac{\sum(e^{-\mu*(t - s)}*perf_i(s))}{\sum(e^{-\mu*(t - s)})}\) Here, \(perf_i(s)\) is the performance of model \(i\) at time \(s\), and \(\mu\) is the decay rate. Due to this exponential decay, higher weights are assigned to more recent performance.
[0073] As a more advanced approach, the weight update unit can also implement online learning algorithms (e.g., exponentially weighted average predictors, hedge algorithms, etc.). These algorithms can adapt to the temporal variations in the performance of each model and emphasize different models in different market regimes.
[0074] For example, the exponentially weighted average predictor is defined by the following formula: \(w_i(t + 1)=w_i(t)*exp(\eta*r_i(t))\) \(w_i(t + 1)=\frac{w_i(t + 1)}{\sum(w_j(t + 1))}\) Here, \(w_i(t)\) is the weight of model \(i\) at time \(t\), \(r_i(t)\) is the reward of model \(i\) at time \(t\) (e.g., 1 for a correct prediction and 0 for an incorrect prediction), and \(\eta\) is the learning rate.
[0075] The hedge algorithm is defined by the following formula: \(w_i(t + 1)=w_i(t)*(1-\varepsilon+\varepsilon*\frac{r_i(t)}{r_{avg}(t)})\) Here, \(\varepsilon\) is the hedge parameter, and \(r_{avg}(t)\) is the average reward at time \(t\).
[0076] The weight update unit also provides the following functions: 1. Adaptive learning rate: Adjust the learning rate according to the performance variation 2. Regularization: Apply regularization to prevent the over-concentration of weights 3. Exploration and exploitation balance: Balance the exploration of new models and the exploitation of known good models 4. Weight constraints: Apply constraints so that the weights fall within a specific range 5. Weight smoothing: Apply smoothing to prevent sudden fluctuations in weights
[0077] The noise reduction component removes noise from the input data and improves the signal-to-noise ratio. Wavelet decomposition can analyze signals in both the time and frequency domains, so it is suitable for dealing with noise at various scales contained in financial time series data.
[0078] The noise reduction component includes a wavelet setting unit, a wavelet decomposition unit, a threshold calculation unit, a threshold processing unit, and a signal reconstruction unit. The wavelet setting unit sets the wavelet type (e.g., Daubechies, Symlet, Coiflet, etc.) and the decomposition level to be used. The selection of the wavelet type depends on the characteristics of the signal. The following wavelet types are commonly used for financial time series data: 1. Daubechies wavelet: There are different orders (db1, db2, db4, db8, etc.), and the balance between smoothness and locality is different. db1 is the most local but the least smooth, and the higher the order, the smoother but the lower the locality. For financial time series data, db4 or db8 is commonly used. 2. Symlet wavelet: A symmetric version of the Daubechies wavelet with less phase distortion. Symlets also have different orders (sym2, sym4, sym8, etc.), and the higher the order, the smoother. For financial time series data, sym4 or sym8 is commonly used. 3. Coiflet wavelet: It has high vanishing moments and is suitable for approximating smooth signals. There are also different orders of Coiflet (such as coif1, coif2, coif3), and the higher the order, the smoother it becomes. For financial time series data, coif2 or coif3 is generally used. 4. Discrete Meyer wavelet: It is defined in the frequency domain and has good time-frequency localization. It is suitable for analyzing financial time series data.
[0079] The selection of wavelet type depends on the characteristics of the signal (such as smoothness, discontinuity, etc.) and the purpose of the analysis (such as noise reduction, feature extraction, etc.). A common approach is to try multiple wavelet types and select the one that provides the best results.
[0080] The selection of the decomposition level depends on the length of the signal and the characteristics of the noise. Generally, levels from 3 to 5 are used. The higher the level, the more low-frequency noise can be removed, but the distortion of the signal may also increase. The upper limit of the decomposition level is restricted by the length of the signal. Specifically, the maximum decomposition level should not exceed log2(N) (where N is the length of the signal).
[0081] The wavelet setting part provides the following functions: 1. Automatic wavelet selection: Select the optimal wavelet type based on the characteristics of the signal 2. Automatic level selection: Select the optimal decomposition level based on the length of the signal and the characteristics of the noise 3. Parameter optimization: Optimize the wavelet parameters using cross-validation or information criteria 4. Adaptive setting: Adapt the wavelet setting according to the temporal changes of the signal 5. Multi-wavelet analysis: Analyze the signal using multiple wavelet types and integrate the results
[0082] The wavelet decomposition unit performs wavelet decomposition on the time series data. Wavelet decomposition decomposes the signal into components of different scales. Specifically, the signal is decomposed into an approximation coefficient (low-frequency component) and detail coefficients (high-frequency components). In an n-level decomposition, one approximation coefficient and multiple detail coefficients are generated: [a_n, d_n, d_{n - 1},..., d_1] = DWT(x, n) Here, a_n is the approximation coefficient at level n, d_i is the detail coefficient at level i, DWT is the discrete wavelet transform, x is the input signal, and n is the decomposition level.
[0083] The approximation coefficient represents the low-frequency component of the signal and captures the overall trend. The detail coefficients represent the high-frequency components of the signal and capture fine fluctuations and noise. Generally, since noise mainly appears in the detail coefficients, noise can be reduced by applying threshold processing to the detail coefficients.
[0084] The wavelet decomposition unit provides the following functions: 1. Efficient implementation: Use the fast wavelet transform (FWT) algorithm to improve computational efficiency 2. Boundary processing: Provide boundary processing methods (such as periodic extension, symmetric extension, etc.) to minimize distortion at the boundaries of the signal 3. Parallel computing: Parallelize the decomposition calculation using multiple CPU cores or GPUs 4. Incremental decomposition: Incrementally update the decomposition every time a new data point arrives 5. Multi-level analysis: Analyze detail coefficients at different levels to identify noise at different scales
[0085] The threshold calculation unit calculates an adaptive threshold to apply to the detail coefficients. The selection of the threshold affects the balance between the effect of noise reduction and the distortion of the signal. Common threshold selection methods include the following: 1. Universal threshold: T = σ * sqrt(2 * log(N)) Here, σ is the noise level (usually estimated using the median absolute deviation (MAD) of the detail coefficients), and N is the data length. The MAD is calculated by the following formula: σ = median(|d_i - median(d_i)|) / 0.6745 Here, d_i is the detail coefficient, and 0.6745 is the normalization coefficient of the median absolute deviation of the normal distribution. 2. SURE threshold (Stein's Unbiased Risk Estimator): Select the threshold that minimizes the risk function. The SURE threshold is calculated as the value that minimizes the risk function defined by the following formula: SURE(T) = N - 2 * #{i: |d_i| ≦ T} + Σ(min(|d_i|, T)^2) Here, #{i: |d_i| ≦ T} is the number of detail coefficients with absolute values smaller than T. 3. Minimax threshold: Select the threshold that minimizes the risk in the worst case. The minimax threshold is approximated by the following formula: T = σ * (0.3936 + 0.1829 * log2(N)) 4. Adaptive threshold: Adjust the threshold based on the local characteristics of the detail coefficients. For example, use a low threshold in regions where the local variance of the detail coefficients is large, and use a high threshold in regions where the variance is small. The adaptive threshold is calculated by the following formula: T_i = σ_i * sqrt(2 * log(N)) Here, σ_i is the local standard deviation of the detail coefficients.
[0086] The threshold calculation unit also provides the following functions: 1. Level-dependent threshold: Calculate different thresholds for each decomposition level 2. Signal-dependent threshold: Adjust the threshold based on the characteristics of the signal 3. Noise estimation: Provide multiple methods for estimating the noise level from the detail coefficients 4. Threshold optimization: Optimize the threshold using cross-validation or information criteria 5. Threshold based on confidence interval: Set the threshold based on the confidence interval of the detail coefficients
[0087] The threshold processing unit applies threshold processing to the detail coefficients to remove noise. Common threshold processing methods are as follows: 1. Hard thresholding: Set coefficients smaller than the threshold to zero and keep the other coefficients unchanged. d'_i = d_i * I(|d_i| > T) Here, d_i is the original detail coefficient, d'_i is the processed detail coefficient, T is the threshold, and I() is the indicator function. 2. Soft thresholding: Set coefficients smaller than the threshold to zero and shrink the other coefficients by the amount of the threshold. d'_i = sign(d_i) * max(0, |d_i| - T) Here, sign() is the sign function. 3. Garrote thresholding: An intermediate method between hard thresholding and soft thresholding. d'_i = d_i * max(0, 1 - (T / |d_i|)^2)
[0088] Soft thresholding is commonly used for financial time series data to maintain the continuity of the signal. Hard thresholding preserves the amplitude of the signal but may introduce discontinuities. Garrote thresholding has intermediate characteristics between these two methods.
[0089] The threshold processing unit also provides the following functions: 1. Adaptive Thresholding: Adapt the thresholding method based on the local characteristics of the detail coefficients 2. Hybrid Thresholding: Combine multiple thresholding methods 3. Hierarchical Thresholding: Apply different thresholding methods at different levels 4. Block Thresholding: Apply thresholding to blocks of detail coefficients 5. Probabilistic Thresholding: Process detail coefficients based on a probability model
[0090] The signal reconstruction unit reconstructs the original signal from the processed coefficients. The reconstruction is performed using the Inverse Discrete Wavelet Transform (IDWT): x' = IDWT(a_n, d'_n, d'_{n-1}, ..., d'_1) Here, x' is the reconstructed (denoised) signal, a_n is the approximation coefficient, d'_i is the processed detail coefficient, and IDWT is the Inverse Discrete Wavelet Transform
[0091] The reconstructed signal is a noise-reduced version of the original signal and is used as the input to the prediction model. The signal reconstruction unit also provides the following functions: 1. Efficient Implementation: Improve computational efficiency using the Inverse Fast Wavelet Transform (IFWT) algorithm 2. Boundary Processing: Provide a boundary processing method to minimize distortion at the boundaries of the reconstructed signal 3. Parallel Computation: Parallelize the reconstruction calculation using multiple CPU cores or GPUs 4. Incremental Reconstruction: Incrementally update the reconstruction whenever new processed coefficients become available 5. Selective Reconstruction: Reconstruct the signal using only the coefficients at specific levels or frequency bands
[0092] The noise reduction component also provides the following additional functions: 1. Noise characteristic analysis: Analyze the noise characteristics (distribution, spectrum, autocorrelation, etc.) of the input signal 2. Noise reduction evaluation: Quantitatively evaluate the effect of noise reduction (e.g., signal-to-noise ratio, mean squared error, etc.) 3. Parameter optimization: Optimize the noise reduction parameters using cross-validation or information criteria 4. Adaptive noise reduction: Adapt the noise reduction parameters according to the temporal changes of the signal 5. Multiscale noise reduction: Apply different reduction strategies to noises at different scales
[0093] In addition to these five main components, this system includes an integration module. This module integrates the information from each component and generates the final trading decision. The integration module includes an information integration unit, a decision generation unit, and an execution management unit. The information integration unit integrates the information (predictions, reliability, risk assessment, etc.) from each component. The decision generation unit generates trading decisions (buy, sell, hold, etc.) based on the integrated information. The execution management unit manages the execution of the trading decision and optimizes the execution strategy according to the market situation.
[0094] The information integration unit integrates the information from each component. The information to be integrated includes the following: 1. Predictions from the ensemble model (probabilities of each class) 2. Reliability score of the prediction 3. Importance (attention score) of each time scale 4. Contribution degree of each model 5. Market liquidity indicator 6. Risk assessment
[0095] This information is passed to the decision generation unit. The information integration unit provides the following functions: 1. Information normalization: Normalize information at different scales to make it comparable 2. Information weighting: Weight the information based on the importance of each information source 3. Information verification: Verify the consistency and reliability of information 4. Information aggregation: Aggregate information from multiple information sources 5. Information visualization: Visually represent the integrated information
[0096] The decision-making unit generates a trading decision based on the integrated information. The trading decision includes the direction of the trade (buy, sell, hold), the trading size, and the order type (market order, limit order, etc.).
[0097] The direction of the trade is determined based on the predicted class and the confidence level of the prediction. For example, if the probability of class +1 exceeds a threshold (e.g., 0.6) and the confidence score is high, a buy decision is generated. Similarly, if the probability of class -1 exceeds the threshold and the confidence score is high, a sell decision is generated. Otherwise, a hold decision is generated.
[0098] The trading size is determined based on the confidence level of the prediction, market liquidity, and risk assessment. A common approach is to calculate the trading size using a weighted sum of these factors: size = base_size * (w_1 * confidence + w_2 * liquidity + w_3 * (1 - risk)) Here, base_size is the basic trading size, confidence is the confidence level of the prediction (ranging from 0 to 1), liquidity is the market liquidity (ranging from 0 to 1), risk is the risk assessment (ranging from 0 to 1), and w_1, w_2, w_3 are the weights of each factor.
[0099] The order type is determined based on market liquidity and volatility. For example, in a situation of high liquidity and low volatility, a limit order may be appropriate. Conversely, in a situation of low liquidity and high volatility, a market order may be appropriate.
[0100] The decision-making unit also provides the following functions: 1. Risk management: Evaluate the risk of positions and apply risk limits 2. Portfolio optimization: Optimize trading decisions across multiple assets 3. Transaction cost analysis: Optimize trading decisions considering transaction costs (spread, fees, slippage, etc.) 4. Scenario analysis: Simulate the results of trading decisions under different market scenarios 5. Explanation of decisions: Provide information explaining the basis for trading decisions
[0101] The Execution Management Department manages the execution of trading decisions. This component is responsible for order sending, order tracking, order modification or cancellation, and evaluation of execution quality.
[0102] Order sending is done through the API of the exchange or trading platform. The Execution Management Department optimizes the order sending timing and implements strategies to minimize market impact. For example, strategies such as splitting large orders into multiple small orders or sending orders at staggered times are used.
[0103] Order tracking monitors the status of sent orders (pending, partially filled, fully filled, canceled, etc.). The Execution Management Department modifies or cancels orders as necessary based on the order status.
[0104] Evaluation of execution quality includes comparison of the actual execution price with the expected price, measurement of slippage, evaluation of market impact, etc. These metrics are used to improve and optimize trading strategies.
[0105] The Execution Management Department also provides the following functions: 1. Optimal execution algorithm: Select the optimal execution algorithm according to market conditions 2. Minimize execution costs: Implement strategies to minimize execution costs (spread, market impact, timing cost, etc.) 3. Perform risk management: Manage risks during execution (such as price fluctuations, liquidity depletion, etc.) 4. Analyze execution performance: Analyze execution performance and identify areas for improvement 5. Generate execution reports: Generate reports documenting the details of the execution
[0106] Furthermore, this system includes the following additional modules: 1. Outlier handling module: Detect and handle outliers in market data. This module uses outlier detection algorithms (such as Z-score, interquartile range, isolation forest, etc.) to identify abnormal data points and processes them appropriately (excluding, interpolating, robust transformation, etc.). 2. Reliability assessment module: Evaluate the reliability of predictions and adjust the trading size based on the reliability. This module calculates a reliability score considering factors such as prediction probability, agreement between models, and stability of predictions, and determines the trading size based on it. 3. Liquidity assessment module: Evaluate market liquidity and optimize the trading execution strategy based on liquidity. This module uses indicators such as bid-ask spread, market depth, trading volume, etc. to evaluate liquidity and determines the optimal order type and order size. 4. Seasonality detection module: Detect seasonality and periodicity in market data and adjust predictions based on it. This module uses techniques such as Fourier transform, wavelet transform, autocorrelation analysis, etc. to identify seasonality such as intraday patterns, weekly patterns, monthly patterns, etc. 5. Structural change detection module: Detect structural changes in the market and retrain the model based on it. This module uses change point detection algorithms (such as CUSUM, PELT, BICP, etc.) to identify changes in the market regime. 6. Correlation Analysis Module: Analyzes the correlation between multiple assets and enhances predictions based on it. This module quantifies the relationships between assets using metrics such as Pearson correlation, Spearman correlation, and mutual information. 7. Macroeconomic Integration Module: Considers the macroeconomic factors of the market and adjusts predictions based on them. This module collects information such as economic indicators, central bank policies, and geopolitical events, and analyzes the impact they have on the market. 8. Transaction Analysis Module: Analyzes transaction histories and provides insights for optimizing transaction performance. This module calculates metrics such as win rate, profit-loss ratio, maximum drawdown, and Sharpe ratio, and identifies the strengths and weaknesses of trading strategies.
[0107] The training process of this system consists of the following steps: 1. Extract features from historical market data: Use a multi-resolution time processing engine to extract features at multiple time scales. 2. Remove noise from features using a noise reduction component: Apply wavelet decomposition and threshold processing to improve the signal-to-noise ratio. 3. Update source domain statistics: Use the source statistics calculation part of the domain adaptation module to calculate and save the statistical information of the training data. 4. Update the weights of labels based on class distribution: Use an adaptive label balancing framework to calculate the optimal weights for each class. 5. Train each model using weighted loss: Train each model of the ensemble model architecture using the calculated weights.
[0108] The training process can be executed batch-wise or incrementally. In batch processing, the model is trained once using all the training data. In the incremental approach, the model is updated each time new data becomes available. The latter approach is suitable for continuously adapting to changes in market conditions. In the case of batch processing, the training is executed in the following steps: 1. Divide the training data into mini-batches 2. For each mini-batch, perform the following operations: a. Feature extraction and noise reduction b. Domain adaptation c. Forward propagation (generation of predictions by the model) d. Calculation of weighted loss e. Backward propagation (calculation of gradients) f. Parameter update (using an optimization algorithm) 3. After processing all the mini-batches, evaluate the model using the validation data 4. Repeat steps 1 - 3 until the convergence condition is met In the case of the incremental approach, the training is executed in the following steps: 1. Train an initial model (similar to batch processing) 2. Each time new data becomes available, perform the following operations: a. Feature extraction and noise reduction b. Domain adaptation c. Generation and evaluation of predictions d. Update the model (using an online learning algorithm) 3. Periodically, perform a complete retraining of the model (e.g., when performance degrades, or after a certain period of time) The following optimization techniques are also used in the training process: 1. Learning rate scheduling: Adjust the learning rate according to the progress of training (e.g., cyclic learning rate, cosine annealing, step decay, etc.) 2. Regularization: Techniques to prevent overfitting (e.g., L1 regularization, L2 regularization, dropout, early stopping, etc.) 3. Batch normalization: Reduces internal covariate shift and improves training stability 4. Gradient clipping: Limits the magnitude of gradients to prevent gradient explosion problems 5. Model ensemble: Trains multiple models and combines their predictions The prediction process consists of the following steps: 1. Extract features from current market data: Using a multi-resolution time processing engine, extract features at multiple time scales. 2. Remove noise from features: Using a noise reduction component, remove noise from the extracted features. 3. Update target domain statistics and adapt features: Using a domain adaptation module, calculate statistical information of recent market data and adapt the features. 4. Generate predictions using an ensemble: Using an ensemble model architecture, weight and aggregate the predictions of each model to generate the final prediction. 5. Generate trading decisions based on predictions: Using an integration module, generate trading decisions considering factors such as predictions, confidence levels, and market conditions. 6. Execute trading decisions: Using the execution management unit of the integration module, execute trading decisions and evaluate the execution quality. The prediction process can be executed in real-time or periodically. In the real-time approach, predictions are generated whenever new market data becomes available. In the periodic approach, predictions are generated at fixed intervals (e.g., every 1 second, every 5 seconds, etc.). The appropriate approach depends on the trading strategy and the characteristics of the market. In the case of the real-time approach, the prediction is executed in the following steps: 1. Continuously monitor the market data stream 2. Whenever a new data point arrives, perform the following operations: a. Feature extraction and noise reduction b. Domain adaptation c. Generation of predictions d. Generation and execution of trading decisions 3. Continuously monitor the market situation and the performance of the predictions, and adjust the parameters as necessary In the case of a periodic approach, the prediction is performed in the following steps: 1. Collect data at fixed intervals 2. At the end of each interval, perform the following operations: a. Feature extraction and noise reduction b. Domain adaptation c. Generation of predictions d. Generation and execution of trading decisions 3. Periodically evaluate the prediction performance and adjust the parameters as necessary In the prediction process, the following optimization techniques are also used: 1. Caching: Cache the frequently used calculation results to reduce the calculation overhead 2. Parallel processing: Use multiple CPU cores or GPUs to parallelize the calculations 3. Batch processing: Process multiple data points as a batch to improve the throughput 4. Model quantization: Convert the model parameters to a low-precision representation to reduce the memory usage and the calculation cost 5. Model distillation: Transfer the knowledge of a large model to a small model to improve the calculation efficiency The system is periodically updated based on the latest data, labels, and model performance. This update process includes the following steps: 1. Update of the domain adaptation module: Calculate the statistical information of recent market data and update the target domain statistics. 2. Update of the label balancer: Update the weights of each class based on the recent class distribution and model performance. 3. Update of the ensemble weights: Update the weights of the models based on the recent performance of each model. 4. Retraining of the model: Retrain the model using the latest data as needed. The update frequency is set based on market dynamics and computational resource constraints. A common approach is to update different components at different frequencies. For example, the domain adaptation module and the ensemble weights are updated frequently (e.g., every few minutes), while the retraining of the model is performed less frequently (e.g., every few hours or days). The update process is executed according to the following schedule: 1. Real-time updates (from milliseconds to seconds): - Feature extraction and noise reduction - Generation of predictions - Generation and execution of trading decisions 2. Short-term updates (in minutes): - Update of target domain statistics - Update of ensemble weights - Update of confidence scores - Update of liquidity evaluation 3. Medium-term updates (in hours): - Update of label weights - Incremental update of the model - Update of outlier detection parameters - Update of seasonal patterns 4. Long-term updates (in days or weeks): - Complete retraining of the model - Optimization of Hyperparameters - Update of Feature Selection - Performance Evaluation of the Entire System In the update process, the following techniques are also used: 1. Online Learning: Incrementally update the model whenever new data becomes available 2. Transfer Learning: Transfer the knowledge of the previous model to the new model 3. Active Learning: Select the most informative data points for labeling and training 4. Continuous Learning: Continuously learn new patterns while maintaining the performance of the model 5. Prevention of Catastrophic Forgetting: Use techniques to retain previous knowledge when learning new data
Industrial Applicability
[0109] The present invention is widely applicable to financial institutions, hedge funds, prop trading companies, algorithmic trading companies, etc. that conduct high-frequency trading. In particular, it is useful in trading environments facing problems such as label imbalance, domain shift, and high noise levels.
[0110] The main application fields of the present invention are as follows: 1. High-frequency trading in the stock market: The stock market is characterized by high liquidity and low trading costs and is suitable for high-frequency trading. This system can be used to predict short-term fluctuations in stock prices and obtain profits from market inefficiencies. In particular, it can be applied to trading in different market segments such as large-cap stocks, mid-cap stocks, and small-cap stocks. It can also be applied to trading of derivative products such as ETFs, index futures, and options. 2. High-Frequency Trading in the Futures Market: The futures market offers high leverage and the possibility of 24-hour trading, making it an attractive target for high-frequency trading. This system can be used to predict short-term fluctuations in futures prices and profit from market inefficiencies. In particular, it is applicable to trading in financial futures (such as stock index futures, interest rate futures, currency futures, etc.), commodity futures (such as energy, metals, agricultural products, etc.), and futures in emerging markets. 3. High-Frequency Trading in the Foreign Exchange Market: The foreign exchange market is the world's largest financial market, offering high liquidity and the possibility of 24-hour trading. This system can be used to predict short-term fluctuations in exchange rates and profit from market inefficiencies. In particular, it is applicable to trading in major currency pairs (such as EUR / USD, USD / JPY, GBP / USD, etc.), cross currency pairs, and emerging market currencies. 4. High-Frequency Trading in the Cryptocurrency Market: The cryptocurrency market offers high volatility and the possibility of 24-hour trading, making it an emerging field for high-frequency trading. This system can be used to predict short-term fluctuations in cryptocurrency prices and profit from market inefficiencies. In particular, it is applicable to trading in Bitcoin, Ethereum, other major cryptocurrencies, and trading pairs with stablecoins. 5. Market Making: Market makers simultaneously present both buy and sell orders and profit from the bid-ask spread. This system can be used to determine appropriate bid and ask prices and manage inventory risk. In particular, it is useful for market making in low-liquidity markets and emerging markets. 6. Statistical Arbitrage Trading: Statistical arbitrage trading identifies price imbalances between related financial instruments and then profits from them. This system can be used to detect these imbalances and execute appropriate trading strategies. In particular, it is useful for identifying arbitrage opportunities between ETFs and their constituent stocks, futures and physicals, options and underlying assets, etc.
[0111] Furthermore, each component of the present invention is also applicable to other time series prediction tasks. For example: 1. Risk Management: Prediction and management of market risk, credit risk, liquidity risk, etc. In particular, it can be applied to the calculation and prediction of risk indicators such as Value at Risk (VaR), Conditional Value at Risk (CVaR), stress testing, and scenario analysis. It can also be used for the decomposition of risk factors and contribution analysis. 2. Portfolio Optimization: Optimization of asset allocation and development of rebalancing strategies. In particular, it can be applied to the construction of mean-variance optimization, risk parity, minimum variance portfolio, maximum Sharpe ratio portfolio, etc. It can also be used for portfolio optimization considering transaction costs and market impact. 3. Market Monitoring: Detection and monitoring of abnormal market activities. In particular, it can be applied to the detection of illegal acts such as market manipulation, insider trading, and front-running. It also functions as an early warning system for market liquidity depletion, sharp price movements, and abnormal trading volumes. 4. Economic Indicator Prediction: Prediction of economic indicators such as GDP growth rate, inflation rate, unemployment rate, etc. In particular, it can be applied to the prediction of central bank policy decisions, government fiscal policies, international trade trends, etc. It can also be used for the analysis of the impact of these indicators on the financial market. 5. Sentiment Analysis: Extraction and prediction of market sentiment from news, social media, analyst reports, etc. In particular, it can be applied to quantify the sentiment of market participants using technologies such as text mining, natural language processing, and sentiment analysis, and predict the impact it has on the market. 6. Optimization of Algorithmic Trading: Improvement of the performance of existing algorithmic trading strategies. In particular, it can be applied to parameter optimization of execution algorithms (such as VWAP, TWAP, implementation shortfall, etc.), improvement of order splitting strategies, and optimization of timing strategies. 7. Cross-asset analysis: Analysis and prediction of relationships between different asset classes. In particular, it can be applied to capture changes in the correlation structure between stocks and bonds, currencies and commodities, developed and emerging markets, etc., and to adjust trading strategies based on this. 8. Macroeconomic scenario analysis: Prediction of market trends under different macroeconomic scenarios. In particular, it can be applied to predict asset prices under different scenarios such as inflation / deflation, economic growth / recession, monetary easing / tightening, etc.
[0112] Furthermore, each component of the present invention can also be applied to other time series prediction tasks. For example: 1. Risk management: Prediction and management of market risk, credit risk, liquidity risk, etc. In particular, it can be applied to the calculation and prediction of risk indicators such as Value at Risk (VaR), Conditional Value at Risk (CVaR), stress testing, scenario analysis, etc. It can also be used for the decomposition of risk factors and contribution analysis. 2. Portfolio optimization: Optimization of asset allocation and development of rebalancing strategies. In particular, it can be applied to the construction of mean-variance optimization, risk parity, minimum variance portfolio, maximum Sharpe ratio portfolio, etc. It can also be used for portfolio optimization considering transaction costs and market impact. 3. Market monitoring: Detection and monitoring of abnormal market activities. In particular, it can be applied to the detection of illegal activities such as market manipulation, insider trading, front-running, etc. It also functions as an early warning system for market liquidity depletion, sharp price movements, abnormal trading volumes, etc. 4. Economic indicator prediction: Prediction of economic indicators such as GDP growth rate, inflation rate, unemployment rate, etc. In particular, it can be applied to predictions such as central bank policy decisions, government fiscal policies, international trade trends, etc. It can also be used for the analysis of the impact of these indicators on the financial market. 5. Sentiment Analysis: Extracting and predicting market sentiment from news, social media, analyst reports, etc. In particular, it can be applied to quantify the sentiment of market participants using technologies such as text mining, natural language processing, and sentiment analysis, and to predict the impact it has on the market. 6. Optimization of Algorithmic Trading: Improving the performance of existing algorithmic trading strategies. In particular, it can be applied to parameter optimization of execution algorithms (such as VWAP, TWAP, implementation shortfall, etc.), improvement of order splitting strategies, and optimization of timing strategies. 7. Cross-Asset Analysis: Analyzing and predicting relationships between different asset classes. In particular, it can be applied to capture changes in the correlation structure between stocks and bonds, currencies and commodities, developed markets and emerging markets, etc., and to adjust trading strategies based on this. 8. Macro-Economic Scenario Analysis: Predicting market trends under different macro-economic scenarios. In particular, it can be applied to predicting asset prices in different scenarios such as inflation / deflation, economic growth / recession, monetary easing / tightening, etc.
[0113] The present invention is expected to contribute to the improvement of the efficiency and liquidity of the financial market by significantly improving the prediction accuracy, adaptability, robustness, and efficiency in a high-frequency trading environment. In addition, the technology of the present invention may also be applicable not only to the financial market but also to other fields where time series prediction is important (for example, medical, energy, transportation, etc.).
Claims
1. 1. An adaptive intelligence system for high frequency trading, comprising: A multi-resolution time processing engine that processes market data at multiple time scales simultaneously; An adaptive label balancing framework that dynamically adjusts label importance according to changing market conditions; A domain adaptation module that continuously monitors and adapts to distribution shifts in market data; An ensemble model architecture that dynamically adjusts the contributions of multiple predictive models based on their recent performance; a noise reduction component using wavelet decomposition to address high noise levels in financial data; A high-frequency trading adaptive intelligence system comprising:
2. the multi-resolution time processing engine comprises a time scale setting unit that sets a plurality of time scales, a feature extractor generating unit that generates a feature extractor corresponding to each time scale, and a feature integration unit that integrates features extracted by each feature extractor; The adaptive label balancing framework includes a class distribution recorder for recording a class distribution history, a weight holder for holding a current weight for each class, a performance tracker for tracking a performance index for each class, and a weight calculator for calculating an optimal weight based on a distribution factor and a performance factor; 2. The system of claim 1, wherein the domain adaptation module comprises a source statistics calculation unit that calculates and stores statistical information of a source domain, a target statistics calculation unit that calculates and stores statistical information of a target domain, and a feature adaptation unit that adapts features using the statistical information of the source domain and the target domain.
3. The ensemble model architecture includes a model storage unit that stores a plurality of models including a multi-layer perceptron (MLP) model, a long short-term memory (LSTM) network model, and a Mamba model; a weight storage unit that stores weights of each model; a performance recording unit that records a performance history of each model; a prediction aggregation unit that weights and aggregates predictions of each model; and a weight update unit that updates the weights of each model based on recent performance; the noise reduction component comprises a wavelet setting unit for setting a wavelet type and a decomposition level, a wavelet decomposition unit for performing wavelet decomposition on time series data, a threshold calculation unit for calculating an adaptive threshold, a threshold processing unit for applying threshold processing to detail coefficients, and a signal reconstruction unit for reconstructing a signal from the processed coefficients; 3. The system of claim 1 or 2, further comprising an integration module including an information integration unit that integrates information from each component, a decision generation unit that generates a trading decision based on the integrated information, and an execution management unit that manages the execution of the trading decision.
Citation Information
Cited By
Thermal power plant data analysis and diagnosis method
CN120804986A
Urban traffic flow prediction method based on multiple time-space characteristics of road network
CN120913415A
Method and system for regulating and controlling explosion-proof dynamic threshold value of robot
CN120941423A
Remote sensing image multi-scale segmentation method based on frequency spectrum information processing and Mama space modeling
CN120953609A
Face sketch generation method and system based on Mama and wavelet convolution
CN120997036A