Multimodal system and method for finance marketplace prediction using hierarchical aggregation

The Hierarchical Aggregation Financial Prediction Framework addresses the limitations of existing methods by integrating news data and sector-level correlations with time series data, achieving substantial improvements in prediction accuracy and adaptability.

JP2025094148AActive Publication Date: 2025-06-24NYU-YO-KU ZENERAL GURU-PU INKU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025046388
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Existing financial asset price prediction methods fail to effectively incorporate news data and sector-level correlations, leading to suboptimal prediction accuracy and lack of adaptability to market conditions.

Method used

The Hierarchical Aggregation Financial Prediction Framework integrates time series data and news data using a hierarchical temporal news aggregation mechanism, relevance-weighted news aggregation, sector-level correlation integration, and adaptive modality fusion to enhance prediction accuracy and adaptability.

Benefits of technology

This approach significantly improves prediction accuracy by reducing the mean absolute percentage error (MAPE) by over 50%, enhances adaptability to market conditions, and provides transparent and interpretable prediction results.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To provide a multimodal system and method that improve the precision of a financial asset price prediction.SOLUTION: A multimodal system includes: a first processing unit for processing time-series data; a second processing unit for processing news data; and an integrating unit for integrating the output by the first processing unit with the output by the second processing unit. The second processing unit includes: a hierarchical time news aggregating mechanism that aggregates news data at a plurality of different time windows; and a relevancy weighting news collecting mechanism that performs weighting based on the relevancy of news. The integrating unit includes: a sector-level correlation integrating mechanism for adjusting a prediction based on a correlation in a sector level; and an adaptive modality fusing mechanism for dynamically adjusting the weighting of each modality in accordance with a market situation. Accordingly, the time-series data and the news data are efficiently integrated with each other, and thus a highly precise asset price prediction is accomplished.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an artificial intelligence system for financial asset price prediction, and particularly to improving prediction accuracy by a multimodal approach that combines time series data and text data. More specifically, it relates to a financial asset price prediction system and method that combines hierarchical news aggregation using multiple time windows, weighting based on news relevance, integration of sector-level correlation relationships, and adaptive modality fusion according to market conditions.

Background Art

[0002] Predicting asset prices in the financial market is an important issue for investors, traders, fund managers, and financial institutions. Accurate price prediction enables strategic planning, optimal investment portfolio management, and risk assessment. Therefore, the development of more accurate prediction models has always attracted high attention in the financial industry.

[0003] Conventional asset price prediction methods mainly rely on numerical data such as price time series, trading volume, order book data, and technical analysis indicators. These methods range from traditional approaches using technical indicators such as moving averages, relative strength index (RSI), and Bollinger Bands, to statistical methods such as autoregressive integrated moving average (ARIMA) models and generalized autoregressive conditional heteroskedasticity (GARCH) models, and even to the latest approaches using machine learning algorithms and deep learning models.

[0004] In the machine learning approach, algorithms such as support vector machines (SVM), random forests, and gradient boosting decision trees (e.g., XGBoost) are widely used. These algorithms extract features from past price data and technical indicators and learn patterns for predicting future price movements.

[0005] In deep learning approaches, architectures such as long short-term memory (LSTM) networks, gated recurrent units (GRU), convolutional neural networks (CNN), and transformers are used. These models can capture the complex temporal dependencies in time series data and show high prediction performance, especially when a large amount of data is available.

[0006] However, these conventional methods mainly rely on numerical data only and do not fully consider the news flow and information flow that have a significant impact on the behavior of market participants and market sentiment. News flow plays an important role in price formation, and the development of a multimodal approach that combines text data and numerical data is highly relevant.

[0007] In recent years, with the progress of natural language processing (NLP) technology, the ability to analyze text data such as financial news and social media posts and evaluate market sentiment has improved. In particular, the emergence of pre-trained language models (e.g., BERT, RoBERTa, GPT) has made it possible to accurately understand the meaning and context of text, and specialized models (e.g., FinBERT) adapted from these models to the financial domain have also been developed. As a result, attempts to incorporate text data into asset price prediction have been increasing.

[0008] In Non-Patent Document 1, a multimodal approach targeting the Russian securities market has been proposed. In this study, a prediction model combining candlestick time series data and text news flow data was developed, and it was shown that adding the text modality reduces the mean absolute percentage error (MAPE) by 55%. However, in this study as well, only a single time window (yesterday's news) was considered, and the impact of news at different time scales was not considered. Also, weighting based on news relevance and adaptive modality fusion according to market conditions have not been realized.

[0009] Furthermore, existing multimodal approaches lack a mechanism to incorporate sector-level correlations into predictions. In the financial market, it is known that strong correlations exist among assets within the same sector, and accuracy improvement can be expected by leveraging this information in predictions. For example, multiple companies within the technology sector are often affected by similar market factors, so price movements are often correlated.

[0010] As described above, in the prior art, (1) a method for effectively capturing the impact of news at different time scales, (2) a weighted aggregation method based on news relevance, (3) a framework for utilizing sector-level correlations in predictions, and (4) a method for dynamically adjusting the importance of each modality according to market conditions, have not been fully realized. By solving these problems, it is possible to significantly improve the accuracy of financial asset price prediction.

Prior Art Documents

Non-Patent Documents

[0011]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0012] An object of the present invention is to solve the above problems of the prior art and provide a system and method for improving the accuracy of financial asset price prediction.

[0013] Specifically, an object of the present invention is to solve the following problems:

[0014] 1. To provide a method for effectively capturing the impact of news at different time scales: In the prior art, the impact of news was evaluated using a single time window (e.g., only the news of the previous day). However, the impact of news changes over time, and there may be differences between short-term and long-term impacts. The present invention aims to provide a method for capturing the impact of news on different time scales using multiple time windows.

[0015] 2. Realization of a weighted aggregation method based on news relevance: In the prior art, simple aggregation (e.g., averaging) was performed without considering the relevance of news. However, not all news has the same relevance to the target asset. The present invention aims to realize a weighted aggregation method that evaluates the relevance of news and gives greater influence to more relevant news.

[0016] 3. Construction of a framework for utilizing sector-level correlation relationships in prediction: In the prior art, individual assets were often predicted independently, and sector-level correlation relationships were not considered. The present invention aims to construct a framework that utilizes the correlation relationships among assets within the same sector to improve the robustness of prediction.

[0017] 4. Provision of a method for dynamically adjusting the importance of each modality according to market conditions: In the prior art, fixed weighting was often used in the integration of time series data and news data. However, the relative importance of each modality changes depending on market conditions. The present invention aims to provide a method for dynamically adjusting the importance of each modality according to market conditions (e.g., high or low volatility, importance of news, etc.).

[0018] 5. Realization of a practical and highly accurate financial asset price prediction system integrating the above elements: The object of the present invention is to realize a practical and highly accurate financial asset price prediction system that integrates the above elements. This system aims to significantly improve the prediction accuracy compared to the prior art by effectively integrating time series data and news data and considering market conditions and sector-level correlation relationships.

Means for Solving the Problems

[0019] To solve the above problems, the present invention provides a new artificial intelligence system called "Hierarchical Aggregation Financial Prediction Framework" (hereinafter referred to as HierAggFin). This system is composed of the following main components:

[0020] 1. Hierarchical Temporal News Aggregation Mechanism: A mechanism that captures the impact of news at different time scales using multiple time windows (e.g., 1 day, 3 days, 7 days, 30 days). This mechanism calculates a weighted average by applying an exponential decay based on temporal proximity to the news vectors within each time window. This enables capturing the influence of more recent news with high weights while also considering the long-term impact of news.

[0021] 2. Relevance-Weighted News Aggregation Mechanism: A mechanism that evaluates the relevance of news articles to the target asset and gives greater influence to more relevant news. The evaluation of relevance is based on elements such as the presence of company names / tickers in the title, the frequency of keyword occurrences in the text, and the position of company mentions.

[0022] 3. Sector-Level Correlation Integration Mechanism: A mechanism that improves the robustness of predictions by leveraging the correlation relationships among assets within the same sector. This mechanism calculates the correlation coefficients between the target asset and other assets within the same sector and adjusts the prediction of the target asset based on the predictions of assets with significant correlation relationships.

[0023] 4. Adaptive Modality Fusion Mechanism: A mechanism that dynamically adjusts the weights of time-series data and news data based on market volatility and the importance of news. For example, in a high-volatility market, increase the weight of time-series data, and increase the weight of news data when important news events occur.

[0024] By combining these components, time-series data and news data can be effectively integrated to achieve high-precision asset price prediction.

[0025] According to one aspect of the present invention, a multimodal system for predicting financial asset prices is provided. This system includes a first processing unit for processing time-series data, a second processing unit for processing news data, and an integration unit for integrating the outputs of the first processing unit and the second processing unit. The second processing unit includes a hierarchical time news aggregation mechanism for aggregating news data in a plurality of different time windows, and a relevance weighting news aggregation mechanism for weighting based on the relevance of news. The integration unit includes a sector-level correlation integration mechanism for adjusting the prediction based on the sector-level correlation relationship, and an adaptive modality fusion mechanism for dynamically adjusting the weights of each modality according to the market situation.

[0026] According to another aspect of the present invention, a multimodal method for predicting financial asset prices is provided. This method includes a step of processing time-series data, a step of processing news data, and a step of integrating the processing results of the time-series data and the processing results of the news data. The step of processing the news data includes a hierarchical time news aggregation step of aggregating news data in a plurality of different time windows, and a relevance weighting news aggregation step of weighting based on the relevance of news. The step of integrating includes a sector-level correlation integration step of adjusting the prediction based on the sector-level correlation relationship, and an adaptive modality fusion step of dynamically adjusting the weights of each modality according to the market situation.

Advantages of the Invention

[0027] The present invention provides the following effects:

[0028] 1. A significant improvement in prediction accuracy: The multimodal approach of the present invention enables a reduction of more than 50% in the mean absolute percentage error (MAPE). This allows investors and traders to make decisions based on more accurate price predictions. For example, if the MAPE of a conventional prediction model based only on time series was about 0.4%, it may be reduced to 0.2% or less with the approach of the present invention. This means that for a stock price of 10,000 yen, the prediction error decreases from 40 yen to 20 yen or less, bringing significant economic benefits especially in large-scale transactions and high-frequency trading.

[0029] 2. An improvement in the prediction accuracy of the price movement direction: The prediction accuracy of the price movement direction, whether it is rising or falling, is improved, enhancing the success rate of trading strategies. In a conventional prediction model based only on time series, the accuracy of direction prediction is often about 52%, but it may be improved to 55% or more with the approach of the present invention. This brings a significant difference in profits in the long run even in a simple binary prediction of "rising / falling".

[0030] 3. An improvement in the adaptability to changes in market conditions: The prediction model is adaptively adjusted according to market volatility and the importance of news, so stable prediction performance can be maintained in different market environments. A conventional fixed model cannot cope with changes in the market environment (for example, the transition from a low-volatility market to a high-volatility market), and the prediction accuracy may decrease significantly. However, with the approach of the present invention, stable prediction performance can be maintained even in such situations.

[0031] 4. An improvement in the transparency and interpretability of the prediction basis: The hierarchical structure and explicit weighting mechanism make it easier to understand the factors on which the prediction results are based. This enables users to easily evaluate the reliability of the prediction results. For example, information such as "This prediction is mainly strongly influenced by short-term news (1-day window)" or "This prediction is based on the trends of correlated assets within the sector" can be provided.

[0032] 5. Applicability to Different Markets and Languages: The proposed framework can be applied to different markets and languages by replacing the appropriate language model (e.g., Japanese BERT for the Japanese market, BERT / RoBERTa for the English market, Chinese BERT for the Chinese market, etc.). Thus, it can be widely utilized as a prediction model in the global financial market.

[0033] 6. Improvement in Computational Efficiency: Due to the hierarchical structure, news aggregation for each time window can be processed in parallel, making it possible to efficiently process a large amount of news data. This enables near-real-time prediction updates and can be used for high-frequency trading and immediate investment decisions.

[0034] 7. Improvement in Scalability: With the modular design, news data processing can be performed in parallel streams, allowing the system to scale according to the increase in the amount of news. Thus, it can also handle a large number of news sources and news feeds updated frequently.

Modes for Carrying Out the Invention

[0035] Hereinafter, embodiments of the present invention will be described in detail.

[0036] System Configuration This system is composed of a data collection unit, a preprocessing unit, a time series processing unit, a news processing unit, a prediction integration unit, and an output unit. The data collection department collects time-series data (prices, trading volumes, and other market indicators) from the financial market and also collects news articles from various sources. The time-series data is collected from the APIs of stock exchanges, financial data providers (e.g., Bloomberg, Refinitiv, Quick), or public data sources (e.g., Yahoo Finance, Alpha Vantage). The news articles are collected from various sources such as financial news websites, corporate press releases, social media, and experts' blogs. The collected data is sent to the preprocessing department. The preprocessing department performs processing such as normalization of time-series data, removal of outliers, and imputation of missing values, and also performs processing such as cleaning, tokenization, and vectorization of news articles. In the preprocessing of time-series data, the following processing is performed: 1. Imputation of missing values: Impute missing values using linear interpolation, forward fill, or machine learning-based imputation methods. 2. Outlier processing: Identify and process outliers using moving average, median filtering, or statistical methods (e.g., outlier detection based on Z-score). 3. Scaling: Scale each feature to an appropriate range (e.g., from 0 to 1, or mean 0 and standard deviation 1). 4. Feature extraction: Extract features such as relative price changes (returns), volatility, and momentum from candlestick data. 5. Application of time window: Form input vectors using data from a certain period in the past (e.g., 5 business days, 10 business days, etc.). In the preprocessing of news articles, the following processing is performed: 1. Cleaning: Removal of HTML elements, normalization of special characters, removal of stop words, and processing of punctuation marks, etc. 2. Tokenization: Tokenize the text using the tokenizer of a language model (such as BERT, RoBERTa, etc.). 3. Vectorization: Convert news articles into vector representations using pre-trained language models. Specifically, the final-layer hidden state of the [CLS] token or the average of the hidden states of all tokens may be used to obtain the representation of the entire article. 4. Filtering: Use keyword matching, entity recognition, or a dedicated classifier to select news articles related to the target asset. The preprocessed data is sent to the time series processing unit and the news processing unit respectively. The time series processing unit uses a deep learning model such as a long short-term memory (LSTM) network to generate price predictions from time series data. Since LSTM can capture the long-term dependencies of time series data, it is suitable for predicting financial time series. The specific configuration of the time series processing unit is as follows: 1. Input layer: Receive preprocessed time series features (e.g., relative price changes over the past 5 business days, volatility, etc.). 2. LSTM layer: A layer consisting of multiple LSTM cells that learns the temporal dependencies of time series data. Usually, two or more LSTM layers are stacked. 3. Dropout layer: Apply dropout regularization to prevent overfitting. 4. Fully connected layer: Receive the output of the LSTM and convert it into the final prediction value. 5. Output layer: Output the prediction value (e.g., relative price change on the next business day). The news processing unit includes a hierarchical time news aggregation mechanism and a relevance weighted news aggregation mechanism, and generates price predictions from news data. The specific configuration of the news processing unit is as follows: 1. Input layer: Receive preprocessed news vectors. 2. Hierarchical time news aggregation layer: Aggregate news vectors in multiple time windows. 3. Relevance weighted news aggregation layer: Perform weighting based on the relevance of news. 4. Fully connected layer: Receive the aggregated news vectors and convert them into intermediate representations. 5. Output layer: Outputs the predicted value (e.g., the relative price change on the next business day). The prediction integration unit includes a sector-level correlation integration mechanism and an adaptive modality fusion mechanism, and integrates the predictions from the time-series processing unit and the news processing unit to generate a final prediction. The specific configuration of the prediction integration unit is as follows: 1. Input layer: Receives the predicted values from the time-series processing unit and the news processing unit. 2. Sector-level correlation integration layer: Adjusts the prediction based on the correlation within the sector. 3. Adaptive modality fusion layer: Dynamically adjusts the weights of each modality based on the market situation. 4. Output layer: Outputs the final predicted value (e.g., the relative price change on the next business day). The output unit provides information such as the predicted price, price movement direction, prediction confidence level, etc. to the user. The output unit provides the following information: 1. Predicted price: The predicted closing price on the next business day (or a specified future point in time). 2. Price movement direction: The prediction of increase or decrease. 3. Prediction confidence level: An indicator showing the confidence level of the prediction (e.g., probability value or confidence interval). 4. Explanation information: The main factors contributing to the prediction (important news, trends of correlated assets, etc.). 5. Visualization: Graphical display of the prediction results, the trend of past prediction accuracy, etc.

[0037] Hierarchical Time News Aggregation Mechanism The hierarchical time news aggregation mechanism is one of the core features of the present invention. This mechanism aims to capture the impact of news at different time scales using multiple time windows. In the financial market, the impact of news changes over time. For example, important news such as a company's quarterly earnings announcement may cause significant price fluctuations immediately after the announcement, but its impact decays over time, and the influence of other factors may become greater after a few days. On the other hand, news such as industry-wide trends or regulatory changes may continue to affect prices over a long period. To capture the impact of news at such different time scales, the hierarchical time news aggregation mechanism defines multiple time windows (e.g., 1 day, 3 days, 7 days, 30 days) and aggregates the news vectors within each time window. This enables capturing both short-term market reactions and long-term market sentiment. Specifically, based on the prediction target date, for the news vectors within each time window, a weighted average is calculated by applying an exponential decay based on temporal proximity. This allows giving higher weights to more recent news while also capturing the impact of long-term news. The hierarchical time news aggregation process is executed in the following steps: Step 1: Define multiple time windows (e.g., 1 day, 3 days, 7 days, 30 days) based on the prediction target date. For example, if the prediction target date is April 10, 2024, - 1-day window: News on April 9, 2024 - 3-day window: News from April 7 to April 9, 2024 - 7-day window: News from April 3 to April 9, 2024 - 30-day window: News from March 11 to April 9, 2024 are targeted. Step 2: Select the news vectors within each time window. For example, in the 1-day window, only the news from the day before the target prediction day is selected, and in the 3-day window, the news from the three days before the target prediction day is selected. Here, the news vector is obtained by converting a news article into a vector representation using a pre-trained language model (e.g., BERT, RoBERTa). Usually, the dimension of the news vector is determined by the language model used. For example, it is 768 dimensions for BERT-base and 1024 dimensions for BERT-large. Step 3: Calculate the weights based on the temporal proximity of each news. Specifically, apply an exponential decay based on the time difference between the target prediction day and the news release day. For example, if the time difference is d days, the weight w is calculated by the following formula: w = exp(-λ * d) Here, λ is the decay rate parameter, and usually a value in the range of 0.1 to 0.5 is used. This parameter controls how much the importance of the news decays over time. The larger λ is, the faster the decay, and the smaller λ is, the slower the decay. For example, when λ = 0.2, the weight of the news from 1 day ago is exp(-0.2 * 1) = 0.819, and the weight of the news from 5 days ago is exp(-0.2 * 5) = 0.368. The decay rate parameter λ can be adjusted according to the volatility and liquidity of the asset, the characteristics of the market, etc. For example, in high-volatility assets or high-liquidity markets, since the impact of news is reflected quickly, a larger λ value may be appropriate. On the other hand, in low-volatility assets or low-liquidity markets, since the impact of news is reflected gradually, a smaller λ value may be appropriate. It is also possible to use different decay rate parameters for each time window. For example, by using a larger λ value for short-term windows (1 day, 3 days) and a smaller λ value for long-term windows (7 days, 30 days), short-term reactions and long-term trends can be appropriately captured. Step 4: Calculate the average of the weighted news vectors for each time window. Specifically, for n news vectors v_1, v_2, ..., v_n within the time window t and their corresponding weights w_1, w_2, ..., w_n, the aggregated vector a_t is calculated by the following formula: a_t = (w_1 * v_1 + w_2 * v_2 + ... + w_n * v_n) / (w_1 + w_2 + ... + w_n) Here, * represents scalar multiplication, and + represents vector addition. The denominator (w_1 + w_2 + ... + w_n) is a normalization coefficient, which makes the sum of the weights equal to 1. This ensures that the scale of the aggregated vector remains constant regardless of the number of news items. For example, if there are three news vectors v_1, v_2, v_3 within the time window t and the corresponding weights are w_1 = 0.9, w_2 = 0.7, w_3 = 0.5, the aggregated vector a_t is calculated as follows: a_t = (0.9 * v_1 + 0.7 * v_2 + 0.5 * v_3) / (0.9 + 0.7 + 0.5) = (0.9 * v_1 + 0.7 * v_2 + 0.5 * v_3) / 2.1 If there is no news within a specific time window (for example, if there is no news related to that asset on a specific day), the aggregated vector for that time window is set to a zero vector or a learnable default vector. Step 5: Concatenate the aggregated vectors of each time window to generate the final hierarchical news representation. For example, if the aggregated vectors for the 1-day window, 3-day window, 7-day window, and 30-day window are a_1, a_3, a_7, a_30 respectively, the final hierarchical news representation h is concatenated as follows: h = [a_1, a_3, a_7, a_30] Here, [a, b] represents the concatenation of vectors a and b. For example, when each aggregated vector is 768-dimensional, the final hierarchical news representation h will be 768 * 4 = 3072 dimensions. With this hierarchical approach, both short-term market reactions (1-day window) and long-term market sentiment (30-day window) can be captured, enabling a more comprehensive evaluation of the impact of news. For example, the aggregated vector of the short-term window can capture the impact of recent news events (e.g., earnings announcements, product launches), and the aggregated vector of the long-term window can capture long-term impacts such as industry trends, regulatory changes, and macroeconomic factors. Furthermore, not only can the aggregated vectors of each time window be simply concatenated, but it is also possible to dynamically weight them using an attention mechanism. In this case, the importance of different time windows can be learned according to market conditions and the characteristics of the assets. For example, in a high-volatility market, the importance of the short-term window increases, and in a low-volatility market, the importance of the long-term window increases, enabling adaptive behavior.

[0038] Relevance-Weighted News Aggregation Mechanism The relevance-weighted news aggregation mechanism aims to evaluate the relevance of each news article to the target asset and give greater influence to more relevant news. In the financial market, not all news has the same relevance to the target asset. For example, the news "Company A launches a new product" is likely to have a direct impact on the stock price of Company A, while the news "The growth rate of the entire industry is slowing down" may only have an indirect impact. Also, among news related to the same company, the degree of relevance to that company varies depending on the way it is mentioned in the title or the text. To consider such differences in relevance, the relevance-weighted news aggregation mechanism evaluates the relevance of each news article to the target asset and gives greater influence to more relevant news. The relevance score is calculated based on the following elements: 1. Presence of company name / ticker in the title: If the title of a news article contains a company name or ticker symbol, the news is likely to be directly related to the target company. For example, a title like "Toyota Motor Corporation Announces New EV" is likely to have a direct impact on the stock price of Toyota Motor Corporation (7203). Therefore, a high score is given when a company name or ticker is present in the title. 2. Frequency of keyword occurrences in the text: If keywords related to the target company (such as company name, product name, service name, major business field, etc.) frequently appear in the news article, the news is likely to be related to the target company. For example, in news about the "smartphone market", if keywords such as "Apple", "iPhone", "iOS", etc. frequently appear, the news may affect the stock price of Apple Inc. (AAPL). The score is calculated based on the frequency of keyword occurrences, but a decreasing function is used instead of a simple linear relationship to prevent over-weighting. 3. Location of company mention: If a company is mentioned at the beginning of a news article, it is likely that the company is the main subject of the news. For example, if a company is mentioned in the first paragraph of the article, it is likely that the company is the main topic of the news. On the other hand, if it is mentioned later in the article, it may be mentioned as a related company rather than the main topic. Therefore, an additional score is given when a company is mentioned at the beginning of the article (e.g., the first paragraph). The relevance weighted news aggregation process is executed in the following steps: Step 1: Calculate the relevance score for each news article. Specifically, the following formula is used: relevance_score = title_score + keyword_score + position_score Here, - title_score: 3.0 if the company name or ticker exists in the title, 0.0 otherwise - keyword_score: min(2.0, 0.5 * keyword_count), where keyword_count is the number of occurrences of relevant keywords in the article - position_score: 1.0 if the company name or ticker exists in the first paragraph of the article, 0.0 otherwise For example, if the company name is included in the title of the news article (title_score = 3.0), the relevant keyword appears 5 times in the text (keyword_score = min(2.0, 0.5 * 5) = 2.0), and the company name is included in the first paragraph (position_score = 1.0), the relevance score is 3.0 + 2.0 + 1.0 = 6.0. It is also possible to use more sophisticated methods for calculating the relevance score. For example, methods such as calculating the semantic similarity between the news article and the company profile using natural language processing techniques, or training a model to predict the relevance score using supervised learning. Step 2: Assign weights to the news vectors based on the relevance score. Specifically, for the news vector v and the relevance score r, the weighted vector v' is calculated by the following formula: v' = (1 + r) * v Here, (1 + r) is the weight coefficient based on the relevance score, and the higher the relevance score, the larger the value. For example, when the relevance score r = 6.0, the weight coefficient is (1 + 6.0) = 7.0, and the news vector v is expanded 7 times. In this weighting method, even when the relevance score is 0, the weight coefficient becomes 1 and the news vector is retained as it is. As a result, even news with low relevance is not completely ignored and can have a certain influence. Instead of weighting based on the relevance score, a method of normalizing the relevance score through the softmax function can also be considered. In this case, since the sum of the weights of all news becomes 1, the scale of the aggregated vector is kept constant regardless of the number of news. Step 3: Aggregate the weighted news vectors to generate the final news representation. Specifically, for n weighted news vectors v'_1, v'_2, ..., v'_n and the corresponding relevance scores r_1, r_2, ..., r_n, the aggregated vector a is calculated by the following formula: a = (v'_1 + v'_2 +... + v'_n) / n Or, when using the weighted average normalized by the relevance score: a = ((1 + r_1) * v'_1 + (1 + r_2) * v'_2 +... + (1 + r_n) * v'_n) / ((1 + r_1) + (1 + r_2) +... + (1 + r_n)) With this relevance weighting approach, news directly related to the target asset has a greater influence, and the impact of noisy news can be reduced. For example, news such as "Company A announces a new product" has a high relevance score and will have a greater influence. On the other hand, news such as "The growth rate of the entire industry is slowing down" has a low relevance score and its influence is limited. Furthermore, it is also possible to filter news based on the distribution of relevance scores. For example, by excluding news with a relevance score below a certain threshold (e.g., 2.0), noisy news can be completely eliminated. However, such hard filtering needs to be applied carefully because it may also exclude useful information contained in low-relevance news.

[0039] Sector-level correlation integration mechanism The sector-level correlation integration mechanism aims to improve the robustness of predictions by leveraging the correlation relationships among assets within the same sector. In the financial market, companies within the same sector often have similar business models and are affected by similar market factors, so there is often a strong correlation in price movements. For example, companies in the technology sector (such as Apple, Microsoft, Google, etc.) are often affected by common factors such as technology trends, regulatory changes, and consumer demand, so there is often a correlation in the movement of their stock prices. Similarly, similar correlation relationships are observed in other sectors such as the financial sector, energy sector, and healthcare sector. By leveraging such sector-level correlation relationships in predictions, the prediction accuracy of individual assets can be improved. For example, if it is predicted that the stock prices of multiple companies within a certain sector will rise, it is highly likely that the stock prices of other companies within that sector will also rise. The sector-level correlation integration process is executed in the following steps: Step 1: Identify the sector of the target asset. For example, if Company A belongs to the "technology" sector, identify other companies that belong to the same "technology" sector. Sector information is determined based on the classification of stock exchanges, GICS (Global Industry Classification Standard), or an in-house classification standard. For example, in the case of the Japanese market, classifications such as the 33-industry classification of the Tokyo Stock Exchange and the TOPIX 17 series can be used. In the case of the US market, the GICS classification of the S&P 500 (11 sectors such as Information Technology, Finance, and Healthcare) is widely used. Step 2: Calculate the correlation coefficient with other assets within the same sector. The correlation coefficient is calculated using past price data. Specifically, the Pearson correlation coefficient is calculated using the time-series data of the daily returns of the target asset and other assets. The Pearson correlation coefficient is an index that measures the strength of the linear correlation between two variables and takes values in the range from -1 to 1. The closer it is to 1, the stronger the positive correlation; the closer it is to -1, the stronger the negative correlation. When it is close to 0, it indicates a weak correlation. For the calculation of the correlation coefficient, data for a certain past period (e.g., 60 business days, 120 business days, 250 business days, etc.) is used. If the period is too short, it is easily affected by temporary correlation relationships, and if it is too long, there is a possibility of not capturing structural changes. Therefore, it is important to select an appropriate period. Also, instead of the simple Pearson correlation coefficient, it is also possible to use more robust correlation indicators (e.g., Spearman's rank correlation coefficient, Kendall's tau coefficient) or dynamic correlation models that capture time-varying correlation structures (e.g., Dynamic Conditional Correlation (DCC) model). Step 3: Identify assets with significant correlation relationships. Specifically, identify assets whose absolute value of the correlation coefficient exceeds a predetermined threshold (e.g., 0.3). This allows only assets with a strong correlation relationship with the target asset to be considered. It is important to adjust the selection of the threshold according to the characteristics of the market and the sector. Generally, thresholds in the range of 0.3 to 0.5 are often used. If the threshold is too low, weak correlation relationships are also considered and noise increases, and if it is too high, only strong correlation relationships are considered and information may be lost. In addition to the absolute value of the correlation coefficient, it is also important to consider the statistical significance (p-value) of the correlation. For example, even if the absolute value of the correlation coefficient exceeds 0.3, if the p-value exceeds 0.05 (not statistically significant), it may be determined not to consider that correlation. Step 4: Calculate the weight adjustment based on the correlation coefficient. Specifically, adjustments are made to the predicted return of the target asset based on the predicted returns of the correlated assets. The adjustment amount is proportional to the strength of the correlation coefficient, and the stronger the correlation, the greater the adjustment. The adjustment amount adjustment is calculated by the following formula: adjustment = sum(correlation_i * predicted_return_i * weight_i) / sum(weight_i) Here, - correlation_i: The correlation coefficient between the target asset and correlated asset i - predicted_return_i: The predicted return of correlated asset i - weight_i: The weight of correlated asset i (usually the absolute value of the correlation coefficient) - sum(): The sum for all correlated assets For example, if target asset A has correlation coefficients of 0.7, 0.5, and 0.3 with three assets B, C, and D within the sector, and the predicted returns of these assets are +2%, +1%, and -1% respectively, the adjustment amount is calculated as follows: adjustment = (0.7 * 2% * 0.7 + 0.5 * 1% * 0.5 + 0.3 * (-1%) * 0.3) / (0.7 * 0.7 + 0.5 * 0.5 + 0.3 * 0.3) = (0.98% + 0.25% - 0.09%) / (0.49 + 0.25 + 0.09) = 1.14% / 0.83 = 1.37% In this example, it is suggested that the predicted return of the target asset be adjusted by +1.37% based on the predicted returns of other assets within the sector. Step 5: Apply the adjustment to the original prediction to generate the final prediction. Specifically, the following formula is used: adjusted_return = original_return + dampening_factor * adjustment Here, - original_return: The original predicted return of the target asset - adjustment: The amount of adjustment based on the predicted return of the correlated asset - dampening_factor: The dampening factor to control the impact of the adjustment (usually in the range of 0.1 to 0.5) The dampening factor is a parameter used to control the impact of the adjustment, and values in the range of 0.1 to 0.5 are usually used. The larger the dampening factor, the stronger the impact of the sector-level correlation relationship, and the smaller the impact. For example, if the original predicted return of the target asset is original_return = +0.5%, the adjustment amount is adjustment = +1.37%, and the dampening factor is dampening_factor = 0.3, the adjusted predicted return is calculated as follows: adjusted_return = 0.5% + 0.3 * 1.37% = 0.5% + 0.411% = 0.911% In this example, by considering the sector-level correlation relationship, the predicted return is adjusted from +0.5% to +0.911%. It is important to adjust the attenuation coefficient according to the characteristics of the market and the sector. For example, when the correlation within the sector is strong, it is conceivable to use a large attenuation coefficient, and when the correlation is weak, a small attenuation coefficient is used. Moreover, it is also possible to adjust the attenuation coefficient dynamically according to the market situation and the characteristics of the sector instead of using a fixed value. For example, when the market volatility is high or the correlation within the sector is particularly strong, an adjustment such as using a large attenuation coefficient is considered. This sector-level correlation integration approach enables prediction not only of individual assets but also of trends across the entire sector, improving the robustness of the prediction. For example, even when there is little news about individual assets, predictions can be made based on the trends of the entire sector. Also, the impact of outliers and temporary noise can be reduced.

[0040] Adaptive Modality Fusion Mechanism The purpose of the adaptive modality fusion mechanism is to dynamically adjust the importance of each modality (time series data and news data) according to the market situation. The relative importance of time series data and news data changes depending on the state and conditions of the market. For example, in a high-volatility market, short-term price fluctuations become important and the weight of time series data increases, while when important news events occur, the weight of news data increases. Conventional fixed weighting methods cannot cope with such changes in market conditions and may lead to a decrease in prediction accuracy. The adaptive modality fusion mechanism realizes more accurate prediction by dynamically adjusting the weights of each modality according to the market situation. The adaptive modality fusion process is executed in the following steps: Step 1: Evaluate the volatility of the market. Specifically, the volatility level is calculated using indicators such as the VIX (Volatility Index) or the standard deviation of past price movements. The volatility level is normalized to the range from 0 to 1. For calculating volatility, it is common to use the standard deviation of daily returns over a certain past period (e.g., 10 business days, 20 business days, 30 business days, etc.). For example, by calculating the standard deviation of daily returns over the past 20 business days and converting it to a percentile rank based on the distribution over the past year, a volatility level in the range from 0 to 1 can be obtained. Also, when using a market-wide volatility indicator such as the VIX, a volatility level in the range from 0 to 1 can be obtained by normalizing its value based on the past distribution. For example, conversions such as low volatility (0.0 to 0.3) when the VIX is less than 15, medium volatility (0.3 to 0.7) when it is between 15 and 25, and high volatility (0.7 to 1.0) when it exceeds 25 can be considered. Step 2: Evaluate the quantity and importance of news. Specifically, the following indicators are calculated: - News quantity: A value normalized by comparing the number of news articles related to the target asset with the past average (range from 0 to 1) - News importance: A value normalized by taking the average of the relevance scores of news articles (range from 0 to 1) For calculating the news quantity, the number of news articles related to the target asset published over a certain past period (e.g., 5 business days, 10 business days, etc.) is used. By converting this value to a percentile rank based on the past distribution (e.g., the distribution of the number of news articles over the past year), a news quantity indicator in the range from 0 to 1 can be obtained. For example, if there are 10 news articles related to the target asset published in the past 5 business days and it corresponds to the 80th percentile in the distribution of the number of news articles per 5 business days over the past year, the news quantity indicator is 0.8. To calculate the news importance, the average value of the relevance scores of news articles related to the target asset published in a certain period in the past is used. By converting this value into a percentile rank based on the past distribution, a news importance indicator in the range of 0 to 1 can be obtained. For example, if the average value of the relevance scores of news articles related to the target asset published in the past 5 business days is 4.5 and it corresponds to the 70th percentile in the distribution of the average relevance scores in the past year, the news importance indicator will be 0.7. Step 3: Adjust the basic weights of each modality based on the market situation. Specifically, the following rules are applied: (a) In the case of a high-volatility market (volatility level > 0.8): - Weight of time-series data: 0.7 - Weight of news data: 0.3 (b) When an important news event is occurring (news importance > 0.7 and news volume > 0.6): - Weight of time-series data: 0.3 - Weight of news data: 0.7 (c) In the case of normal market conditions: - Weight of time-series data: 0.5 - Weight of news data: 0.5 These rules can be adjusted according to the characteristics of the market and the asset. For example, in a specific sector or asset class, the impact of news may be stronger, or the time-series pattern may be more important. Also, instead of discrete rules, it is possible to calculate the weights using a continuous function. For example, a function like the following can be used: ts_weight = 0.5 + 0.3 * volatility - 0.3 * news_importance news_weight = 1 - ts_weight Here, volatility is the volatility level (from 0 to 1), and news_importance is the news importance index (from 0 to 1). In this function, the higher the volatility, the more the weight of the time series data increases, and the higher the news importance, the more the weight of the time series data decreases (i.e., the weight of the news data increases). Step 4: Combine the time series prediction and the news prediction using the adjusted weights. Specifically, use the following formula: combined_prediction = (ts_weight * time_series_prediction + news_weight * news_prediction) / (ts_weight + news_weight) Here,[[]] - ts_weight: The weight of the time series data - time_series_prediction: The prediction based on the time series data - news_weight: The weight of the news data - news_prediction: The prediction based on the news data The denominator (ts_weight + news_weight) is a normalization factor to ensure that the sum of the weights is 1. This keeps the scale of the combined prediction constant regardless of the weight values. For example, if the prediction based on the time series data is time_series_prediction = +1.2%, the prediction based on the news data is news_prediction = +0.5%, the weight of the time series data is ts_weight = 0.7, and the weight of the news data is news_weight = 0.3, the combined prediction is calculated as follows: combined_prediction = (0.7 * 1.2% + 0.3 * 0.5%) / (0.7 + 0.3) = (0.84% + 0.15%) / 1.0 = 0.99% In this example, assuming a high-volatility market, the weight of time-series data is large, and the combined prediction is closer to the time-series prediction value. With this adaptive modality fusion approach, the optimal modality combination is selected according to the market situation, and the prediction accuracy is improved. For example, in a high-volatility market, time-series data becomes more important, and when important news events occur, news data becomes more important. Furthermore, this adaptive fusion mechanism can be extended not only to the fusion of two modalities (time series and news), but also to the fusion of more modalities (e.g., time series, news, social media, macroeconomic indicators, etc.). In this case, by dynamically adjusting the importance of each modality according to the market situation, a more comprehensive prediction becomes possible.

[0041] Overall processing flow of the system

[0153] The overall processing flow of the system of the present invention is as follows:

[0154] 1. Collection of input data - Time-series data of asset prices (candlestick data: opening price, closing price, high price, low price) - Trading volume data - Related news articles - Sector information - Market indicators (volatility indicators such as VIX)

[0155] Time series data is collected from the API of a stock exchange, a financial data provider, or a public data source. The data is usually provided at a daily, hourly, or minute-level granularity. Although this invention mainly uses daily data, it is also applicable to higher-frequency data (e.g., hourly, minute-level).

[0156] News articles are collected from various sources such as financial news websites, corporate press releases, social media, and experts' blogs. News articles contain information such as the publication date and time, title, text, author, source, etc.

[0157] Sector information is collected based on the classification of a stock exchange, GICS (Global Industry Classification Standard), or an independent classification standard. Sector information includes corporate sector classification, sub-sector classification, industry classification, etc.

[0158] Market indicators are indicators that show the overall state of the market, such as the VIX (Volatility Index), interest rates, exchange rates, commodity prices, etc. These indicators are used for evaluating market conditions and adjusting adaptive modality fusion.

[0159] 2. Data preprocessing - Normalization of time series data: Normalize each feature to have a mean of 0 and a standard deviation of 1 - Completion of missing values: Use methods such as linear interpolation or forward filling to complete missing values - Handling of outliers: Use methods such as moving average or median filtering to handle outliers - Cleaning of news articles: Removal of HTML elements, normalization of special characters, removal of stop words, etc. - Tokenization of news articles: Tokenize the text using the tokenizer of a language model (such as BERT, RoBERTa, etc.) - Vectorization of news articles: Convert news articles into vector representations using a pre-trained language model In the preprocessing of time series data, first, the missing values are filled. The missing values are filled using linear interpolation, forward fill, backward fill, or machine learning-based filling methods (e.g., k-nearest neighbor method, random forest). Next, the outliers are processed. The outliers are identified and processed using moving average, median filtering, or statistical methods (e.g., outlier detection based on Z-score, modified interquartile range (IQR) method). Next, the scaling of each feature is performed. There are methods such as standardization (mean 0, standard deviation 1), min-max scaling (from 0 to 1), and robust scaling (based on median and interquartile range). In the present invention, standardization is mainly used, but other scaling methods can also be used according to the characteristics of the data. Finally, feature extraction is performed. From the candlestick data, features such as relative price change (return), volatility (e.g., standard deviation of returns over a certain past period), momentum (e.g., cumulative value of returns over a certain past period), and technical indicators (e.g., moving average, RSI, MACD) are extracted. These features are used as the input to the time series processing unit. In the preprocessing of news articles, first, cleaning is performed. Cleaning includes removing HTML elements, normalizing special characters, removing stop words, and processing punctuation. Next, tokenization is performed. For tokenization, a tokenizer of a language model (such as BERT, RoBERTa) is used. The tokenizer splits the text into words or subwords and converts them into integer IDs (token IDs) corresponding to each token. Next, vectorization is performed. For vectorization, a pre-trained language model (such as BERT, RoBERTa) is used. The language model receives the tokenized text as input and outputs the hidden state corresponding to each token. To obtain the representation of the entire article, the final layer hidden state of the [CLS] token (a special token added at the beginning of the sentence) may be used, or the average of the hidden states of all tokens may be taken. Finally, filtering is performed. In filtering, keyword matching, entity recognition, or a dedicated classifier is used to select news articles related to the target asset. In keyword matching, it is checked whether keywords such as company names, ticker symbols, product names, and service names are included in the news articles. In entity recognition, natural language processing techniques are used to identify entities such as company names and organization names included in the news articles. In a dedicated classifier, a machine learning model (e.g., logistic regression, random forest, neural network) is used to classify whether a news article is related to the target asset. 3. Processing of Time-Series Data - Feature extraction: Calculate relative price changes (returns) from candlestick data - Application of time window: Form an input vector using the return vector of the past 5 business days - Application of time-series prediction model: Predict the return of the next business day using a deep learning model such as LSTM In the processing of time-series data, first, feature extraction is performed. Relative price changes (returns) are calculated from candlestick data. Returns can be calculated as continuous returns (log returns) or discrete returns (simple returns). Continuous returns are calculated as ln(P_t / P_{t-1}), and discrete returns are calculated as (P_t / P_{t-1}) - 1. Here, P_t is the price at time t (usually the closing price). In the present invention, mainly discrete returns are used, but continuous returns can also be used depending on the characteristics of the data. Next, the application of the time window is performed. An input vector is formed using the return vector for a certain period in the past (e.g., 5 business days, 10 business days, etc.). For example, when using the return vector of the past 5 business days, the input vector will be 5-dimensional. Also, when calculating returns for each of the four features (opening price, closing price, high price, low price) of the candlestick data, the input vector will be 5 * 4 = 20-dimensional. Finally, apply the time series prediction model. Use a deep learning model such as LSTM to predict the return of the next business day. The LSTM model receives an input vector, learns the temporal dependencies through hidden states, and predicts the return of the next business day. The structure of the LSTM model is as follows: - Input layer: Receives an input vector (e.g., 20 dimensions) - LSTM layer 1: 128 units, tanh as the activation function, dropout rate 0.2 - LSTM layer 2: 64 units, tanh as the activation function, dropout rate 0.2 - Fully connected layer: 32 units, ReLU as the activation function - Output layer: 1 unit (prediction of the return of the next business day) 4. Processing of news data - Filtering of news articles: Select news articles related to the target asset - Hierarchical temporal news aggregation: Aggregation of news data in multiple time windows - Relevance-weighted news aggregation: Weighting based on the relevance of news - Application of the news prediction model: Predict the return of the next business day using the processed news vector In the processing of news data, first filter the news articles. To select news articles related to the target asset, use keyword matching, entity recognition, or a dedicated classifier. The filtered news articles are sent to the hierarchical temporal news aggregation mechanism. Next, perform hierarchical temporal news aggregation. Use multiple time windows (e.g., 1 day, 3 days, 7 days, 30 days) to aggregate the news vectors within each time window. For the news vectors within each time window, calculate a weighted average with an exponentially decaying weight based on temporal proximity. The aggregated vectors are sent to the relevance-weighted news aggregation mechanism. Next, perform relevance-weighted news aggregation. Evaluate the relevance of news articles to the target assets and give greater influence to more relevant news. The relevance score is calculated based on factors such as the presence of company names / tickers in the title, the frequency of keyword occurrences in the text, and the position of company mentions. Finally, apply the news prediction model. Use the processed news vectors to predict the return on the next business day. The structure of the news prediction model is as follows: - Input layer: Receive the aggregated news vectors (e.g., 3072 dimensions) - Fully connected layer 1: 512 units, ReLU as the activation function, dropout rate 0.3 - Fully connected layer 2: 128 units, ReLU as the activation function, dropout rate 0.3 - Fully connected layer 3: 32 units, ReLU as the activation function - Output layer: 1 unit (prediction of the return on the next business day) 5. Integration of Predictions - Sector-level correlation integration: Adjust the prediction based on the correlation within the sector - Adaptive modality fusion: Generate the final prediction based on the market situation In the integration of predictions, first perform sector-level correlation integration. Adjust the prediction based on the correlation within the sector. Calculate the correlation coefficient between the target asset and other assets in the same sector, and adjust the prediction of the target asset based on the predictions of assets with significant correlation. Next, perform adaptive modality fusion. Dynamically adjust the weights of each modality (time series data and news data) according to the market situation. Determine the weights of time series data and news data based on market volatility and the importance of news, and generate the final prediction. 6. Output - Predicted price: Prediction of the closing price on the next business day - Price movement direction: Prediction of increase or decrease - Prediction confidence: An indicator showing the confidence level of the prediction - Explanation information: The main factors contributing to the prediction (important news, trends of correlated assets, etc.) In the output, first, the predicted price is calculated. To calculate the predicted price from the predicted return, the following formula is used: predicted_price = current_price * (1 + predicted_return) Here, current_price is the current price (usually the latest closing price), and predicted_return is the predicted return. Next, the price movement direction is determined. If the predicted return is positive, it is predicted to rise; if negative, it is predicted to fall. Next, the confidence level of the prediction is calculated. The confidence level of the prediction is calculated based on the probability distribution of the output of the prediction model (e.g., softmax output) or the distribution of the prediction error. For example, based on the distribution of past prediction errors, the confidence interval of the predicted value can be calculated. Finally, explanation information is generated. Identify the main factors contributing to the prediction (important news, trends of correlated assets, etc.) and provide them to the user. For example, information such as "This prediction is mainly strongly influenced by short-term news (1-day window)" or "This prediction is based on the trends of correlated assets within the sector" can be provided.

[0042] Details of implementation The system of the present invention can be implemented using the Python language. As the main libraries, TensorFlow / Keras or PyTorch are used for constructing deep learning models, Transformers are used for utilizing language models, NumPy and Pandas are used for data processing, and Scipy is used for statistical calculations. The LSTM model of the time series processing unit can be implemented with the following Keras code: ```python from tensorflow.keras.models import Sequential from tensorflow.keras.layers import LSTM, Dense, Dropout def create_time_series_model(input_shape): model = Sequential() model.add(LSTM(128, activation='tanh', return_sequences=True, input_shape=input_shape)) model.add(Dropout(0.2)) model.add(LSTM(64, activation='tanh')) model.add(Dropout(0.2)) model.add(Dense(32, activation='relu')) model.add(Dense(1)) model.compile(optimizer='adam', loss='mse') return model # Input shape: (number of time steps, number of features) # Example: 4 features (open, close, high, low) for 5 days time_series_model = create_time_series_model((5, 4)) ``` As the language model for the news processing part, pre-trained models such as BERT and RoBERTa are used. It can be implemented as follows using the Transformers library: ```python from transformers import BertModel, BertTokenizer import torch # Loading pre-trained model and tokenizer tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') bert_model = BertModel.from_pretrained('bert-base-uncased') def get_bert_embedding(text, max_length=512): # Tokenize the text inputs = tokenizer(text, return_tensors='pt', max_length=max_length, padding='max_length', truncation=True) # Apply the BERT model with torch.no_grad(): outputs = bert_model(**inputs) # Use the last hidden state of the [CLS] token embedding = outputs.last_hidden_state[:, 0, :].numpy() return embedding ``` The hierarchical temporal news aggregator can be implemented as a Python class as follows: ```python import numpy as np from datetime import datetime, timedelta class HierarchicalTemporalNewsAggregator: def __init__(self, time_windows=[1, 3, 7, 30], decay_rate=0.2): self.time_windows = time_windows self.decay_rate = decay_rate def aggregate(self, news_vectors, news_timestamps, reference_date): aggregated_vectors = [] for window in self.time_windows: # Select news within the time window window_start = reference_date - timedelta(days=window) window_vectors = [] window_weights = [] for vector, timestamp in zip(news_vectors, news_timestamps): if window_start <= timestamp <= reference_date: # Calculate the weight based on the time difference days_diff = (reference_date - timestamp).days weight = np.exp(-self.decay_rate * days_diff) window_vectors.append(vector) window_weights.append(weight) if window_vectors: # Calculate the weighted average window_vectors = np.array(window_vectors) window_weights = np.array(window_weights) weighted_sum = np.sum(window_vectors * window_weights[:, np.newaxis], axis=0) normalized_vector = weighted_sum / np.sum(window_weights) aggregated_vectors.append(normalized_vector) else: # If there is no news, use a zero vector aggregated_vectors.append(np.zeros(news_vectors[0].shape)) # Concatenate the aggregated vectors for each time window return np.concatenate(aggregated_vectors) ``` The relevance-weighted news aggregation mechanism can be implemented as the following Python class: ```python class RelevanceWeightedNewsAggregator: def __init__(self): pass def compute_relevance_score(self, news_item, company_name, ticker, keywords): # Presence of company name / ticker in the title title = news_item.get('title', '').lower() title_score = 3.0 if company_name.lower() in title or ticker.lower() in title else 0.0 # Frequency of keyword occurrences in the text body = news_item.get('body', '').lower() keyword_count = sum(1 for kw in keywords if kw.lower() in body) keyword_score = min(2.0, 0.5 * keyword_count) # Position of company mention first_paragraph = body.split('\n')[0] if '\n' in body else body[:500] position_score = 1.0 if company_name.lower() in first_paragraph or ticker.lower() in first_paragraph else 0.0 return title_score + keyword_score + position_score def aggregate(self, news_vectors, news_items, company_name, ticker, keywords): if not news_vectors: return np.zeros(768) # Empty vector # Calculation of relevance score relevance_scores = [self.compute_relevance_score(item, company_name, ticker, keywords) for item in news_items] # Weighting Based on Relevance weighted_vectors = [(1 + score) * vector for score, vector in zip(relevance_scores, news_vectors)] # Calculation of Weighted Average weights = [1 + score for score in relevance_scores] weighted_sum = np.sum([w * v for w, v in zip(weights, weighted_vectors)], axis=0) normalized_vector = weighted_sum / np.sum(weights) return normalized_vector ```

[0043] The sector-level correlation integration mechanism can be implemented as a Python class as follows: ```python class SectorLevelCorrelationIntegrator: def __init__(self, correlation_threshold=0.3, dampening_factor=0.3): self.correlation_threshold = correlation_threshold self.dampening_factor = dampening_factor def integrate(self, original_return, sector_returns, correlation_matrix, ticker): # Identify assets with significant correlation correlated_assets = [] for other_ticker, corr in correlation_matrix[ticker].items(): if abs(corr) >= self.correlation_threshold and other_ticker != ticker: correlated_assets.append((other_ticker, corr)) if not correlated_assets: return original_return # Calculate weight adjustment based on correlation coefficient adjustment = 0.0 total_weight = 0.0 for other_ticker, corr in correlated_assets: weight = abs(corr) adjustment += corr * sector_returns[other_ticker] * weight total_weight += weight adjustment / = total_weight # Apply the adjustment to the original prediction adjusted_return = original_return + self.dampening_factor * adjustment return adjusted_return ``` The adaptive modality fusion mechanism can be implemented as a Python class as follows: ```python class AdaptiveModalityFusion: def __init__(self): pass def fuse(self, time_series_prediction, news_prediction, volatility, news_importance, news_volume): # Basic weights ts_weight = 0.5 news_weight = 0.5 # Adjust weights based on market conditions if volatility > 0.8: # High volatility market ts_weight = 0.7 news_weight = 0.3 elif news_importance > 0.7 and news_volume > 0.6: # Important news event ts_weight = 0.3 news_weight = 0.7 # Weight normalization total_weight = ts_weight + news_weight ts_weight / = total_weight news_weight / = total_weight # Fusion of predictions fused_prediction = ts_weight * time_series_prediction + news_weight * news_prediction return fused_prediction, {'ts_weight': ts_weight, 'news_weight': news_weight} ``` The news prediction model can be implemented with the following Keras code: ```python def create_news_model(input_dim): model = Sequential() model.add(Dense(512, activation='relu', input_dim=input_dim)) model.add(Dropout(0.3)) model.add(Dense(128, activation='relu')) model.add(Dropout(0.3)) model.add(Dense(32, activation='relu')) model.add(Dense(1)) model.compile(optimizer='adam', loss='mse') return model # Input dimension: Dimension of the vector after hierarchical temporal news aggregation # Example: 4 time windows x 768 dimensions (BERT-base) news_model = create_news_model(4 * 768) ``` The entire system is implemented by combining these components. Each component can be developed, tested, and updated independently. Also, when new language models or time series models are developed, it can be easily upgraded by replacing the corresponding components.

Industrial Applicability

[0044] The system of the present invention can be used in the following application examples:

[0045] 1. Investment decision-making support system: It can be used as a decision-making support tool for individual investors and fund managers when making investment decisions. The system provides the price predictions of multiple assets along with the main factors (such as important news and trends of correlated assets) that form the basis of the predictions, thus assisting in making more information-based investment decisions. For example, when an investor is considering investing in a specific stock, the system provides the predicted price, price movement direction, and prediction confidence level for the next business day. Additionally, by providing explanatory information such as "This prediction is strongly influenced by short-term news (1-day window)" or "This prediction is based on the trends of correlated assets within the sector", investors can understand the basis of the prediction and make investment decisions with more confidence.

[0046] 2. Algorithm Trading System: By integrating it as part of an automated trading system, a trading strategy based on multimodal data can be realized. The system can generate buy and sell signals based on the prediction of price movement direction and automatically execute trades. For example, when the system predicts a price increase of a specific stock with a high confidence level, the automated trading system generates a signal to purchase that stock. Conversely, when a price decrease is predicted, a signal to sell the held stock is generated. Also, it is possible to adjust the trading volume based on the prediction confidence level. Adjustments such as setting a large trading volume for trades based on high-confidence predictions and a small trading volume for trades based on low-confidence predictions can be considered.

[0047] 3. Risk Management System: It can be used as a risk management system for financial institutions and hedge funds. The system can evaluate the risk of a portfolio based on the predicted price movements of multiple assets and propose countermeasures for risk reduction. For example, if the system predicts a decline in the price of stocks in a particular sector, it can propose reducing the exposure to that sector. Also, by considering the correlation between assets within the portfolio, it can make proposals to maximize the effect of diversification. Furthermore, it can also be utilized for generating stress test scenarios. For example, scenarios such as "price movements in the event of an important news event" or "price movements in a high-volatility market" can be generated to evaluate the risk of the portfolio.

[0048] 4. Market Monitoring System: It can be used as a market monitoring system for regulatory authorities and exchanges. The system can monitor the deviation between the predicted price and the actual price in order to detect abnormal price movements and the possibility of market manipulation. For example, if the system detects that the price of a particular stock deviates significantly from the model's prediction, it may indicate the possibility of market manipulation or trading based on undisclosed information. By detecting such abnormalities, regulatory authorities can initiate an investigation and maintain the fairness and transparency of the market.

[0049] 5. Economic Analysis System: It can be used as a tool for economists and policymakers to analyze economic trends. The system can provide insights into the overall health and future trends of the economy through predicting the trends in the financial market. For example, the system can predict the stock price indices of multiple sectors and estimate the overall economic growth rate and the stage of the business cycle based on those predictions. Also, by predicting the impact of specific policy changes (e.g., interest rate changes, regulatory changes) on the financial market, policymakers can evaluate the effects of policies in advance.

[0050] 6. Financial Education Tool: It can be used as a tool for investment education and improving financial literacy. By explaining the main factors underlying the predictions, the system can help users understand the trends and price formation mechanisms in the financial market. For example, by providing an explanation such as "This stock price prediction is mainly based on the recent quarterly earnings announcements (news factor) and past momentum (time series factor)", users can understand the factors affecting price formation. Also, by providing an explanation such as "The stock prices of this sector have a high correlation with each other, and news about one company also affects the stock prices of other companies", users can understand the interrelationships in the market.

Claims

1. 1. A multi-modal system for forecasting financial asset prices, comprising: a first processing unit for processing the time series data with an LSTM network to generate a price forecast; A second processing unit including a hierarchical time news aggregation mechanism that aggregates news data over a number of different time windows (1 day, 3 days, 7 days, 30 days) and a relevance weighting news aggregation mechanism that weights news based on their relevance; An integration unit including a sector-level correlation integration mechanism that adjusts forecasts based on sector-level correlations and an adaptive modality fusion mechanism that dynamically adjusts the weights of each modality according to market conditions; A multimodal system comprising:

2. A news data processing method for predicting financial asset prices, comprising the steps of: a hierarchical temporal news aggregation step for computing a weighted average of news vectors in a number of different time windows with exponential decay based on temporal proximity; a relevance weighting news aggregation step of calculating a relevance score based on the presence of company names or tickers in the titles of news articles, the frequency of keywords in the news articles, and the location of company mentions, and weighting the news vectors based on the relevance score; converting news articles into vector representations using pre-trained language models (BERT, RoBERTa, GPT) and concatenating aggregate vectors of different time windows to generate a final hierarchical news representation; A news data processing method comprising:

3. 1. An adaptive synthesis method for financial asset price forecasting, comprising: a sector-level correlation integration step of calculating correlation coefficients between the target asset and other assets in the same sector and adjusting the forecast of the target asset based on the forecast of assets whose absolute value of the correlation coefficient exceeds a predetermined threshold; an adaptive modality fusion step that dynamically adjusts the weights of the time series data and the news data based on market volatility, news volume, and news importance, increasing the weight of the time series data in high volatility markets and increasing the weight of the news data when important news events are occurring; Calculating the adjusted forecast according to the formula "adjusted_return = original_return + dampening_factor * adjustment"; 13. An adaptive integration method comprising:

Citation Information

Patent Citations

  • Information processing device, control method therefor, program, and learned model

    JP2021163073A

  • Methods and systems for predicting market behavior based on news and sentiment analysis

    US20130138577A1