Real-time investment decision-making system and method based on multi-modal fusion

Through the investment decision-making system of multimodal fusion and reinforcement learning, the problem of insufficient adaptability of the existing system in multi-source data processing and market status is solved, a comprehensive understanding of the market and dynamic risk control are achieved, and the accuracy and robustness of decision-making are improved.

CN120707303APending Publication Date: 2025-09-26CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511101881.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing investment decision-making systems find it difficult to integrate multi-source heterogeneous data, lack dynamic adaptability to market conditions and intelligent risk control capabilities, resulting in large losses in extreme market environments.

Method used

It adopts a multimodal sentiment information fusion module, a market status judgment module, a reinforcement learning portfolio management module, a multi-strategy collaborative decision-making module and a risk adaptive control module. Through the weighted fusion of social media sentiment data, public opinion events and market structured data, combined with reinforcement learning and multi-strategy collaborative decision-making, it dynamically adjusts the investment portfolio configuration and risk control parameters.

Benefits of technology

It achieves a comprehensive understanding of multi-dimensional market information, enhances the system's adaptability to market changes, improves the accuracy and robustness of decision-making, and reduces the risk of loss under extreme market conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707303A_ABST
    Figure CN120707303A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of finance, in particular to a real-time investment decision-making system and method based on multi-modal fusion, and the method comprises the steps: integrating social media emotion data, public opinion event classification data and market structured data through a multi-modal emotion information fusion module, generating a comprehensive emotion representation vector, the market state judgment module judges market states based on historical market data and generates time-weighted market state representations, the reinforcement learning combination management module generates asset allocation decisions according to the market state representations and the time-weighted market state representations, and the multi-strategy collaborative decision module integrates reinforcement learning, decision trees and regression analysis strategies and generates collaborative decision signals. The risk self-adaptive control module predicts the market fluctuation rate, assesses the investment risk and adjusts decision signal execution parameters, the decision signal generation module generates a final investment decision signal and outputs the final investment decision signal to the transaction execution system, and the transaction execution system performs multi-modal information fusion and a time-sensitive attention mechanism. And the comprehensive understanding capability and the time sequence change adaptive capability of the market are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of financial technology, and in particular to a real-time investment decision-making system and method based on multimodal fusion, which are specifically applied to investment portfolio management, quantitative trading strategy formulation, and market risk control. Background Art

[0002] With the increasing globalization and informatization of financial markets, investment decision-making faces the dual challenges of processing multi-source, heterogeneous data and adapting to complex market environments. Traditional investment decision-making systems primarily rely on single data sources, such as price time series data or fundamental analysis, which makes it difficult to fully capture the multidimensional market information. Furthermore, existing systems often lack adaptability when faced with drastic market fluctuations, making it difficult to adjust decision-making strategies in a timely manner to address market fluctuations.

[0003] In existing technologies, most investment decision-making systems adopt a single modal analysis method, either focusing only on price trends or only analyzing text sentiment, and lacking comprehensive consideration of multi-source data; at the same time, existing systems mostly use static asset allocation models, which are difficult to adapt to dynamic changes in market conditions; in addition, existing systems usually use fixed parameter settings for risk control and lack adaptive capabilities, which leads to large losses in extreme market environments.

[0004] Therefore, there is an urgent need for a real-time investment decision-making system that can integrate multimodal data, dynamically adapt to changes in market conditions, and have intelligent risk control capabilities. Summary of the Invention

[0005] The purpose of the present invention is to provide a real-time investment decision-making system and method based on multimodal fusion, aiming to solve the technical problems existing in the existing investment decision-making system in multi-source heterogeneous data processing, market status adaptability and risk control.

[0006] The present invention proposes a real-time investment decision-making system based on multimodal fusion, comprising:

[0007] A multimodal sentiment information fusion module is used to obtain social media sentiment data, public opinion event classification data, and market structured data, and perform weighted fusion processing on the data to generate a comprehensive sentiment representation vector;

[0008] A market status judgment module, connected to the multimodal sentiment information fusion module, is used to extract features based on historical market data, judge whether the current market status is a low volatility market or a high volatility market, and generate a time-weighted market status representation based on the market status;

[0009] a reinforcement learning portfolio management module, connected to the multimodal sentiment information fusion module and the market status judgment module, configured to receive the comprehensive sentiment representation vector and the time-weighted market status representation, generate asset allocation decisions based on a reinforcement learning framework, and output an investment portfolio allocation ratio;

[0010] A multi-strategy collaborative decision-making module is connected to the reinforcement learning combination management module and is used to integrate the decision results of the reinforcement learning strategy, decision tree strategy and regression analysis strategy, dynamically calculate the weight of each strategy, and generate a collaborative decision signal;

[0011] A risk adaptive control module, connected to the multi-strategy collaborative decision-making module, is used to predict current market volatility, assess investment risk levels, and dynamically adjust the execution parameters of the collaborative decision-making signal according to the risk level;

[0012] A decision signal generation module is connected to the multi-strategy collaborative decision module and the risk adaptive control module, and is used to receive the collaborative decision signal and the adjusted execution parameters, generate a final investment decision signal, and output the final investment decision signal to the transaction execution system.

[0013] Preferably, the multimodal emotion information fusion module includes:

[0014] Sentiment data collection unit, used to obtain unstructured text data from social media, obtain event classification information from news and public opinion analysis, and extract key variable data from financial statements;

[0015] An emotion vector encoding unit, connected to the emotion data acquisition unit, is used to convert the unstructured text data into emotion vectors through a pre-trained model, convert the event classification information into a category vector, and normalize the key variable data to generate a numerical vector;

[0016] An emotional information classification and integration unit, connected to the emotional vector encoding unit, is used to collect relevant emotional vectors for each public opinion event, and weight multiple emotional vectors of the same event according to their importance to generate a fused emotional vector;

[0017] The timeliness adjustment unit is connected to the emotional information classification and integration unit, and is used to apply a time decay function to reduce the weight of outdated emotional information, set a weight enhancement mechanism for emergencies, and generate a timeliness-weighted comprehensive emotional representation vector.

[0018] Preferably, the market status judgment module includes:

[0019] Feature extraction unit, used to extract features from time series market data through CNN+LSTM network and generate market feature vectors;

[0020] a market classification unit, connected to the feature extraction unit, for calculating classification probability values ​​of a low volatility market and a high volatility market based on the market feature vector, and determining a current market state;

[0021] a time sensitivity judgment unit, connected to the market classification unit, for analyzing the time correlation strength of the market feature vector and judging whether the market feature is time sensitive or moment sensitive;

[0022] The feature updating unit is connected to the time sensitivity judgment unit and is used to adopt a progressive weighted updating method for time-sensitive features and a multi-dimensional fusion updating method for moment-sensitive features to generate a time-weighted market status representation.

[0023] Preferably, the reinforcement learning combination management module includes:

[0024] The state space construction unit is used to define the portfolio state space, including multi-dimensional features such as holding ratio, return performance and risk indicators, design the action space, and construct state transition rules;

[0025] The strategy network unit is connected to the state space construction unit and is used to generate asset allocation decisions through the actor network, evaluate the current market state value through the critic network, and provide decision quality feedback;

[0026] A dynamic configuration execution unit, connected to the strategy network unit, is used to generate an asset allocation decision sequence, calculate the expected reward and risk value of each decision point, and dynamically adjust the execution step size and execution frequency based on market conditions;

[0027] The parameter update optimization unit is connected to the dynamic configuration execution unit and is used to calculate the loss rate and cumulative reward value based on actual transaction feedback and update the Actor network and Critic network parameters.

[0028] Preferably, the multi-strategy collaborative decision-making module includes:

[0029] A multi-strategy generation unit, used to generate multiple investment strategy signals based on different algorithms, including reinforcement learning strategy signals, decision tree strategy signals, and regression analysis strategy signals;

[0030] a strategy evaluation unit, connected to the multi-strategy generation unit, for recording the historical performance of each strategy under different market conditions and calculating the applicability of each strategy to the current market condition;

[0031] A weight calculation unit, connected to the strategy evaluation unit, for dynamically allocating weights to each strategy through a logistic regression model to generate a weighted coefficient;

[0032] The decision fusion unit is connected to the multi-strategy generation unit and the weight calculation unit, and is used to integrate the multi-strategy decision signals based on the weight coefficients, eliminate conflicts between strategies, and generate collaborative decision signals.

[0033] Preferably, the risk adaptive control module includes:

[0034] Volatility prediction unit, used to predict future market volatility based on historical market data and current market conditions;

[0035] a risk assessment unit, connected to the volatility prediction unit, for comprehensively considering volatility, liquidity, and tail risk to generate a risk score;

[0036] a parameter adjustment unit, connected to the risk assessment unit, for dynamically adjusting decision execution parameters according to the risk score, including adjusting the execution step size, execution frequency, and position limit;

[0037] The risk hedging unit is connected to the parameter adjustment unit and is used to activate the automatic risk hedging mechanism in high-risk situations and generate a hedging strategy.

[0038] Preferably, the action space set by the state space construction unit includes the asset adjustment amplitude, adjustment direction and adjustment frequency; the reward function of the strategy network unit combines the three dimensions of yield, volatility and maximum drawdown; the dynamic configuration execution unit automatically reduces the decision step in a high volatility market and sets upper and lower limit constraints on the configuration ratio adjustment amplitude; the parameter update optimization unit introduces a long-term value factor to balance short-term returns and long-term stability.

[0039] Preferably, the feature update unit adopts a multi-head attention mechanism, including short-term attention heads, medium-term attention heads and long-term attention heads, which focus on market characteristics on time scales of 1-3 days, 5-20 days and 30-90 days respectively, and the attention weight is automatically adjusted based on the relevance of each time scale to the current decision.

[0040] Preferably, the system further includes a multi-modal cycle optimization module for:

[0041] De-noise and standardize historical transaction data, classify and organize them by modality type, and construct the initial feature matrix and label mapping relationship;

[0042] Design a dedicated reinforcement learning environment for each modality, define the modality-specific state space and action space, and design reward functions tailored to the characteristics of different modalities;

[0043] Train the initial modality-specific reinforcement learning model, randomly shuffle the training data to improve generalization ability, and fuse different modality training sets to generate comprehensive training data;

[0044] Collect real-time transaction feedback as new training samples, evaluate model performance, strengthen the training intensity of specific modalities in a targeted manner, and dynamically adjust the weight ratio of each modality in the fusion process.

[0045] The real-time investment decision-making method based on multimodal fusion includes the following steps:

[0046] Obtaining social media sentiment data, public opinion event classification data, and market structured data, and performing weighted fusion processing on the data to generate a comprehensive sentiment representation vector;

[0047] Extracting features based on historical market data to determine whether the current market state is a low volatility market or a high volatility market, and generating a time-weighted market state representation based on the market state;

[0048] receiving the comprehensive sentiment representation vector and the time-weighted market state representation, generating an asset allocation decision based on a reinforcement learning framework, and outputting an investment portfolio allocation ratio;

[0049] Integrate the decision results of reinforcement learning strategy, decision tree strategy and regression analysis strategy, dynamically calculate the weight of each strategy, and generate collaborative decision signals;

[0050] Predicting current market volatility, assessing investment risk levels, and dynamically adjusting execution parameters of the collaborative decision-making signal based on the risk levels;

[0051] The collaborative decision signal and the adjusted execution parameters are received, a final investment decision signal is generated, and the final investment decision signal is output to a transaction execution system.

[0052] The beneficial effects of the present invention include:

[0053] 1. Through a multimodal information fusion mechanism, we achieve comprehensive analysis of social media text data, news and public opinion events, and structured market data, improving our ability to fully understand the market;

[0054] 2. Through the time-sensitive attention mechanism, the system can dynamically adjust feature weights based on the time sensitivity of market state characteristics, enhancing the system's adaptability to market temporal changes;

[0055] 3. Through reinforcement learning portfolio management and multi-strategy collaborative decision-making, dynamic adjustment of investment portfolios and intelligent coordination of strategies are achieved, significantly improving the accuracy and robustness of decision-making;

[0056] 4. Through the risk adaptive control mechanism, the system can dynamically adjust decision execution parameters according to market risk levels, effectively reducing the risk of loss under extreme market conditions;

[0057] 5. Through the multimodal cyclic optimization mechanism, continuous optimization of training data and adaptive iteration of the model are achieved, enabling the system to continuously learn and adapt to market changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 Schematic diagram of the overall structure of the real-time investment decision-making system based on multimodal fusion of the present invention;

[0059] Figure 2 Schematic diagram of the structure of the multimodal emotion information fusion module of the present invention;

[0060] Figure 3 This is a structural diagram of the market status judgment module of the present invention;

[0061] Figure 4 This is a schematic diagram of the structure of the reinforcement learning combination management module of the present invention;

[0062] Figure 5 This is a schematic diagram of the structure of the multi-strategy collaborative decision-making module of the present invention;

[0063] Figure 6 This is a schematic diagram of the structure of the risk adaptive control module of the present invention;

[0064] Figure 7 Schematic diagram of the structure of the multi-modal cycle optimization module of the present invention;

[0065] Figure 8 This is a flow chart of the real-time investment decision-making method based on multimodal fusion of the present invention. DETAILED DESCRIPTION

[0066] Please refer to the attached Figure 1-8 , the specific implementation of the present invention is further described in detail below with reference to the accompanying drawings.

[0067] like Figure 1 As shown, the real-time investment decision-making system based on multimodal fusion provided by the present invention includes a multimodal emotional information fusion module 1, a market status judgment module 2, a reinforcement learning portfolio management module 3, a multi-strategy collaborative decision-making module 4, a risk adaptive control module 5, a decision signal generation module 6 and a multimodal cycle optimization module 7.

[0068] Reference Figure 2 The multimodal sentiment information fusion module 1 is primarily used to acquire social media sentiment data, public opinion event classification data, and market structured data, and then performs weighted fusion processing on these data to generate a comprehensive sentiment representation vector. This module includes a sentiment data acquisition unit 11, a sentiment vector encoding unit 12, a sentiment information classification and integration unit 13, and a timeliness adjustment unit 14.

[0069] The sentiment data collection unit 11 is used to obtain multi-source heterogeneous data from different channels. In one embodiment of the present invention, the unit obtains unstructured text data from social media platforms such as Twitter, Weibo, and Reddit, obtains classification information of public opinion events from news sources such as Bloomberg and Reuters, and obtains structured market data from financial databases such as Yahoo Finance. Preferably, the frequency of social media data collection is 5 minutes / time, the frequency of news public opinion event collection is 30 minutes / time, and the frequency of market structured data collection is 1 minute / time. For example, in the analysis of Tesla stock investment decisions, the system can collect tweets with the "TSLA" tag on Twitter, news reports about Tesla on Bloomberg, and market data such as Tesla's stock price and trading volume on Yahoo Finance in real time.

[0070] The emotion vector encoding unit 12 is used to convert different types of raw data into a unified vector representation. For unstructured text data, this unit uses the BERT pre-training model for encoding processing to generate a 768-dimensional text representation vector. Specifically, for a text sequence , generate the text encoding vector in the following way:

[0071] ,

[0072] in: is the text encoding vector, with a dimension of 768; Represents the BERT pre-training model encoding function; is the input text sequence; The first word unit; is the sequence length; Represents a 768-dimensional real number vector space. In practical applications, for example, when processing a news headline like "Tesla's second-quarter earnings report exceeded analysts' expectations," the BERT model generates a 768-dimensional vector that captures its semantic information.

[0073] For event classification information, it is converted into a discrete category vector, and each category corresponds to an independent dimension. For example, for event type , its category vector is expressed as:

[0074] ,

[0075] in: For event type The category vector of ; is the total number of event types; Elements in the category vector are 1, and the rest are 0. In an investment decision-making scenario, event types may include "earnings announcement," "management change," "product launch," etc. For example, if Tesla releases a new product, the event might be classified as "product launch" (type 3), in which case the third element in its category vector is 1, and the rest are 0.

[0076] For structured market data, normalization is performed to generate a numerical vector:

[0077] ,

[0078] in: For the Normalized values ​​of market indicators; is the original value; and The normalized value is 0.5, which is the minimum and maximum value of the indicator in the past 30 trading days. For example, for Tesla's daily trading volume data, if the daily trading volume is 10 million shares, the maximum trading volume in the past 30 days is 15 million shares, and the minimum trading volume is 5 million shares.

[0079] The sentiment information classification and integration unit 13 is used to integrate the relevant sentiment vectors for each public opinion event. In practical applications, the same public opinion event may be discussed on multiple social media platforms, generating multiple sentiment vectors. This unit first identifies all sentiment vectors related to a specific event and then performs a weighted integration based on importance. Specifically, for event $i$, its fused sentiment vector is calculated as follows:

[0080] ,

[0081] in: For events The fused emotion vector; For events The number of related sentiment vectors; For the Related sentiment vectors; is its weight coefficient; Indicates that all For example, if Tesla releases new battery technology, this event might be discussed simultaneously on Twitter, Weibo, and Reddit. The system collects the sentiment vectors of related discussions on these three platforms, assigns them weights (e.g., Twitter: 0.5, Weibo: 0.3, Reddit: 0.2), and then sums them to create a fused sentiment vector.

[0082] Weight coefficient Comprehensive calculation based on information source reliability, influence range and emotional intensity:

[0083] ,

[0084] in: Score the reliability of the information source (O-1); Score the impact range (0-1); rate the intensity of the emotion (0-1); 、 、 is the weight coefficient and satisfies Preferably, , In practical applications, for example, news from Bloomberg may receive a higher source reliability score (0.9), while posts from a small investment forum may only receive a reliability score of 0.3.

[0085] The timeliness adjustment unit 14 is used to consider the timeliness of sentiment information. In financial markets, the value of information decays over time, and emergencies often have a higher impact. This unit uses a time decay function to reduce the weight of outdated sentiment information and implements a weight increase mechanism for emergencies. Specifically, the timeliness weight is calculated as follows:

[0086] ,

[0087] in: For the moment The timeliness weight of the information released; is the current time (in hours); The time when the information is released (in hours); is the time decay factor, which controls the rate at which the weight decays over time; is the emergency indicator function, which takes the value 1 when the event is an emergency and 0 when it is a normal event; is the emergency weight enhancement coefficient; represents an exponential decay function. Preferably, hour, For example, if Tesla suddenly announces a recall of some models, the information weight of this emergency will be 50% higher than that of ordinary information. For information that is released 10 hours later than the current time, its basic timeliness weight is .

[0088] Finally, the timeliness-weighted comprehensive sentiment representation vector is calculated as follows:

[0089] ,

[0090] in: is the comprehensive emotion representation vector; is the number of all relevant events; For events Time of occurrence; For events The fused emotion vector; Indicates that all This calculation ensures that recent and unexpected events receive higher weight in decision making. For example, in a Tesla investment decision, the system might consider multiple events, including the recent earnings release, product announcements, and industry news, but would give significantly more weight to the recent quarterly earnings release (which occurred two hours ago) than to an analyst rating adjustment from a week ago.

[0091] Reference Figure 3 The market state determination module 2 is primarily used to extract features based on historical market data, determine whether the current market state is low volatility or high volatility, and generate a time-weighted market state representation based on the market state. This module includes a feature extraction unit 21, a market classification unit 22, a time sensitivity determination unit 23, and a feature update unit 24.

[0092] The feature extraction unit 21 is used to extract features from the time series market data. The present invention adopts the CNN+LSTM network structure to capture local features and long-term dependencies at the same time. Specifically, for the input time series market data , first extract local features through a one-dimensional convolutional layer:

[0093] ,

[0094] in: is the convolution feature vector at time t; Represents a one-dimensional convolution operation; represents the market data sequence from time tk to time t, forming a sliding window; k is the convolution kernel size, indicating the window length. In practical applications, for example, for Tesla stock, the system uses the closing price, trading volume, volatility, and other data from the last three days (k=3) as the convolution window to extract short-term price patterns.

[0095] Then, the convolutional features are fed into the LSTM network to capture temporal dependencies:

[0096] ,

[0097] in: For LSTM at time The hidden state vector of For LSTM at time The unit state vector of For the moment The hidden state vector of For the moment The cell state vector of LSTM() represents the forward propagation function of the LSTM network. LSTM networks are able to capture long-term dependencies, such as seasonal patterns in Tesla’s stock price or long-term responses to macroeconomic events.

[0098] Finally, the market feature vector is:

[0099] ,

[0100] in: is the market characteristic vector; For LSTM at the last moment of the sequence The hidden state vector of contains the information of the entire time series data. Preferably, the convolution kernel size is 3 and the LSTM hidden layer dimension is 128. This means that the final market feature vector is 128-dimensional, which can efficiently encode market state information.

[0101] The market classification unit 22 is used to determine the current market state based on the market feature vector. This unit uses a fully connected network to calculate the probability distribution of the market state:

[0102] ,

[0103] in: is the probability distribution of market states; is the market state, with values ​​of {low volatility, high volatility}; is the weight matrix, with dimension (2 state categories, 128-dimensional feature vector); is the bias vector, dimension is 2; for Activation function, converts the output into a probability distribution. (High Volatility) When , it is judged as a high volatility market, otherwise it is a low volatility market. This threshold setting is based on historical market data analysis and can effectively distinguish between normal fluctuations and abnormal fluctuations.

[0104] For example, during a market panic, the system may detect that the feature vector of Tesla stock indicates a high volatility probability of 0.75, which exceeds the threshold of 0.6, and therefore judges the current market state as a high volatility market, which will trigger the system to adopt a more conservative investment strategy.

[0105] The time sensitivity determination unit 23 is used to analyze the time correlation of market features. In financial markets, some features are highly dependent on historical data (time-sensitive), while others are more dependent on market conditions at a specific moment (moment-sensitive). This unit determines the time sensitivity of feature vectors by calculating the autocorrelation coefficient of the feature vector over the time series:

[0106] ,

[0107] in: Time delay The autocorrelation coefficient of for The eigenvector of the moment; is the average value of the feature vector in the time dimension, calculated as ; is the delay step, usually set to 1; Indicates from arrive The sum of Indicates from arrive The sum of . When , it is judged as a time-sensitive feature; otherwise, it is a moment-sensitive feature. Preferably, the threshold ,This threshold is set based on the characteristics of financial time series data, and can effectively distinguish ,features with temporal continuity and sudden features.

[0108] For example, the trading volume of Tesla stock typically has strong temporal continuity (i.e., today's trading volume is highly correlated with yesterday's trading volume), and its autocorrelation coefficient $R(1)$ may be 0.8, exceeding the threshold of 0.5, and therefore is judged as a time-sensitive feature. However, the impact of Tesla's breaking news events may be moment-sensitive, and its autocorrelation coefficient may be only 0.2.

[0109] The feature updating unit 24 is used to select different feature updating strategies according to the time sensitivity judgment result. For time-sensitive features, a progressive weighted update method is adopted:

[0110] ,

[0111] in: is the updated feature vector; is the historical feature vector; is the current feature vector; is the weight coefficient, which controls the retention ratio of historical information. This weighting allows the system to retain its memory of historical information while moderately absorbing new information. For example, for Tesla’s trading volume characteristics, the system assigns 70% of the weight to historical trading volume patterns and 30% to the most recently observed trading volume, thereby achieving a smooth transition.

[0112] For time-sensitive features, a multi-dimensional fusion update method is adopted:

[0113] ,

[0114] in: is the updated feature vector; is the historical feature vector; is the current feature vector; Represents a vector concatenation operation, which connects two vectors in dimension; is the attention weight matrix, learned through the attention mechanism. For example, when Tesla releases a major product announcement, the system will combine this emergency information with historical information and then assign weights to different dimensions through the attention mechanism, so that decision-makers pay more attention to important information dimensions.

[0115] As described in claim 8, the feature updating unit 24 uses a multi-head attention mechanism, including a short-term attention head, a medium-term attention head, and a long-term attention head, which focus on market features on time scales of 1-3 days, 5-20 days, and 30-90 days, respectively. The multi-head attention mechanism is calculated as follows:

[0116] ,

[0117] ,

[0118] in: For the The output of an attention head; Calculate the attention function; is the query matrix; is the bond matrix; is the value matrix; 、 、 For the The parameter matrix of the attention heads, used to project the input into a lower-dimensional space; Represents a splicing operation; is the number of attention heads; is the output projection matrix; is the multi-head attention calculation function. Preferably, the number of short-term, medium-term, and long-term attention heads are 4, 6, and 2, respectively, for a total of 12 attention heads.

[0119] For example, for Tesla stock, the short-term attention head may focus on the price fluctuations of the last three days, the medium-term attention head may focus on the trend formation of the last ten days, and the long-term attention head may focus on quarterly cyclical patterns. This multi-timescale design enables the system to simultaneously capture market information of different time spans.

[0120] Attention weights are automatically adjusted based on the relevance of each time scale to the current decision:

[0121] ,

[0122] in: For the The weight of each attention head; Score its relevance, is an exponential function; The exponential sum of scores for all attention heads, used for normalization; is the total number of attention heads. The relevance score is calculated as follows:

[0123] ,

[0124] in: For the The relevance score of each attention head; is the projection result of the query vector; is the transpose of the key vector projection; is the dimension of the key vector; It is a normalization factor used to prevent the dot product result from being too large and causing the softmax gradient to disappear.

[0125] For example, in analyzing the market reaction after Tesla’s product launch, if short-term price changes are highly correlated with the product launch, the short-term attention head may receive a higher weight; while when analyzing the long-term impact of Tesla’s quarterly financial report, the long-term attention head may receive a higher weight.

[0126] Finally, the time-weighted market state representation is calculated as:

[0127] ,

[0128] in: It is a time-weighted market state representation vector that contains market information at multiple time scales and is used for subsequent decision making.

[0129] Reference Figure 4The reinforcement learning portfolio management module 3 is primarily used to receive the comprehensive sentiment representation vector and the time-weighted market state representation, generate asset allocation decisions based on the reinforcement learning framework, and output the portfolio allocation ratio. This module includes a state space construction unit 31, a policy network unit 32, a dynamic configuration execution unit 33, and a parameter update optimization unit 34.

[0130] The state space construction unit 31 is used to define the portfolio state space, action space, and state transition rules. The state space includes multi-dimensional features such as holding ratio, return performance, and risk indicators:

[0131] ,

[0132] in: for The state vector at the moment; is the asset holding ratio vector, which represents the weight distribution of each asset in the investment portfolio; is the historical yield series, including the past The rate of return in a time window; It is a volatility indicator, which indicates the degree of volatility of the market or portfolio; is the comprehensive emotion representation vector, which comes from the multimodal emotion information fusion module; It represents the market status and comes from the market status judgment module; [] represents the vector concatenation operation.

[0133] For example, for a technology stock portfolio that includes Tesla, the state space may include: Tesla holding ratio (30%), Apple holding ratio (25%), Microsoft holding ratio (20%), cash ratio (25%), daily return series of the past 5 days, 30-day volatility index, Tesla-related sentiment representation vector, and current market state representation.

[0134] The action space includes the asset adjustment range, adjustment direction and adjustment frequency:

[0135] ,

[0136] in: for The action vector at the moment; is the adjustment range, which indicates the change in the holding ratio; To adjust the direction, it means increasing or decreasing the position; To adjust the frequency, it means the time interval between transaction executions. , that is, the maximum single adjustment shall not exceed 20% to avoid excessive transaction costs and market shocks; adjustment direction , respectively indicating reducing positions, holding positions unchanged and increasing positions; adjustment frequency , indicating the number of adjustments per day.

[0137] For example, the system may generate an action: Increase Tesla stock holdings by 10% ( , ), and perform adjustments 3 times a day ( This fine-grained action space design enables the system to precisely control trading behavior and adapt to different market environments.

[0138] The state transition rule defines how the action affects the state at the next moment:

[0139] ,

[0140] ,

[0141] ,

[0142] in: is the position ratio vector after the action is executed; is the current holding ratio vector; is the position adjustment amount (magnitude multiplied by direction); is the updated historical yield series; Indicates the earliest return in the removed sequence; is the rate of return at the current moment; is the updated volatility indicator; is the smoothing coefficient, which controls the impact of historical volatility; is the absolute value of the current rate of return, indicating the current fluctuation range. This setting allows the volatility calculation to have appropriate historical memory while being able to reflect the latest market fluctuations in a timely manner.

[0143] For example, if Tesla's current holding ratio is 30%, the system decides to increase it by 5% ( , ), the holding ratio after execution will be updated to 35%. At the same time, the system will update the yield series and volatility indicators to reflect the latest market performance.

[0144] The strategy network unit 32 is used to generate asset allocation decisions and evaluate the value of market conditions. This unit adopts the Actor-Critic architecture, including the Actor network and the Critic network. The Actor network generates asset allocation decisions:

[0145] ,

[0146] in: In state Take action The probability distribution of Represents an Actor Network, typically a multi-layer neural network. For example, when determining whether a Tesla stock holding should be adjusted, the Actor Network calculates the probabilities of increasing, decreasing, or maintaining holdings based on current market conditions, sentiment analysis, and historical performance.

[0147] The Critic Network evaluates the current market value:

[0148] ,

[0149] in: Status ; Critic() represents the Critic network, also a multi-layer neural network. The Critic network assesses the potential contribution of the current market state to long-term returns, helping the system determine whether the current market environment is suitable for trading.

[0150] The reward function combines three dimensions: rate of return, volatility, and maximum drawdown:

[0151] ,

[0152] in: In state Next action After transfer to state Rewards received; is the current rate of return, calculated as the relative change in the asset value after the action is performed; is volatility, which indicates the risk level; is the maximum drawdown, calculated as the largest percentage drop from the historical peak to the current value; 、 and are weight coefficients that control the importance of returns, volatility risk, and tail risk in rewards. , , This weight setting reflects the balanced consideration of returns and risks, especially the emphasis on tail risk (maximum drawdown).

[0153] For example, if Tesla stock generates a 2% daily return after adjustment, with a daily volatility of 1.5% and a maximum drawdown of 3%, the reward calculation is: This negative reward indicates that despite the positive gain, the high drawdown makes the decision overall negative.

[0154] The dynamic configuration execution unit 33 is used to generate an asset configuration decision sequence and dynamically adjust the execution parameters. At each decision point, this unit generates a decision sequence ,in Indicates the The allocation ratio of assets, is the amount of assets. At the same time, calculate the expected reward and risk value of each decision point:

[0155] ,

[0156] ,

[0157] in: for The expected reward at the moment; For all possible actions sum; In state Take action probability; For the corresponding reward; Risk for The value at risk at the moment, calculated as the standard deviation of the reward; is the square of the difference between the reward and the expected reward. These calculations help the system evaluate the risk-return characteristics of decisions and guide subsequent parameter adjustments.

[0158] Automatically reduce decision steps in highly volatile markets:

[0159] ,

[0160] in: This is the adjustment range in a highly volatile market; is the standard adjustment range; is a scaling factor used to reduce transaction size in highly volatile markets. This setting allows the system to take a more cautious adjustment strategy in highly volatile markets, reducing risk. For example, during periods of extreme market volatility, a plan to adjust Tesla's holdings by 10% might be reduced to 5% to reduce trading risk.

[0161] At the same time, set upper and lower limits for the configuration ratio adjustment range:

[0162] ,

[0163] in: For assets Configuration ratio; and Assets The minimum and maximum allocation ratio limits. Preferably, for stock assets, , For bond assets, , For cash assets, , These limits ensure that the portfolio is appropriately diversified and avoids excessive concentration in a single asset. For example, even if the system is extremely bullish on Tesla, its proportion in the portfolio will not exceed 40% to control the risk of a single stock.

[0164] The parameter update optimization unit 34 is used to update model parameters based on actual transaction feedback. This unit uses the approximate policy optimization method (PPO algorithm) to update the actor network and the temporal difference learning to update the critic network. The actor network parameters are updated as follows:

[0165] ,

[0166] in: is the updated Actor network parameter; Actor network parameters before updating; is the learning rate, which controls the step size of parameter update; is the policy objective function About parameters The policy objective function is calculated as follows:

[0167] ,

[0168] in: is the policy objective function; Express expectations in the time dimension; is the strategy ratio, which represents the probability ratio of the new strategy to the old strategy; is the advantage function, which represents the difference between actual return and expected return; is the clipping function, which limits the policy ratio to within the scope; is the cropping parameter. Preferably, , This pruning mechanism prevents policy updates from being too large and improves training stability.

[0169] For example, in Tesla stock decisions, if the new strategy recommends a 3x higher buy action than the old strategy, the trimming mechanism will limit this ratio to 1.2x to avoid drastic strategy changes.

[0170] The critic network parameters are updated as follows:

[0171] ,

[0172] in: is the updated critic network parameter; is the critic network parameter before updating; is the learning rate; is the value loss function About parameters The gradient of . The value loss function is calculated as follows:

[0173] ,

[0174] in: is the value loss function; Express expectations in the time dimension; is the state of the Critic network Estimated value of is the target value; is the square of the difference between the estimated value and the target value. .

[0175] Introducing long-term value factors to balance short-term gains and long-term stability:

[0176] ,

[0177] in: is the target value; is the current moment profit; is the discount factor, controlling for the attenuation of the importance of future earnings; Estimate the value of the next state; is the long-term value weight; For the future The average return of time steps, where Indicates the future The sum of the returns at each time step, is an average operation. Preferably, (Trading day). This setting allows the system to consider the impact of decisions on long-term performance while pursuing short-term gains, improving the robustness of the strategy.

[0178] For example, when making investment decisions at Tesla, the system not only considers possible price fluctuations for the day, but also considers the average expected returns over the next 20 trading days to avoid sacrificing long-term value for short-term gains. This is particularly important when evaluating long-term events such as quarterly financial reports and product launches.

[0179] Reference Figure 5 The multi-strategy collaborative decision-making module 4 is mainly used to integrate the decision results of the reinforcement learning strategy, decision tree strategy, and regression analysis strategy, dynamically calculate the weights of each strategy, and generate a collaborative decision signal. This module includes a multi-strategy generation unit 41, a strategy evaluation unit 42, a weight calculation unit 43, and a decision fusion unit 44.

[0180] The multi-strategy generation unit 41 is used to generate multiple investment strategy signals based on different algorithms. In addition to the reinforcement learning strategy, this unit also includes the decision tree strategy and the regression analysis strategy. The decision tree strategy constructs decision rules based on market characteristics:

[0181] ,

[0182] in: is a decision tree model; It is a feature space, which includes multi-dimensional features such as market status and technical indicators; is the decision space, which includes trading decisions such as buy, sell, and hold. Preferably, the decision tree uses the CART algorithm, with a maximum depth of 5 and a minimum number of leaf node samples of 100. These parameter settings balance the complexity and generalization ability of the model.

[0183] For example, a decision tree might generate the following rules: If Tesla's Relative Strength Index (RSI) is below 30 and sentiment analysis is positive, then buy; if the RSI is above 70 and market volatility is high, then sell; otherwise, hold the position. This rule-based decision-making approach provides intuitive and interpretable trading signals.

[0184] Regression analysis strategies use linear regression or ARIMA models to predict market trends:

[0185] ,

[0186] in: for The price forecast value at the moment; is the intercept term; is the weighted sum of the explanatory variables, where For the The regression coefficients of the variables, for The explanatory variable value at time t, is the model order; is a random error term, which obeys normal distribution; Express The regression model preferably adopts the RollingWindow method with a window size of 60 trading days, and uses the most recent data for rolling training to maintain the timeliness of the model.

[0187] For example, the system might use regression analysis to predict Tesla's future price trend, taking into account factors such as historical prices, trading volume, and market sentiment. If the forecast indicates an upward price trend, a buy signal is generated; otherwise, a sell signal is generated.

[0188] The strategy evaluation unit 42 is used to record the historical performance of each strategy under different market conditions. This unit maintains a strategy performance matrix ,in is the number of strategies, is the number of market states. Matrix elements Representation Strategy In market status The historical performance score is calculated as follows:

[0189] ,

[0190] in: For strategy In market status Historical performance ratings under; For strategy In market status The number of historical decisions under is the corresponding time point set; Indicates the sum of all time points in the set; for The rate of return at the moment; for Volatility at the moment; and are weight coefficients, which control the importance of returns and risks in the scoring. , ,This weight setting makes the assessment more focused on risk ,control.

[0191] For example, if an RL strategy generates an average return of 2% and a volatility of 3% in a high volatility market, its performance score is calculated as: This negative score indicates that the strategy underperforms in high-volatility markets, primarily due to the excessive volatility it generates.

[0192] At the same time, calculate the applicability of each strategy to the current market status:

[0193] ,

[0194] in: For strategy Suitability rating; Indicates that all Sum of market states; For strategy In market status Historical performance ratings under; For the current market status and historical market conditions The similarity is calculated as follows:

[0195] ,

[0196] Among them: Sim Status and similarity; For market status The eigenvector of is the square of the Euclidean distance between two eigenvectors; is the kernel width parameter, which controls the sensitivity of similarity calculation; is an exponential function. Preferably, .

[0197] For example, suppose the current market state has a similarity of 0.8 to a historical high-volatility market state and a similarity of 0.2 to a historical low-volatility market state. The reinforcement learning strategy has a performance score of -2.5% in high-volatility markets and 3% in low-volatility markets. Its current suitability score is: , indicating that the strategy may not be suitable for the current market conditions.

[0198] The weight calculation unit 43 is used to dynamically assign weights to each strategy. This unit uses the softmax function to calculate the strategy weights:

[0199] ,

[0200] in: is the weight of strategy i; The index value for scoring the suitability of strategy i; The sum of all M strategy applicability score index values ​​is used for normalization; This softmax calculation ensures that the sum of all policy weights is 1, and policies with better performance receive higher weights.

[0201] In order to improve the stability of weight calculation, smoothing is introduced:

[0202] ,

[0203] in: is the weight after smoothing; is the historical weight; is the currently calculated weight; is a smoothing coefficient that controls the speed of weight update. Preferably, ,This setting avoids excessive fluctuation of strategy weights and enhances the ,stability of the decision.

[0204] For example, assuming the historical weight of the reinforcement learning strategy is 0.5 and the currently calculated weight is 0.3, the smoothed weight is: (1- This smoothing process avoids drastic changes in strategy weights and improves the stability of the system.

[0205] The decision fusion unit 44 is used to integrate the weighted multi-strategy decision signals. This unit first normalizes the decision signals of each strategy and then performs weighted averaging based on the strategy weights:

[0206] ,

[0207] in: is the fused decision signal; Indicates that all The sum of strategies; For strategy The smoothing weight of For strategy Normalize() represents the normalization operation, which maps the decision signal to a uniform scale.

[0208] For example, suppose the system integrates a reinforcement learning strategy (weight 0.46, signal is buy, strength 0.8), a decision tree strategy (weight 0.34, signal is hold, strength 0.6), and a regression analysis strategy (weight 0.2, signal is buy, strength 0.5). The normalized fusion decision signal is: , indicating a medium-strength buy signal.

[0209] To eliminate conflicts between strategies, this unit uses a hierarchical voting mechanism to ensure consistency in decision-making direction. When the weighted decision signal strength exceeds the preset threshold, the final collaborative decision signal is generated:

[0210] ,

[0211] in: It is the final collaborative decision signal; is the fused decision signal; is the absolute value of the decision signal, indicating the signal strength; is the decision threshold, and the transaction is executed only when the signal strength exceeds the threshold. =0.3. This threshold setting ensures that trading decisions are executed only when the majority of strategies reach a strong consensus, reducing unnecessary trading frequency and lowering transaction costs.

[0212] In the above example, the fusion decision signal strength is 0.468, exceeding the threshold of 0.3, so the system generates a final buy signal. If the fusion signal strength is lower than the threshold, the system will maintain the current position unchanged to avoid frequent trading.

[0213] Reference Figure 6 The risk adaptive control module 5 is mainly used to predict the current market volatility, assess the investment risk level, and dynamically adjust the decision execution parameters according to the risk level. This module includes a volatility prediction unit 51, a risk assessment unit 52, a parameter adjustment unit 53, and a risk hedging unit 54.

[0214] The volatility prediction unit 51 is used to predict future market volatility based on historical market data and current market conditions. This unit uses the GARCH (1,1) model to capture the volatility clustering effect:

[0215] ,

[0216] in: for The conditional variance at time t, i.e. the square of the predicted volatility; is a constant term, which represents the underlying level of volatility; is the ARCH term coefficient, which controls the impact of the previous market shock on the current volatility; is the square of the residual at the previous moment, indicating the intensity of the market shock; is the GARCH term coefficient, which controls the impact of the volatility at the previous moment on the current volatility; is the conditional variance of the previous moment. Preferably, the parameters are determined by the maximum likelihood estimation method, and the initial value is set to , , ,These parameter settings are based on empirical values ​​from financial market volatility modeling and are suitable for most market environments. The stability of the model is ensured, and the sum close to 1 reflects the high persistence of volatility in financial markets.

[0217] For example, when predicting the daily volatility of Tesla stock, if the squared residual of the previous day is 0.0004 (corresponding to a 2% price change) and the conditional variance of the previous day is 0.0009 (corresponding to a 3% volatility), then the conditional variance of the current day's forecast is: , the corresponding volatility is 2.86%.

[0218] To improve the accuracy of forecasts, this unit also considers the impact of market sentiment on volatility:

[0219] ,

[0220] in: is the adjusted volatility forecast value; is the underlying volatility predicted by the GARCH model, calculated as ; is the emotional intensity index of the emotional representation vector, with a value range of [-1,1], and its absolute value Indicates emotional intensity, regardless of direction; is the influence coefficient, which controls the influence of emotional factors on volatility prediction. , this setting reflects the moderate impact of market sentiment on volatility.

[0221] For example, if the base volatility forecast is 2.86% and the sentiment intensity is 0.8 (indicating strong positive or negative sentiment), the adjusted volatility is: The adjustment reflects that strong market sentiment typically leads to higher volatility.

[0222] The risk assessment unit 52 is used to comprehensively assess the investment risk level. This unit considers various risk factors, including volatility, liquidity risk, and tail risk:

[0223] ,

[0224] in: is the total risk score, and its value range is usually [0,1]; is the adjusted volatility forecast value; is the benchmark volatility (such as the average of the VIX index); is the relative volatility, which indicates the current volatility level relative to the benchmark; is the liquidity risk indicator, with a value range of [0,1]; is the tail risk indicator, with a value range of [0,1]; 、 and is the weight coefficient and satisfies Preferably, , , This weight distribution reflects the dominance of volatility risk in the overall risk assessment, while also taking into account the impact of liquidity risk and tail risk.

[0225] For example, if Tesla's adjusted volatility is 3.55%, the benchmark volatility is 2%, the relative volatility is 1.775, the liquidity risk indicator is 0.3, and the tail risk indicator is 0.4, then the total risk score is: , the part exceeding 1 is truncated, and the final risk score is 1.0, indicating an extremely high risk level.

[0226] The liquidity risk indicator is calculated as follows:

[0227] ,

[0228] in: is a liquidity risk indicator; is the number of assets; Indicates that all The sum of the assets; For assets bid-ask spread; for its price; The relative spread indicates the relative magnitude of transaction costs. For example, if the bid-ask spread for Tesla stock is $0.20 and the price is $200, the relative spread is 0.1%. For portfolios containing multiple assets, the average relative spread is used as an indicator of liquidity risk.

[0229] Tail risk indicators are based on the conditional value at risk (CVaR) calculation:

[0230] ,

[0231] in: is the tail risk indicator; Is the confidence level The conditional value at risk under , calculated as the expected value of losses in excess of value at risk (VaR); is a yield sequence. Preferably, , which focuses on the average loss in the worst 5% of trading days. For example, if Tesla stock loses an average of 4% in the worst 5% of trading days, its CVaR (0.95) is 4%, which better captures potential losses in extreme market conditions.

[0232] The parameter adjustment unit 53 is used to dynamically adjust the decision execution parameters according to the risk score. The parameters adjusted by this unit include the execution step size, execution frequency and position limit:

[0233] ,

[0234] ,

[0235] ,

[0236] in: is the adjusted execution step, which indicates the change in asset ratio for each transaction; is the original execution step length; is the risk adjustment factor, which decreases as risk increases; To adjust the coefficient, control the impact of risk on step size; is the adjusted execution frequency, which indicates the number of transactions per day; is the original execution frequency; The frequency adjustment amount is rounded down; is the frequency adjustment coefficient; is the adjusted maximum position limit; is the original maximum position limit; is the position adjustment factor, which decreases as risk increases; is the position adjustment coefficient. Preferably, , These parameter settings enable the system to significantly reduce transaction size and frequency in high-risk environments and enhance risk control capabilities.

[0237] For example, if the original execution step size is 10% and the risk score is 0.8, the adjusted execution step size is: , meaning that in high-risk environments, the system will significantly reduce the size of each trade. Similarly, if the original execution frequency was five times per day, it might be reduced to three after adjustment; and if the original maximum position limit was 40%, it might be reduced to 20% after adjustment.

[0238] The risk hedging unit 54 is used to activate the automatic risk hedging mechanism in high-risk situations. When the risk score exceeds the warning threshold, this unit generates a hedging strategy:

[0239] ,

[0240] in: For hedging strategies; Score the overall risk; The hedging mechanism is activated only when the risk score exceeds this threshold. This setting ensures that the hedging mechanism is activated only when the risk increases significantly, avoiding the cost losses caused by excessive hedging.

[0241] The specific contents of the hedging instructions include: 1. Increasing the allocation ratio of defensive assets, such as government bonds, gold, etc.; 2. Establishing futures or options positions opposite to current positions; 3. Appropriately increasing the proportion of cash holdings to reduce market exposure.

[0242] Preferably, the hedge ratio is proportional to the risk score:

[0243] ,

[0244] in: is the hedge ratio, which indicates the proportion of assets that need to be hedged; As a basic hedge ratio, it ensures that at least 20% of assets are hedged even if the risk score just exceeds the threshold; The function ensures that the hedge ratio does not exceed 80% to maintain a certain market exposure. For example, if the risk score is 0.9 and the hedge trigger threshold is 0.8, the hedge ratio is: , indicating that the system will hedge 30% of the assets.

[0245] In a Tesla investment, this might mean establishing PUT option protection on 30% of the holdings, or adding 30% to defensive asset allocations to reduce the risk of market declines.

[0246] The decision signal generation module 6 is primarily responsible for receiving the collaborative decision signal and adjusted execution parameters, generating a final investment decision signal, and outputting it to the trade execution system. This module first verifies the collaborative decision signal to ensure it complies with risk control requirements and investment constraints. It then generates specific trading instructions based on the adjusted execution parameters, including the direction, size, and timing of the trade. Finally, it formats the final investment decision signal into a standard trading instruction and outputs it to the trade execution system.

[0247] In practice, the generation of the final investment decision signal takes into account a variety of factors, including market liquidity, transaction costs, and execution slippage. Preferably, for large-scale transactions, a batch execution strategy is adopted to reduce market impact; for low-liquidity assets, a maximum trading volume limit is set to avoid excessive impact on market prices; at the same time, transaction costs and slippage are estimated based on historical trading data, and the execution strategy is optimized to minimize the impact of transaction costs.

[0248] For example, for a large purchase decision of Tesla stock, the system may divide the transaction into three batches, with each batch executed at an interval of 30 minutes to reduce market impact; at the same time, the maximum trading volume is set to no more than 5% of the average trading volume of the day to avoid excessive impact on market prices; in addition, the system will estimate the best trading hours based on historical data (such as 1 hour after opening and 1 hour before closing, when liquidity is usually higher), and optimize the execution timing to reduce slippage and transaction costs.

[0249] Reference Figure 7 The multimodal loop optimization module 7 is mainly used to optimize the system's training data and model parameters, improving the system's learning ability and adaptability. This module implements a complete loop optimization process, including four main steps: data processing, environment construction, training optimization, and adaptive adjustment.

[0250] First, historical transaction data is denoised and standardized, categorized and organized by modality, and an initial feature matrix and label mapping relationship is constructed. The data cleaning process includes steps such as outlier detection and processing, missing value filling, and data normalization to ensure the quality and consistency of training data.

[0251] For example, for Tesla's historical transaction data, the system will detect and process price jumps, abnormal trading volumes, and other situations; for social media data, the system will filter out spam and irrelevant content and extract valuable emotional information; for news data, the system will deduplicate and extract key information to ensure the quality of training data.

[0252] Next, we design a dedicated reinforcement learning environment for each modality, define the modality-specific state space and action space, and design reward functions tailored to the characteristics of each modality. Data from different modalities exhibit distinct characteristics, such as high-dimensional sparsity in text data and strong temporal dependencies in market price data. Therefore, this module designs modality-specific state spaces and reward functions to better capture the characteristics of each type of data.

[0253] For example, for the text modality, the state space may include word frequency features, sentiment polarity, and topic distribution; for the price modality, the state space may include technical indicators, volatility, and trend characteristics. Accordingly, the reward function for the text modality may place more emphasis on sentiment prediction accuracy, while the reward function for the price modality may place more emphasis on trend prediction accuracy.

[0254] Next, the initial modality-specific reinforcement learning model is trained. The training data is randomly shuffled to improve generalization, and training sets from different modalities are fused to generate synthetic training data. The training process uses sample importance sampling techniques to improve learning efficiency in key scenarios. Appropriate regularization methods are also used to prevent overfitting and enhance the model's generalization capabilities.

[0255] For example, the system may give higher sampling weights to days when Tesla's stock fluctuates significantly, so that the model can better learn how to deal with extreme market conditions; at the same time, by adding regularization terms or early stopping techniques, the model can be prevented from overfitting the training data.

[0256] Finally, real-time trading feedback is collected as new training samples to evaluate model performance, specifically strengthening the training intensity of specific modalities and dynamically adjusting the weight ratio of each modality during the fusion process. This process enables continuous learning and adaptive optimization of the system, enabling it to continuously adapt to market changes.

[0257] In terms of dynamic adjustment of modal weights, this module makes adjustments based on the prediction accuracy of each modality:

[0258] ,

[0259] in: For modal The new weight of For modal The old weight of For modal The prediction accuracy of , which is usually evaluated based on the performance on the validation set; is the average accuracy of all modalities, calculated as Accuracy ,in is the total number of modes; To adjust the step size, control the speed of weight update; is the relative accuracy difference, indicating the modality Performance relative to the average. Preferably, ,This setting enables the system to gradually adjust the modal weights while maintaining ,sufficient stability.

[0260] For example, if the prediction accuracy of the text modality is 80% and the prediction accuracy of the price modality is 70%, the average accuracy is 75%. The old weight of the text modality is 0.5, and the new weight is calculated as: This indicates that the text modality receives a slight increase in weight due to its higher accuracy.

[0261] like Figure 8 As shown in FIG, the real-time investment decision-making method based on multimodal fusion includes the following steps:

[0262] Step S1: Acquire social media sentiment data, public opinion event classification data, and market structured data, perform weighted fusion processing on the data, and generate a comprehensive sentiment representation vector;

[0263] Step S2: extracting features based on historical market data, determining whether the current market state is a low volatility market or a high volatility market, and generating a time-weighted market state representation based on the market state;

[0264] Step S3: receiving the comprehensive sentiment representation vector and the time-weighted market state representation, generating an asset allocation decision based on a reinforcement learning framework, and outputting an investment portfolio allocation ratio;

[0265] Step S4: Integrate the decision results of the reinforcement learning strategy, decision tree strategy, and regression analysis strategy, dynamically calculate the weight of each strategy, and generate a collaborative decision signal;

[0266] Step S5: predicting the current market volatility, assessing the investment risk level, and dynamically adjusting the execution parameters of the collaborative decision signal according to the risk level;

[0267] Step S6: Receive the collaborative decision signal and the adjusted execution parameters, generate a final investment decision signal, and output the final investment decision signal to the transaction execution system.

[0268] In a preferred embodiment of the present invention, the weighted fusion processing in step S1 includes: converting unstructured text data into sentiment vectors through a pre-trained model, converting event classification information into category vectors, and normalizing key variable data to generate numerical vectors; collecting relevant sentiment vectors for each public opinion event, and weighting multiple sentiment vectors of the same event according to importance to generate a fused sentiment vector; applying a time decay function to reduce the weight of outdated sentiment information, setting a weight enhancement mechanism for emergencies, and generating a comprehensive sentiment representation vector weighted by timeliness.

[0269] The feature extraction and market status judgment in step S2 include: extracting features from time-series market data through the CNN+LSTM network to generate a market feature vector; calculating the classification probability values ​​of low-volatility markets and high-volatility markets based on the market feature vector to judge the current market status; analyzing the time correlation strength of the market feature vector to judge whether the market characteristics are time-sensitive or moment-sensitive; using a progressive weighted update method for time-sensitive features and a multi-dimensional fusion update method for moment-sensitive features to generate a time-weighted market status representation.

[0270] The asset allocation decision generation in step S3 includes: defining the portfolio state space, including multi-dimensional features such as holding ratio, return performance and risk indicators, designing the action space, and constructing state transition rules; generating asset allocation decisions through the Actor network, evaluating the current market state value through the Critic network, and providing decision quality feedback; generating an asset allocation decision sequence, calculating the expected reward and risk value of each decision point, and dynamically adjusting the execution step size and execution frequency based on the market state; calculating the loss rate and cumulative reward value based on actual transaction feedback, and updating the Actor network and Critic network parameters.

[0271] The multi-strategy collaborative decision-making in step S4 includes: generating multiple investment strategy signals based on different algorithms, including reinforcement learning strategy signals, decision tree strategy signals and regression analysis strategy signals; recording the historical performance of each strategy under different market conditions, and calculating the applicability of each strategy to the current market conditions; dynamically allocating the weight of each strategy through a logistic regression model to generate a weighting coefficient; integrating the multi-strategy decision signals based on the weighting coefficient, eliminating conflicts between strategies, and generating a collaborative decision signal.

[0272] The risk adaptive control in step S5 includes: predicting future market volatility based on historical market data and current market conditions; generating a risk score by comprehensively considering volatility, liquidity, and tail risk; dynamically adjusting decision execution parameters based on the risk score, including adjusting the execution step size, execution frequency, and position limit; and activating an automatic risk hedging mechanism in high-risk situations to generate a hedging strategy.

[0273] Through the above steps, the present invention realizes a real-time investment decision-making system and method based on multimodal fusion, effectively solving the technical problems existing in the existing investment decision-making system in multi-source heterogeneous data processing, market status adaptability and risk control, and significantly improving the accuracy, robustness and risk control capabilities of investment decisions.

[0274] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A real-time investment decision-making system based on multimodal fusion, characterized by: include: A multimodal sentiment information fusion module is used to obtain social media sentiment data, public opinion event classification data, and market structured data, and perform weighted fusion processing on the data to generate a comprehensive sentiment representation vector; A market status judgment module, connected to the multimodal sentiment information fusion module, is used to extract features based on historical market data, judge whether the current market status is a low volatility market or a high volatility market, and generate a time-weighted market status representation based on the market status; a reinforcement learning portfolio management module, connected to the multimodal sentiment information fusion module and the market status judgment module, configured to receive the comprehensive sentiment representation vector and the time-weighted market status representation, generate asset allocation decisions based on a reinforcement learning framework, and output an investment portfolio allocation ratio; A multi-strategy collaborative decision-making module is connected to the reinforcement learning combination management module and is used to integrate the decision results of the reinforcement learning strategy, decision tree strategy and regression analysis strategy, dynamically calculate the weight of each strategy, and generate a collaborative decision signal; A risk adaptive control module, connected to the multi-strategy collaborative decision-making module, is used to predict current market volatility, assess investment risk levels, and dynamically adjust the execution parameters of the collaborative decision-making signal according to the risk level; A decision signal generation module is connected to the multi-strategy collaborative decision module and the risk adaptive control module, and is used to receive the collaborative decision signal and the adjusted execution parameters, generate a final investment decision signal, and output the final investment decision signal to the transaction execution system.

2. The real-time investment decision-making system based on multimodal fusion according to claim 1 is characterized in that: The multimodal emotion information fusion module includes: Sentiment data collection unit, used to obtain unstructured text data from social media, obtain event classification information from news and public opinion analysis, and extract key variable data from financial statements; An emotion vector encoding unit, connected to the emotion data acquisition unit, is used to convert the unstructured text data into emotion vectors through a pre-trained model, convert the event classification information into a category vector, and normalize the key variable data to generate a numerical vector; An emotional information classification and integration unit, connected to the emotional vector encoding unit, is used to collect relevant emotional vectors for each public opinion event, and weight multiple emotional vectors of the same event according to their importance to generate a fused emotional vector; The timeliness adjustment unit is connected to the emotional information classification and integration unit, and is used to apply a time decay function to reduce the weight of outdated emotional information, set a weight enhancement mechanism for emergencies, and generate a timeliness-weighted comprehensive emotional representation vector.

3. The real-time investment decision-making system based on multimodal fusion according to claim 1 is characterized in that: The market status judgment module includes: Feature extraction unit, used to extract features from time series market data through CNN+LSTM network and generate market feature vectors; a market classification unit, connected to the feature extraction unit, for calculating classification probability values ​​of a low volatility market and a high volatility market based on the market feature vector, and determining a current market state; a time sensitivity judgment unit, connected to the market classification unit, for analyzing the time correlation strength of the market feature vector and judging whether the market feature is time sensitive or moment sensitive; The feature updating unit is connected to the time sensitivity judgment unit and is used to adopt a progressive weighted updating method for time-sensitive features and a multi-dimensional fusion updating method for moment-sensitive features to generate a time-weighted market status representation.

4. The real-time investment decision-making system based on multimodal fusion according to claim 1 is characterized in that: The reinforcement learning combination management module includes: The state space construction unit is used to define the portfolio state space, including multi-dimensional features such as holding ratio, return performance and risk indicators, design the action space, and construct state transition rules; The strategy network unit is connected to the state space construction unit and is used to generate asset allocation decisions through the actor network, evaluate the current market state value through the critic network, and provide decision quality feedback; A dynamic configuration execution unit, connected to the strategy network unit, is used to generate an asset allocation decision sequence, calculate the expected reward and risk value of each decision point, and dynamically adjust the execution step size and execution frequency based on market conditions; The parameter update optimization unit is connected to the dynamic configuration execution unit and is used to calculate the loss rate and cumulative reward value based on actual transaction feedback and update the Actor network and Critic network parameters.

5. The real-time investment decision-making system based on multimodal fusion according to claim 1 is characterized in that: The multi-strategy collaborative decision-making module includes: A multi-strategy generation unit, used to generate multiple investment strategy signals based on different algorithms, including reinforcement learning strategy signals, decision tree strategy signals, and regression analysis strategy signals; a strategy evaluation unit, connected to the multi-strategy generation unit, for recording the historical performance of each strategy under different market conditions and calculating the applicability of each strategy to the current market condition; A weight calculation unit, connected to the strategy evaluation unit, for dynamically allocating weights to each strategy through a logistic regression model to generate a weighted coefficient; The decision fusion unit is connected to the multi-strategy generation unit and the weight calculation unit, and is used to integrate the multi-strategy decision signals based on the weight coefficients, eliminate conflicts between strategies, and generate collaborative decision signals.

6. The real-time investment decision-making system based on multimodal fusion according to claim 1 is characterized in that: The risk adaptive control module includes: Volatility prediction unit, used to predict future market volatility based on historical market data and current market conditions; a risk assessment unit, connected to the volatility prediction unit, for comprehensively considering volatility, liquidity, and tail risk to generate a risk score; a parameter adjustment unit, connected to the risk assessment unit, for dynamically adjusting decision execution parameters according to the risk score, including adjusting the execution step size, execution frequency, and position limit; The risk hedging unit is connected to the parameter adjustment unit and is used to activate the automatic risk hedging mechanism in high-risk situations and generate a hedging strategy.

7. The real-time investment decision-making system based on multimodal fusion according to claim 4 is characterized in that: The action space set by the state space construction unit includes the asset adjustment amplitude, adjustment direction and adjustment frequency. The reward function of the strategy network unit combines the three dimensions of rate of return, volatility and maximum drawdown. The dynamic configuration execution unit automatically reduces the decision step size in a highly volatile market and sets upper and lower limit constraints on the adjustment amplitude of the configuration ratio. The parameter update optimization unit introduces a long-term value factor to balance short-term returns and long-term stability.

8. The real-time investment decision-making system based on multimodal fusion according to claim 3 is characterized in that: The feature update unit adopts a multi-head attention mechanism, including short-term attention heads, medium-term attention heads and long-term attention heads, which focus on market characteristics on time scales of 1-3 days, 5-20 days and 30-90 days respectively. The attention weight is automatically adjusted based on the relevance of each time scale to the current decision.

9. The real-time investment decision-making system based on multimodal fusion according to claim 1 is characterized in that: Also includes a multimodal loop optimization module for: De-noise and standardize historical transaction data, classify and organize them by modality type, and construct the initial feature matrix and label mapping relationship; Design a dedicated reinforcement learning environment for each modality, define the modality-specific state space and action space, and design reward functions tailored to the characteristics of different modalities; Train the initial modality-specific reinforcement learning model, randomly shuffle the training data to improve generalization ability, and fuse different modality training sets to generate comprehensive training data; Collect real-time transaction feedback as new training samples, evaluate model performance, strengthen the training intensity of specific modalities in a targeted manner, and dynamically adjust the weight ratio of each modality in the fusion process.

10. A real-time investment decision-making method based on multimodal fusion, using the system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Obtaining social media sentiment data, public opinion event classification data, and market structured data, and performing weighted fusion processing on the data to generate a comprehensive sentiment representation vector; Extracting features based on historical market data to determine whether the current market state is a low volatility market or a high volatility market, and generating a time-weighted market state representation based on the market state; receiving the comprehensive sentiment representation vector and the time-weighted market state representation, generating an asset allocation decision based on a reinforcement learning framework, and outputting an investment portfolio allocation ratio; Integrate the decision results of reinforcement learning strategy, decision tree strategy and regression analysis strategy, dynamically calculate the weight of each strategy, and generate collaborative decision signals; Predicting current market volatility, assessing investment risk levels, and dynamically adjusting execution parameters of the collaborative decision-making signal based on the risk levels; The collaborative decision signal and the adjusted execution parameters are received, a final investment decision signal is generated, and the final investment decision signal is output to a transaction execution system.

Citation Information

Cited By

  • Market fluctuation early warning method and system based on multi-dimensional index fusion

    CN121504601A

  • Investment decision risk assessment method and system based on multi-modal behavior perception

    CN121563679A

  • Investment decision-making method and system integrating game equilibrium analysis and market emotion recognition

    CN121903669A

  • Weeding robot operation fine control system based on reinforcement learning

    CN122195018A

  • Transaction strategy generation method and system based on graph neural network and reinforcement learning

    CN122367528A