Financial risk prediction method based on multi-modal data fusion
Through the combination of multimodal data fusion and large language models, the limitations of a single data source in financial risk prediction are solved, and comprehensive capture of market dynamics and multi-task prediction are achieved, which improves prediction accuracy and stability.
Patent Information
- Application Number
- CN202510343420.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Existing financial risk prediction technologies mainly rely on a single data source, and it is difficult to fully reflect the multi-dimensional drivers of market fluctuations. Traditional models lack the effective fusion ability of multimodal data when processing unstructured data, resulting in insufficient accuracy and reliability of prediction results.
A multimodal data fusion method is adopted, combined with financial report conference call audio, text minutes, news text and timing transaction data, in-depth feature extraction and fusion analysis are performed through large language models, and a multi-task learning framework is used to predict multiple risk indicators.
It significantly improves the accuracy and stability of financial risk prediction, can capture multi-level information about the company's operations and market environment, improves the accuracy of prediction and the adaptability of the model, and can predict multiple risk indicators at the same time.
Smart Images

Figure CN120258951A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence risk prediction, and particularly to a financial risk prediction method based on multi-modal data fusion. Background Art
[0002] With the rapid development of artificial intelligence technology, the intelligent level of the financial industry has been significantly improved. In the field of financial risk prediction, artificial intelligence technology can not only improve the efficiency of risk identification, but also provide important support for enterprises to optimize investment decisions and improve risk management. However, there are still certain deficiencies in current research and practice in how to efficiently integrate multi-modal data to achieve more accurate financial risk prediction. Existing financial risk prediction technologies mainly focus on the utilization of single data sources, such as historical stock price analysis based on time series or news text sentiment analysis based on natural language processing. Although this method is effective in specific scenarios, it is difficult to fully reflect the multi-dimensional driving factors of market fluctuations. Especially in the face of complex market environments and increasing data diversity, the limitations of single data source prediction methods become more obvious. In addition, a large amount of key information in the financial market exists in the form of unstructured data, such as text transcripts, voice recordings of earnings conference calls, and media news reports. These data contain implicit features such as the intonation, speech rate, and emotional changes of speakers, which are important clues for evaluating market risks. However, traditional models have technical bottlenecks in processing unstructured data and lack the ability to effectively fuse multi-modal data, resulting in prediction results that are difficult to comprehensively and accurately reflect market dynamics. The existing technologies have the following significant defects in the field of financial risk prediction: the limitation of prediction results caused by single data sources, as well as the poor adaptability and stability of prediction models, greatly affecting the accuracy and reliability of prediction results.
[0003] In recent years, large language models (LLMs) have gradually been introduced into the financial field due to their advantages in cross-domain text processing, sentiment analysis, and multi-task learning. Research shows that large language models can efficiently process long texts and generate high-quality summaries, and have the ability to deeply analyze financial data. However, it is difficult to comprehensively capture the multi-dimensional dynamics of the financial market relying solely on a single model. How to integrate multi-modal data into a unified prediction framework and conduct in-depth analysis by combining the advantages of large language models has become the core issue that urgently needs to be solved in the current financial risk prediction field. To sum up, it is very necessary to provide a financial risk prediction method based on multi-modal data fusion, comprehensively utilize multi-modal data, conduct deep feature extraction and fusion analysis through large language models, and significantly improve the accuracy and stability of financial risk prediction through the integration of multi-source data and the design of a multi-task prediction framework, providing strong technical support for financial risk management. Summary of the Invention
[0004] In view of this, the present invention proposes a financial risk prediction method based on multi-modal data fusion, which integrates multiple data sources, comprehensively captures the dynamics and market information of target companies, simultaneously predicts multiple risk indicators, and has good prediction comprehensiveness and accuracy.
[0005] The technical solution of the present invention is implemented as follows: The present invention provides a financial risk prediction method based on multi-modal data fusion, including the following steps:
[0006] S1: Extract the audio feature vectors of the earnings conference calls of the target company;
[0007] S2: Extract the text summary feature vectors of the earnings conference calls of the target company;
[0008] S3: Use a large language model to summarize the text summary of the earnings conference calls of the target company to obtain the embedding vector corresponding to the summary of the text summary;
[0009] S4: Extract the news text features of the target company to obtain news text feature vectors;
[0010] S5: Extract the features of the time-series transaction data of the target company within a period before the target date to obtain the feature vectors of the time-series data;
[0011] S6: Fuse the obtained audio feature vectors, text summary feature vectors, embedding vectors corresponding to the summary of the text summary, news text feature vectors, and feature vectors of the time-series data to obtain a joint representation vector; and
[0012] S7: Input the joint representation vector into a multi-task learning framework for risk indicator prediction.
[0013] On the basis of the above technical solution, preferably, the specific content of step S1 is:
[0014] S11: Use the WenetSpeech pre-trained language model to extract the embedding vector of the earnings conference call audio. For a segment of audio after audio preprocessing where represents the i-th frame of the audio, and n represents the number of frames in the audio. Convert each frame of the audio into a vector representation to obtain the audio embedding vector of the entire segment of audio A c
[0015] S12: Feed the audio embedding vector E ac into the multi-head self-attention module MHSA to further extract the audio feature vector T ac = MHSA(E ac )
[0016] S13: Feed the feature vector T ac into the average pooling layer Average Pooling Layer to obtain the compressed audio feature vector T a , T a = AveragePooling(T ac ).
[0017] Preferably, the vector representation of each frame of audio has a dimension of 512, and the dimension of the compressed audio feature vector T a is also 512; the dimension of the text summary feature vector is 768; the dimension of the embedding vector corresponding to the summary of the text summary is 768; the dimension of the joint representation vector is 512.
[0018] Preferably, the multi-head self-attention module MHSA uses multiple attention heads to process the input audio embedding vectors in parallel, and calculates the attention weights of the audio embedding vectors using the following formula: where the query vector Q = E ac W Q , the key vector K = E ac W K , the value vector V = E ac W V , W Q , W K and W V are linear projection matrices, d k is the dimension of the attention head, and the softmax function is used for normalization; the outputs of each attention head are concatenated, and the feature vector T ac of the audio is obtained through a linear transformation, T ac = Concat(head1,head2,head h )W o , where the subscripts 1,2,,h are the distinguishing marks of different attention heads head, and W o is the combined linear transformation matrix, and Concat is the combination function.
[0019] More preferably, the specific content of step S2 is:
[0020] S21: Preprocess the text summary of the earnings conference call of the target company to obtain a sentence set where represents the L-th sentence in the text summary, L = 1,2,,l, and l represents the number of sentences in the text summary; use the Sentence-BERT pre-trained language model to map each sentence to the embedding vector of the earnings conference call text summary Obtain the vector representation of the text transcript of the entire earnings conference call
[0021] S22: Feed the vector representation of the text transcript of the entire earnings conference call into the multi-head self-attention module MHSA to further extract the feature vector T of the text transcript tc , T tc = MHSA(E tc );
[0022] S23: Feed the feature vector T of the text transcript tc , into the average pooling layer Average Pooling Layer to obtain the compressed feature vector T of the text transcript t = AveragePooling(T tc ).
[0023] More preferably, the specific content of step S3 is as follows:
[0024] S31: Paragraph segmentation. Segment the text transcript of the earnings conference call of the target company according to logical paragraphs to obtain paragraphs p M = p1, p2,, p m , M = 1, 2,, m, where the subscript m represents the number of paragraphs; Input each paragraph p M into the large language model LLM to extract the paragraph-level summary s M = LLM(p M ), s M is the summary of paragraph p M ;
[0025] S32: Merge all the paragraph summaries {s1, s2, s M} into the overall text and input it into the large language model to generate the comprehensive summary S, S = LLM({s1, s2, s M});
[0026] S33: Use the Sentence-BERT pre-trained language model to vectorize the comprehensive summary S to generate the embedding vector T corresponding to the summary of the text transcript l , T l = SBERT(S).
[0027] Even more preferably, the specific content of step S4 is as follows:
[0028] S41: Collect the news text data of the target company in the days before the target transaction date, denoted as N = {n1, n2, n k}, where n K0 represents the K0th news text and k is the total number of news;
[0029] S42: Parse each news item n using the large language model LLM K0 to extract metadata m K0 , where m K0 = LLM(n K0 ); Integrate the metadata of all news items into an overall metadata set M N = {m1, m2,..., m k};
[0030] S43: Find the k historical news groups N that are most similar in terms of metadata to the news text data N from several days before the target trading date in the historical news dataset K0 = N1, N2,..., N k , calculate the semantic relatedness between N and any historical news group N K0 , where f is a function that converts news text into an embedding vector, and select the historical news H with the highest semantic similarity to the news text data N from several days before the target trading date;
[0031] S44: Concatenate N, H, and the market trend-related text after H occurred, and then use the SBERT model to convert the concatenated text into an embedding vector T n , where the embedding vector T n serves as the news text feature vector of the target company for several days before the target trading date.
[0032] More preferably, the specific content of step S5 is as follows:
[0033] S51: Collect the time-series trading data D of the target company for 30 days before the target date, including the daily closing price and trading volume, expressed as D = {(p1, v1), (p2, v2),..., (p d , v d )}, where p F = p1, p2,..., p d represents the closing price on the F-th day, and v F = v1, v2,..., v d represents the trading volume on the F-th day;
[0034] S52: Input the time-series trading data D into the bidirectional long short-term memory network Bi-LSTM. The Bi-LSTM captures the time-series dynamic characteristics of the trading data and outputs a feature vector T containing time-series data v , where T v = BiLSTM(D);
[0035] S53: Capture the dynamic relationship between different time-series features through the vector autoregressive VAR model:
[0036] log(σ 3,t ) = α3 + β 1,1 log(σ -3,t ) + β 1,2 log(σ -7,t ) + β 1.3 log(σ -15,t ) + β 1.4. log(σ -30,t ) + u 3,t ;
[0037] log(σ 7,t ) = α7 + β 2,1 log(σ -3,t ) + β 2,2 log(σ -7,t ) + β 2.3 log(σ -15,t ) + β 2.4. log(σ -30,t ) + u 7,t ;
[0038] log(σ 15,t ) = α 15 + β 3,1 log(σ -3,t ) + β 3,2 log(σ -7,t ) + β 3.3 log(σ -15,t ) + β 3.4. log(σ -30,t ) + u 15,t ;
[0039] log(σ 30,t ) = α 30 + β 4,1 log(σ -3,t ) + β 4,2 log(σ -7,t ) + β 4.3 log(σ -15,t ) + β 4.4. log(σ -30,t ) + u 30,t ;
[0040] σ z,t represents the volatility of the target company's stock price within z = 3, 7, 15, and 30 days, where z ≤ F and z ∈ t; u z,t is a white noise term; β a,b is the coefficient matrix of the dynamic relationship, a, b = 1, 2, 3, 4; α z is the intercept term; the volatility of the stock price is defined as the standard deviation of the return rate of the target company within z days:
[0041] More preferably, the combined representation vector E described in step S6 is fused by the following formula: E = w0 + w1T a + w2T t + w3T l + w4T n + w5T v + ε, where w0 is the bias term, w1, w2, w3, w4, w5 are the fusion weights, and ε is the error term representing random noise.
[0042] Even more preferably, the specific content of step S7 is to use the combined representation vector E generated in step S6 as the input of the multi-task learning framework, and the multi-task learning framework simultaneously predicts the following risk indicators: the volatility σ 3,t 、σ 7,t 、σ 15,t and σ 30,t of the stock price of the target company over time spans of 3 days, 7 days, 15 days, and 30 days, as well as the single-day risk value VAR;
[0043] For the stock price volatility of each time span, an independent first prediction sub-network MLP is constructed, and the prediction result of the stock price volatility is where is the predicted volatility; f MLP,z (·) is the first prediction sub-network MLP for the time span z;
[0044] For the single-day risk value VAR, an independent second prediction sub-network MLP is constructed, with the input being the combined representation vector E and the output being the predicted value of VAR where is the predicted value of VAR; f MLP,VAR (·) is the second prediction sub-network MLP;
[0045] The combined loss function simultaneously optimizes the stock price volatility and VAR prediction tasks, and the formula of the combined loss function is: μ is a weight hyperparameter that balances the prediction errors of stock price volatility and VAR; y j and represent the true and predicted values of the stock price volatility index; represents the mean squared error of the stock price volatility prediction; q represents the quantile threshold of the single-day risk value; V and represent the true and predicted single-day risk values respectively; represents the quantile regression loss function of the single-day risk value prediction task.
[0046] A financial risk prediction method based on multi-modal data fusion provided by the present invention has the following beneficial effects compared with the prior art:
[0047] (1) The present invention constructs a new financial risk prediction framework by integrating multi-modal data, including earnings conference call audio and text, news text, and time series data. This multi-modal feature fusion method significantly improves the ability to perceive financial market risks, can capture multi-level information on company operations and market environment, and solves the problem of insufficient prediction accuracy caused by a single data source in the prior art.
[0048] (2) The prior art often ignores the non-explicit features contained in earnings conference call audio, such as the speaker's intonation, speech rate, and emotion. The present invention introduces a multi-head self-attention mechanism (MHSA) and a large language model (LLM) to jointly extract and deeply analyze the features of audio and text data, which can reveal potential information correlations that are difficult to capture by traditional text analysis methods, thereby improving the prediction accuracy and insight.
[0049] (3) The present invention adopts a multi-task learning framework that can simultaneously predict multiple risk indicators, including stock price volatility over different time spans and the single-day value at risk (VAR). This technical solution not only improves the prediction efficiency of the model but also effectively alleviates the dependence on a single task in traditional methods and enhances the model's adaptability to complex task scenarios. In addition, by introducing a joint loss function, the multi-task learning process is optimized, further improving the prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 It is a schematic flowchart of a financial risk prediction method based on multi-modal data fusion according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0053] Most existing financial risk predictions use a single data source, unable to capture the complex correlations between a company's internal operations and the external market environment, resulting in incomplete and inaccurate predictions, and unable to predict multiple risk indicators simultaneously. In view of this, as Figure 1 shown, the present invention provides a financial risk prediction method based on multi-modal data fusion, including the following steps:
[0054] S1: Extract the audio feature vectors of the earnings conference calls of the target company.
[0055] The specific content of step S1 is as follows:
[0056] S10: Audio preprocessing, segment the earnings conference call audio, divide the earnings conference call audio into frames, each frame has a fixed length, such as 25 milliseconds, with a 10-millisecond overlap between adjacent frames to ensure capturing the temporal features of continuous speech; then normalize the earnings conference call audio, standardize the waveform amplitude to eliminate the influence of recording devices or volume differences; convert the earnings conference call audio format, such as WAV or FLAC, for subsequent feature extraction.
[0057] S11: Use the WenetSpeech pre-trained language model to extract the embedding vectors of the earnings conference call audio. For a segment of audio after audio preprocessing where represents the i-th frame of the audio, and n represents the number of frames in the audio. Convert each frame of audio into a vector representation to obtain the audio embedding vector of the entire segment of audio A c of The pre-trained language model WenetSpeech samples the Transformer architecture. This model has been trained on a large-scale speech dataset and can effectively extract the hidden features of speech, and is applicable to the field of natural language processing.
[0058] S12: Feed the audio embedding vector E ac into the multi-head self-attention module MHSA to further extract the feature vector T ac = MHSA(E ac ); The multi-head self-attention module, namely Multi-Head Self-Attention, is a commonly used technical means in the art. Usually, 8 or 16 attention heads are used to process the input audio embedding vectors in parallel, and each attention head focuses on the global dependencies between audio frames in different ways.
[0059] Among them, the multi-head self-attention module MHSA uses multiple attention heads to process the input audio embedding vectors in parallel, and calculates the attention weights of the audio embedding vectors using the following formula: Among them, the query vector Q = E ac W Q , the key vector K = E ac W K , the value vector V = E ac W V , W Q , W K and W V are linear projection matrices, d k is the dimension of the attention head, and the softmax function is used for normalization; the outputs of each attention head are concatenated, and the feature vector T of the audio is obtained through a linear transformation ac , T ac = Concat(head1, head2, head h )W o , where the subscripts 1, 2,, h are the distinguishing marks of different attention heads head, and W o is the merged linear transformation matrix, and Concat is the merge function.
[0060] S13: Send the feature vector T ac into the average pooling layer Average Pooling Layer to obtain the compressed audio feature vector T a , T a = AveragePooling(T ac ).
[0061] In order to retain sufficient feature information, as a preferred implementation, in the present invention, the vector representation of each frame of audio has a dimension of 512, and the compressed audio feature vector T a also has a dimension of 512; the average value of each dimension of the feature vector T of the audio ac is taken and compressed into a 512-dimensional vector, that is, the operation result of the average pooling layer. This step effectively reduces the feature dimension while retaining the global speech information. The obtained compressed audio feature vector T a can represent implicit information such as intonation, emotion, and speech rate contained in the audio data, providing high-quality input for subsequent multi-modal feature fusion.
[0062] In addition, the dimension of the text summary feature vector is 768; the dimension of the embedding vector corresponding to the summary of the text summary is 768; the dimension of the joint representation vector is 512.
[0063] Step S1 of the present invention makes full use of the unstructured data characteristics of the earnings conference call audio to extract high-dimensional and information-intensive audio features, thus providing important support for the financial risk prediction model.
[0064] S2: Extract the feature vectors of the text minutes of the target company's earnings conference call.
[0065] The specific content of step S2 is as follows:
[0066] S21: Preprocess the text minutes of the target company's earnings conference call to obtain a set of sentences where t c L represents the L-th sentence in the text minutes, L = 1, 2,..., l, and l represents the number of sentences in the text minutes; use the Sentence-BERT pre-trained language model, abbreviated as SBERT, to map each sentence to an embedding vector of 768 dimensions for the text minutes of the earnings conference call to obtain the vector representation of the entire text minutes of the earnings conference call SBERT is a sentence-level semantic representation model based on the BERT architecture. It optimizes the ability to capture semantic similarity between sentences through contrastive learning of sentence pairs in a Siamese network.
[0067] Among them, the content of text preprocessing is to standardize the text minutes of the target company's earnings conference call, clean up redundant characters, unify the format, remove meaningless stop words such as um, uh, etc., and then divide the sentences according to punctuation marks such as periods and semicolons in the text minutes of the target company's earnings conference call to obtain a set of sentences.
[0068] S22: Feed the vector representation of the entire text minutes of the earnings conference call into the multi-head self-attention module MHSA to further extract the feature vector T of the text minutes tc , T tc = MHSA(E tc ); The multi-head self-attention module MHSA mentioned here is similar to the content in the aforementioned step S12, but its role is to capture the global dependencies between sentences. It uses multiple attention heads, each head focusing on different semantic relationships. After concatenating the outputs of each attention head, through a linear transformation, the feature vector T of the text minutes with enhanced context features is generated tc .
[0069] S23: Feed the feature vector T of the text minutes tc , into the average pooling layer Average Pooling Layer to obtain the compressed feature vector T of the text minutes t = AveragePooling(T tc ).
[0070] For the feature vector T of the text minutes tcThe average value of each dimension feature is taken and compressed into a 768-dimensional text summary feature vector, which represents the comprehensive semantic information of the text summary of the entire earnings conference call, fully retaining the semantics and context relationships in the text, including key financial information, semantic sentiment, and logical associations between sentences, providing high-quality text input for subsequent multi-modal feature fusion. The present invention has significant advantages in capturing the implicit text information of the text summary of the earnings conference call by combining context features, providing an important data basis for financial risk prediction.
[0071] S3: Use a large language model to summarize the text summary of the earnings conference call of the target company to obtain the embedding vector corresponding to the summary of the text summary; in this step, by segmenting and summarizing the text summary of the earnings conference call, the core content is extracted, and the embedding vector corresponding to the summary of the text summary is generated for subsequent use.
[0072] The specific content of step S3 is as follows:
[0073] S31: Paragraph segmentation, segment the text summary of the earnings conference call of the target company according to logical paragraphs to obtain paragraphs p M = p1, p2,, p m , M = 1, 2,, m, where the subscript m represents the number of paragraphs; input each paragraph p M into the large language model LLM to extract the paragraph-level summary s M = LLM(p M ), s M is the summary of paragraph p M .
[0074] S32: Combine all the paragraph summaries {s1, s2, s M} into an overall text, input it into the large language model to generate a comprehensive summary S, S = LLM({s1, s2, s M}); after the comprehensive summary is generated, the granularity of the summary can be further adjusted according to needs, such as controlling the length of the summary of a longer text or removing redundant information.
[0075] S33: Use the Sentence-BERT pre-trained language model to vectorize the comprehensive summary S to generate the embedding vector T l , T l = SBERT(S). The embedding vector T l corresponding to the summary of the text summary is a 768-dimensional feature vector, representing the semantic information of the text summary. This vector integrates the core content and global information of the earnings conference text, providing important support for multi-modal feature fusion.
[0076] Through the above steps, after the paragraph segmentation, paragraph-level summarization, and comprehensive summarization of the earnings conference call transcript, the generated embedding vector T l effectively summarizes the key content of the text and can significantly improve the utilization efficiency and prediction accuracy of the financial risk prediction model for text data.
[0077] S4: Extract the news text features of the target company to obtain the news text feature vector. In this step, through semantic analysis and similarity calculation of the news text related to the target company, the news text feature vector before the target trading date is generated.
[0078] The specific content of step S4 is as follows:
[0079] S41: Collect the news text data of the target company in the several days before the target trading date, denoted as N = {n1, n2, n k}, where n K0 represents the K0th news text, and k is the total number of news.
[0080] S42: Use the large language model LLM to parse each news n K0 and extract the metadata m K0 , m K0 = LLM(n K0 ); The extracted metadata includes sentiment tendency (positive, negative or neutral), financial indicators involved in the news (such as revenue, profit, etc.), and other key information related to the target company. Integrate the metadata of all news into an overall metadata set M N = {m1, m2,, m k}.
[0081] S43: Find the k historical news groups N K0 = N1, N2,, N k in the historical news dataset that are most similar to the metadata of the news text data N in the several days before the target trading date, calculate the semantic relevance between N and any historical news group N K0 , where f is a function that converts news text into an embedding vector, and select a group of historical news H with the highest semantic similarity to the news text data N in the several days before the target trading date.
[0082] S44: Concatenate N, H, and the text related to the market trend after H occurs, and then use the SBERT model to convert the concatenated text into an embedding vector T n , and the embedding vector T n is used as the news text feature vector of the target company in the several days before the target trading date.
[0083] Here, there are texts from three sources, namely: 1. News text data N in the days before the target transaction date; 2. Historical news H; 3. Texts related to the market trend after the occurrence of historical news H, that is, news after the occurrence of historical news H. For example, the historical news H on January 1 mentions that Company A has a debt crisis, and a certain news on January 2 mentions that the stock price of Company A has plummeted. The news on January 2 is the text related to the market trend after the occurrence of historical news H.
[0084] Through the above step S4, the news text feature vector T of the target company in the days before the target transaction date is generated. n Effectively integrates the semantic information, sentiment analysis results of the target news, and the similarity relationship with historical news, providing high-quality input features for subsequent multi-modal feature fusion.
[0085] S5: Extract features from the time-series trading data of the target company in a period before the target date to obtain the feature vector of the time-series data. As a preferred implementation method, in this step, the time-series trading data of the target company in the 30 days before the target date is processed to extract key features and capture the dynamic relationship between different time-series features.
[0086] The specific content of step S5 is as follows:
[0087] S51: Collect the time-series trading data D of the target company in the 30 days before the target date, including the closing price and trading volume of each day, expressed as D = {(p1, v1), (p2, v2),...,(p d , v d )}, where p F = p1, p2,..., p d represents the closing price on the Fth day, and v F = v1, v2,..., v d represents the trading volume on the Fth day;
[0088] S52: Input the time-series trading data D into the bidirectional long short-term memory network Bi-LSTM. The bidirectional long short-term memory network Bi-LSTM captures the time-series dynamic characteristics of the trading data and outputs a 128-dimensional feature vector T v , T v = BiLSTM(D); The bidirectional long short-term memory network Bi-LSTM models the time-series data and captures the time-series dynamic characteristics of the trading data.
[0089] S53: Capture the dynamic relationship between different time-series features through the vector autoregressive VAR model:
[0090] log(σ 3,t ) = α3 + β 1,1 log(σ -3,t ) + β1,2 log(σ -7,t ) + β 1.3 log(σ -15,t ) + β 1.4. log(σ -30,t ) + u 3,t ;
[0091] log(σ 7,t ) = α7 + β 2,1 log(σ -3,t ) + β 2,2 log(σ -7,t ) + β 2.3 log(σ -15,t ) + β 2.4. log(σ -30,t ) + u 7,t ;
[0092] log(σ 15,t ) = α 15 + β 3,1 log(σ -3,t ) + β 3,2 log(σ -7,t ) + β 3.3 log(σ -15,t ) + β 3.4. log(σ -30,t ) + u 15,t ;
[0093] log(σ 30,t ) = α 30 + β 4,1 log(σ -3,t ) + β 4,2 log(σ -7,t ) + β 4.3 log(σ -15,t ) + β 4.4. log(σ -30,t ) + u 30,t ;
[0094] σ z,t represents the volatility of the target company's stock price within z = 3, 7, 15, and 30 days, where z ≤ F and z ∈ t; u z,t is a white noise term; β a,b is the coefficient matrix of the dynamic relationship, a, b = 1, 2, 3, 4; α z is the intercept term; the volatility of the stock price is defined as the standard deviation of the return rate of the target company within z days: Here, the values of z for the number of days are 3, 7, 15, and 30, which are only for illustration purposes and not a limitation of the solution itself.
[0095] Through this step S5, the time-series transaction data is converted into a high-dimensional feature vector T v and the parameters of the dynamic relationship, providing key time-series information and dynamic relationship support for financial risk prediction.
[0096] S6: Fuse the obtained audio feature vector, text summary feature vector, embedding vector corresponding to the summary of the text summary, news text feature vector, and feature vector of the time-series data to obtain a joint representation vector.
[0097] The joint representation vector E described in step S6 is fused through the following formula: E = w0 + w1T a + w2T t + w3T l + w4T n + w5T v + ε, where w0 is the bias term, w1, w2, w3, w4, w5 are the fusion weights, and ε is the error term, representing random noise.
[0098] The dimension of the joint representation vector E is fixed at 512 dimensions, representing the comprehensive features of the multi-modal data. This vector retains the important information of each modal feature and can be used for subsequent multi-task prediction.
[0099] S7: Input the joint representation vector into the multi-task learning framework for risk index prediction.
[0100] Specifically, use the joint representation vector E generated in step S6 as the input of the multi-task learning framework. The multi-task learning framework simultaneously predicts the following risk indices: the volatility σ 3,t 、σ 7,t 、σ 15,t and σ 30,t of the stock price of the target company within time spans of 3 days, 7 days, 15 days, and 30 days, as well as the single-day risk value VAR;
[0101] For the volatility of the stock price for each time span, construct an independent first prediction sub-network MLP. The prediction result of the volatility of the stock price is where is the predicted volatility; f MLP,z (·) is the first prediction sub-network MLP for the time span z;
[0102] For the single-day risk value VAR, construct an independent second prediction sub-network MLP. The input is the joint representation vector E, and the output is the predicted value of VAR where is the predicted value of VAR; f MLP,VAR (·) is the second prediction sub-network MLP;
[0103] The combined loss function simultaneously optimizes the stock price volatility and VAR prediction tasks, and the formula for the combined loss function is as follows: μ is a weight hyperparameter that balances the prediction error of stock price volatility and the prediction error of VAR; y j and represent the true and predicted values of the stock price volatility indicator; represents the mean squared error of the stock price volatility prediction; q represents the quantile threshold of the single-day risk value; V and represent the true single-day risk value and the predicted single-day risk value respectively; represents the quantile regression loss function of the single-day risk value prediction task.
[0104] Through the above step S7, the multi-task learning framework can effectively integrate multi-modal features, accurately predict the volatility and single-day risk value of different time spans, and provide reliable support for financial risk management.
[0105] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A financial risk prediction method based on multi-modal data fusion, characterized in that, It includes the following steps: S1: Extract the audio feature vectors of the earnings conference call of the target company; S2: Extract the text summary feature vectors of the earnings conference call of the target company; S3: Use a large language model to summarize the text summary of the earnings conference call of the target company to obtain the embedding vectors corresponding to the summary of the text summary; S4: Extract the news text features of the target company to obtain the news text feature vectors; S5: Extract the features of the time-series transaction data of the target company within a period before the target date to obtain the feature vectors of the time-series data; S6: Fuse the obtained audio feature vectors, text summary feature vectors, embedding vectors corresponding to the summary of the text summary, news text feature vectors, and feature vectors of the time-series data to obtain the joint representation vectors; and S7: Input the joint representation vectors into a multi-task learning framework for risk indicator prediction.
2. The financial risk prediction method based on multi-modal data fusion according to claim 1, wherein The specific content of step S1 is as follows: S11: Use the WenetSpeech pre-trained language model to extract the embedding vectors of the earnings conference call audio. For a segment of audio after audio preprocessing where represents the i-th frame of the audio, n represents the number of frames in the audio, and each frame of audio is converted into a vector representation to obtain the entire audio A c of the audio embedding vector S12: Embed the audio embedding vector E ac into the multi-head self-attention module MHSA to further extract the feature vector T of the audio ac = MHSA(E ac ); S13: Feed the feature vector T ac into the average pooling layer Average Pooling Layer to obtain the compressed audio feature vector T a , T a = AveragePooling(T ac ).
3. A financial risk prediction method based on multi-modal data fusion according to claim 2, characterized in that, The vector representation of each audio frame has a dimension of 512, and the compressed audio feature vector T a also has a dimension of 512; the dimension of the text summary feature vector is 768; the dimension of the embedding vector corresponding to the summary of the text summary is 768; the dimension of the joint representation vector is 512.
4. A financial risk prediction method based on multi-modal data fusion according to claim 2, characterized in that, The multi-head self-attention module MHSA uses multiple attention heads to process the input audio embedding vectors in parallel, and calculates the attention weights of the audio embedding vectors using the following formula: where the query vector Q = E ac W Q , the key vector K = E ac W K , the value vector V = E ac W V , W Q , W K and W V are linear projection matrices, d k is the dimension of the attention head, and the softmax function is used for normalization; the outputs of each attention head are concatenated, and the feature vector T of the audio is obtained through a linear transformation ac , T ac = Concat(head1,head2,...head h )W o , where the subscripts 1,2,...,h are the distinguishing marks of different attention heads head, and W o is the linear transformation matrix after merging, and Concat is the merging function.
5. A financial risk prediction method based on multi-modal data fusion according to claim 4, characterized in that, The specific content of step S2 is as follows: S21: Preprocess the text minutes of the earnings conference call of the target company to obtain a set of sentences where represents the L-th sentence in the text minutes, L = 1, 2,..., l, and l represents the number of sentences in the text minutes; use the Sentence-BERT pre-trained language model to map each sentence t c L to the embedding vector of the earnings conference call text minutes to obtain the vector representation of the entire text minutes of the earnings conference call S22: Feed the vector representation of the text record of the entire earnings conference call into the multi-head self-attention module MHSA to further extract the feature vector T of the text record tc , T tc = MHSA(E tc ); S23: Feed the feature vector T of the text summary tc , into the average pooling layer Average Pooling Layer to obtain the compressed feature vector T of the text summary t = AveragePooling(T tc ).
6. A financial risk prediction method based on multi-modal data fusion according to claim 5, characterized in that, The specific content of step S3 is as follows: S31: Paragraph segmentation. Segment the text record of the earnings conference call of the target company into logical paragraphs to obtain paragraphs p M = p1, p2,..., p m , where M = 1, 2,..., m, and the subscript m represents the number of paragraphs; input each paragraph p M into the large language model LLM to extract the paragraph-level summary s M = LLM(p M ), and s M is the summary of paragraph p M . S32: Combine all the paragraph summaries {s1, s2,... s M} into the overall text and input it into the large language model to generate a comprehensive summary S, S = LLM({s1, s2,... s M}); S33: Vectorize the comprehensive summary S using the Sentence-BERT pre-trained language model to generate the embedding vector T corresponding to the summary of the text minutes l , T l = SBERT(S).
7. A financial risk prediction method based on multi-modal data fusion according to claim 6, characterized in that The specific content of step S4 is as follows: S41: Collect the news text data of the target company in the several days before the target transaction date, denoted as N = {n1, n2,... n k}, where n K0 represents the K0th news text, and k is the total number of news; S42: Use the large language model LLM to parse each news item n K0 and extract the metadata m K0 , where m K0 = LLM(n K0 ); Integrate the metadata of all news items into a single overall metadata set M N = {m1, m2,..., m k}; S43: Find k historical news groups N that are most similar in terms of metadata to the news text data N of several days before the target trading date in the historical news dataset K0 = N1, N2,..., N k , calculate the semantic relevance between N and any historical news group N K0 , where f is a function that converts news text into an embedding vector, and select a group of historical news H with the highest semantic similarity to the news text data N of several days before the target trading date; S44: Concatenate the text related to the market trends after N, H, and H occur, and then use the SBERT model to convert the concatenated text into an embedding vector T n , the embedding vector T n is used as the news text feature vector of the target company in the days before the target transaction date.
8. A financial risk prediction method based on multi-modal data fusion according to claim 7, characterized in that The specific content of step S5 is as follows: S51: Collect the time-series transaction data D of the target company in the 30 days before the target date, including the daily closing price and trading volume, expressed as D = {(p1, v1), (p2, v2),..., (p d , v d )}, where p F = p1, p2,..., p d represents the closing price on the F-th day, and v F = v1, v2,..., v d represents the trading volume on the F-th day; S52: Input the time-series transaction data D into the Bidirectional Long Short-Term Memory Network (Bi-LSTM). The Bi-LSTM captures the time-series dynamic characteristics of the transaction data and outputs a feature vector T containing time-series data v , T v = BiLSTM(D); S53: Capture the dynamic relationship between different time-series features through a vector autoregressive VAR model: log(σ 3,t ) = α3 + β 1,1 log(σ -3,t ) + β 1,2 log(σ -7,t ) + β 1.3 log(σ -15,t ) + β 1.
4. log(σ -30,t ) + u 3,t ; log(σ 7,t ) = α7 + β 2,1 log(σ -3,t ) + β 2,2 log(σ -7,t ) + β 2.3 log(σ -15,t ) + β 2.
4. log(σ -30,t ) + u 7,t ; log(σ 15,t ) = α 15 + β 3,1 log(σ -3,t ) + β 3,2 log(σ -7,t ) + β 3.3 log(σ -15,t ) + β 3.
4. log(σ -30,t ) + u 15,t ; log(σ 30,t ) = α 30 + β 4,1 log(σ -3,t ) + β 4,2 log(σ -7,t ) + β 4.3 log(σ -15,t ) + β 4.
4. log(σ -30,t ) + u 30,t ; σ z,t represents the volatility of the target company's stock price within z = 3, 7, 15, and 30 days, where z ≤ F and z ∈ t; u z,t is the white noise term; β a,b is the coefficient matrix of the dynamic relationship, a, b = 1, 2, 3, 4; α z is the intercept term; the volatility of the stock price is defined as the standard deviation of the return rate of the target company within z days:
9. A financial risk prediction method based on multi-modal data fusion according to claim 8, characterized in that, The combined representation vector E described in step S6 is fused through the following formula: E = w0 + w1T a + w2T t + w3T l + w4T n + w5T v + ε, where w0 is the bias term, w1, w2, w3, w4, w5 are the fusion weights, and ε is the error term representing random noise.
10. A financial risk prediction method based on multi-modal data fusion according to claim 9, characterized in that, Specifically, in step S7, the combined representation vector E generated in step S6 is used as the input to the multi-task learning framework, which simultaneously predicts the following risk metrics: the volatility σ 3,t , σ 7,t , σ 15,t , and σ 30,t of the target company's stock price over time spans of 3 days, 7 days, 15 days, and 30 days, as well as the single-day risk value VAR; For the stock price volatility of each time span, an independent first prediction sub-network MLP is constructed, and the prediction result of the stock price volatility is where is the predicted volatility; f MLP,z (·) is the first prediction sub-network MLP for the time span z; Construct an independent second prediction sub-network MLP for the single-day risk value VAR. The input is the joint representation vector E, and the output is the predicted value of VAR. where is the predicted value of VAR; f MLP,VAR (·) is the second prediction sub-network MLP; The joint loss function simultaneously optimizes the stock price volatility and VAR prediction tasks, and the formula for the joint loss function is: μ is a weight hyperparameter that balances the prediction error of stock price volatility and the prediction error of VAR; y j and represent the true and predicted values of the stock price volatility indicator; represents the mean squared error of the stock price volatility prediction; q represents the quantile threshold of the single-day value at risk; V and represent the true single-day value at risk and the predicted single-day value at risk, respectively; represents the quantile regression loss function for the single-day value at risk prediction task.
Citation Information
Patent Citations
Financial risk prediction method and system based on representation learning
CN116402630A
Multi-feature stock trend prediction method fusing shareholder emotion and stock event
CN116541484A
Financial risk prediction method, device, equipment, medium and program product
CN117011080A
Telephone recording processing method and system based on epidemic situation traffic scheduling process
CN117095703A
Enterprise financial risk prediction method based on multi-source data fusion
CN118780615A
Cited By
Intelligent financial risk early warning method, system and device and storage medium
CN121032676A
Intelligent financial risk early warning method, system, device and storage medium
CN121032676B