Investor emotion divergence quantitative index construction system
Through multi-level data fusion and dynamic weight adjustment mechanism, investor sentiment data is monitored and processed in real time, which solves the problem of insufficient real-time tracking capability of quantitative analysis of sentiment divergence in existing technologies and realizes efficient extraction and rapid response of sentiment features.
Patent Information
- Application Number
- CN202510779553.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have limited real-time tracking capabilities in quantitative analysis of investor sentiment divergence, which affects the timeliness and accuracy of decision-making.
It adopts a multi-level data fusion module, a dynamic weight adjustment mechanism and a real-time feedback unit, obtains sentiment data through distributed crawler technology, uses natural language processing and convolutional neural networks for feature extraction, combines time series weighted aggregation and adaptive weight calculation, monitors and adjusts sentiment feature weights in real time, and uses an improved isolation forest algorithm to identify abnormal fluctuations.
It significantly improves the accuracy and comprehensiveness of sentiment feature extraction, ensures that the sentiment divergence calculation results reflect market changes in a timely manner, and improves the system's response speed and decision-making accuracy.
Smart Images

Figure CN120765384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of financial technology and data analysis technology, and specifically to a system for constructing quantitative indicators for investor sentiment divergence. Background Art
[0002] In the field of financial investment, the analysis and quantification of investor sentiment is an important part of market research. Investor sentiment generally reflects the expectations and attitudes of market participants towards specific assets or the overall market, and its fluctuations can have a significant impact on market prices. However, investor sentiment is not completely consistent, and there may be differences between different investors. Quantifying this degree of disagreement is of great significance for understanding market dynamics, predicting price trends, and formulating investment strategies.
[0003] Reference patent announcement number "CN119004339B" discloses an artificial intelligence-based financial data management method and system, including: collecting financial big data and user behavior data of e-commerce enterprises in a preset historical time period through a data terminal; extracting features from user behavior data and importing a preset clustering model to perform user clustering grouping. The clustering process is based on the feature similarity of user behavior data and forms multiple user groups; for multiple user groups, the financial big data is grouped accordingly to obtain multiple financial data sets.
[0004] As shown in the above-mentioned technologies, existing technologies use cluster analysis and association algorithms to explore the correlation between financial and user behavior characteristics, and use the isolation forest algorithm to detect abnormal data. However, existing technical solutions focus on the classification and anomaly detection of financial data, and do not involve quantitative analysis of investor sentiment divergence. Its processing of emotional characteristics mainly relies on predefined rules and models. When responding to dynamic changes in emotional divergence, its real-time tracking capabilities have certain limitations, which may affect the response speed to changes in investor sentiment, thereby reducing the timeliness and accuracy of decision-making to a certain extent. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the present invention provides a system for constructing a quantitative index of investor sentiment divergence, which solves the problem of limited real-time tracking capability of the existing system for constructing a quantitative index of investor sentiment divergence.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A system for constructing a quantitative indicator of investor sentiment divergence includes a multi-level data fusion module, a dynamic weight adjustment mechanism, and a real-time feedback unit, wherein:
[0007] The multi-level data fusion module includes a basic emotion data acquisition unit, a multi-source information integration layer, and a feature extraction layer, which is used to obtain raw emotion data from multiple data sources and generate emotion feature subsets;
[0008] The dynamic weight adjustment mechanism comprises an adaptive weight calculation unit and a feedback correction loop, which are used for dynamically adjusting the weight coefficients of the subset of emotional features according to market fluctuation indexes and trading volume data;
[0009] The real-time feedback unit comprises a data buffer area, a fast calculation channel and an anomaly detector, which are used for calculating the emotional divergence and identifying abnormal emotional fluctuations.
[0010] Preferably, the basic emotional data acquisition unit acquires original emotional data from social media, news platforms and financial forums through distributed crawler technology, and performs word segmentation, denoising and sentiment labeling operations on the text by using natural language processing technology.
[0011] Preferably, the multi-source information integration layer adopts a weighted aggregation algorithm based on time series, divides the emotional data from different sources into multiple time periods according to the time dimension, and generates a preliminary emotional feature vector.
[0012] Preferably, the multi-source information integration layer adopts a weighted aggregation algorithm based on time series, divides the emotional data from different sources into multiple time periods according to the time dimension, and generates a preliminary emotional feature vector.
[0013] Preferably, the adaptive weight calculation unit analyzes the trend of historical emotional data by using a sliding window incremental update algorithm, and calculates the weight coefficient of each emotional feature in combination with market fluctuation indexes and trading volume data.
[0014] Preferably, the feedback correction loop compares the difference between the current emotional feature distribution and the historical benchmark distribution, calculates the deviation value and uses it as a correction factor to adjust the weight allocation scheme.
[0015] Preferably, the data buffer area adopts a double buffering structure design, in which the main buffer area stores the latest emotional data, and the standby buffer area pre-processes the data stream that is about to enter the main buffer area.
[0016] Preferably, the anomaly detector is based on an improved isolation forest algorithm, which can identify abnormal emotional fluctuations and trigger a warning mechanism within milliseconds.
[0017] Advantages
[0018] The present application provides an investor emotional divergence quantification index construction system. Compared with the prior art, the following advantages are achieved:
[0019] 1. This investor sentiment divergence quantification indicator construction system is based on the design and implementation of a multi-level data fusion module. The basic sentiment data collection unit obtains raw data from channels such as social media, news platforms, and financial forums through distributed crawler technology, and uses natural language processing technology to segment, denoise, and label text. The multi-source information integration layer adopts a weighted aggregation algorithm based on time series to integrate sentiment data from different sources in segments according to the time dimension to generate preliminary sentiment feature vectors. The feature extraction layer uses convolutional neural networks to deeply mine the integrated sentiment features, significantly improving the accuracy and comprehensiveness of sentiment feature extraction.
[0020] 2. The investor sentiment divergence quantitative indicator construction system uses a dynamic weight adjustment mechanism, which includes an adaptive weight calculation unit and a feedback correction loop. The adaptive weight calculation unit analyzes the changing trend of historical sentiment data and combines market volatility index and trading volume data to dynamically adjust the weight coefficient of each sentiment feature. The feedback correction loop automatically corrects the weight distribution scheme by real-time monitoring the deviation between the current sentiment feature distribution and the historical pattern, ensuring that the calculation results of the sentiment divergence can reflect market changes in a timely manner.
[0021] 3. The investor sentiment divergence quantitative indicator construction system is used to improve the system's response speed and accuracy through a real-time feedback unit. The real-time feedback unit includes a data cache, a fast calculation channel and an anomaly detector. The data cache adopts a double-buffer structure design. The main cache is responsible for storing the latest sentiment data, and the backup cache is used to preprocess the data stream that is about to enter the main cache. The fast calculation channel accelerates the calculation process of sentiment divergence through a parallel computing architecture. It has multiple independent computing nodes inside, each of which is responsible for processing a specific type of sentiment feature. The anomaly detector is based on an improved isolation forest algorithm and can identify abnormal sentiment fluctuations within milliseconds and trigger an early warning mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Schematic diagram of the overall architecture of the system of the present invention;
[0023] Figure 2 This is a structural diagram of the multi-level data fusion module of the present invention. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0025] Please refer to Figure 1-2 The present application provides three technical solutions:
[0026] The first embodiment includes a multi-level data fusion module, a dynamic weight adjustment mechanism, and a real-time feedback unit, wherein:
[0027] The multi-level data fusion module includes a basic emotional data acquisition unit, a multi-source information integration layer, and a feature extraction layer, which is used to obtain raw emotional data from multiple data sources and generate a subset of emotional features. The basic emotional data acquisition unit obtains raw emotional data from social media, news platforms, and financial forums through distributed crawler technology, and uses natural language processing technology to perform word segmentation, noise removal, and sentiment labeling operations on the text. The multi-source information integration layer uses a weighted aggregation algorithm based on time series to divide emotional data from different sources into multiple time periods according to the time dimension and generate a preliminary emotional feature vector. The basic emotional data acquisition unit uses distributed crawler technology to obtain raw emotional data from multiple channels such as social media, news platforms, and financial forums. These data sources include but are not limited to Weibo, Twitter, financial news websites, and stock discussion communities. In the basic emotional data acquisition unit, the distributed crawler crawls web page content at a preset time interval and stores the crawled data in the local cache area. Subsequently, natural language processing technology is used to perform word segmentation, noise removal, and sentiment labeling operations on the text. The word segmentation process is based on Chinese or English word segmentation tools. Noise removal is achieved by filtering out irrelevant characters, punctuation marks, and stop words. Sentiment labeling relies on a pre-trained sentiment analysis model that can assign a positive, negative, or neutral sentiment label to each piece of text. After the above processing, the basic emotional data acquisition unit generates a preliminary emotional data set and transmits it to the multi-source information integration layer.
[0028] The feature extraction layer further mines the preliminary emotional feature vector to extract a subset of emotional features with discriminative power. The feature extraction layer uses an improved convolutional neural network structure, in which an attention mechanism module is introduced to enhance the ability to capture key emotional features. The input of the convolutional neural network is the emotional feature vector generated by the multi-source information integration layer. After multiple convolution operations, the network outputs a set of high-dimensional feature vectors. The attention mechanism module performs weighted operations on these feature vectors to highlight important features and suppress irrelevant features, thereby improving the accuracy of feature extraction. Finally, the feature extraction layer generates a refined subset of emotional features and passes it to the dynamic weight adjustment mechanism.
[0029] The multi-source information integration layer receives the emotion data set from the basic emotion data collection unit, and integrates the emotion data from different sources in segments through a weighted aggregation algorithm based on time series. For example, the multi-source information integration layer first divides the received emotion data into multiple time periods according to the time dimension. Each time period corresponds to an emotion feature vector. In order to ensure the accuracy of the integration results, emotion data from different sources are assigned different weight coefficients. These weight coefficients are derived from historical data analysis and are regularly updated to reflect the latest data distribution characteristics. The specific implementation steps of the weighted aggregation algorithm include calculating the standard deviation, mean and correlation parameters of the data from each source, and generating weighting factors based on these parameters. Finally, the multi-source information integration layer outputs a set of preliminary emotion feature vectors and passes them to the feature extraction layer.
[0030] The second embodiment is mainly different from the first embodiment in that: the dynamic weight adjustment mechanism includes an adaptive weight calculation unit and a feedback correction loop, which are used to dynamically adjust the weight coefficients of the sentiment feature subset according to the market volatility index and trading volume data. The adaptive weight calculation unit analyzes the changing trend of the historical sentiment data through a sliding window incremental update algorithm, and calculates the weight coefficient of each sentiment feature in combination with the market volatility index and trading volume data. The dynamic weight adjustment mechanism realizes dynamic adjustment of the sentiment feature weight through the adaptive weight calculation unit and the feedback correction loop. The adaptive weight calculation unit receives the sentiment feature subset from the feature extraction layer, and calculates the weight coefficient of each sentiment feature in combination with the market volatility index and trading volume data. The market volatility index and trading volume data are passed through the adaptive weight calculation unit. The data is acquired in real time through an external interface. The adaptive weight calculation unit uses a sliding window incremental update algorithm to analyze the changing trend of historical sentiment data, thereby dynamically adjusting the weight coefficient. The size of the sliding window can be set according to actual needs. For example, it can be set to sentiment data within the last 100 trading days. The incremental update algorithm significantly reduces the computational complexity by updating only the newly added data each time. The feedback correction loop monitors the deviation between the current sentiment feature distribution and the historical pattern in real time, and automatically corrects the weight distribution scheme according to the degree of deviation. Specifically, the feedback correction loop calculates the deviation value by comparing the difference between the current sentiment feature distribution and the historical benchmark distribution, and introduces it into the weight adjustment process as a correction factor, thereby ensuring that the calculation result of the sentiment divergence can reflect market changes in a timely manner.
[0031] The third implementation method is mainly different from the second implementation method in that: the real-time feedback unit includes a data cache, a fast calculation channel and an anomaly detector, which is used to calculate the emotional divergence and identify abnormal emotional fluctuations. The feedback correction loop calculates the deviation value by comparing the difference between the current emotional feature distribution and the historical benchmark distribution, and uses it as a correction factor to adjust the weight distribution scheme. The data cache adopts a double buffer structure design. The main cache stores the latest emotional data, and the backup cache preprocesses the data stream that is about to enter the main cache. The anomaly detector is based on the improved isolation forest algorithm, which can identify abnormal emotional fluctuations within milliseconds and trigger an early warning mechanism. The real-time feedback unit improves the response speed and accuracy of the system through the data cache, fast calculation channel and anomaly detector. The data cache adopts a double-buffer structure design. The main cache is responsible for storing the latest emotional data, and the backup cache is used to pre-process the data stream that is about to enter the main cache. This double-buffer design effectively avoids blocking problems in the data processing process and improves data processing efficiency. The fast computing channel accelerates the calculation process of emotional divergence through a parallel computing architecture. It has multiple independent computing nodes inside, and each node is responsible for processing a specific type of emotional features. For example, some nodes focus on processing positive emotional features, while other nodes process negative or neutral emotional features. The anomaly detector is based on the improved isolation forest algorithm and can identify abnormal emotional fluctuations within milliseconds. When abnormal emotional fluctuations are detected, the anomaly detector will trigger an early warning mechanism and pass relevant information to the user interface.
[0032] The entire system is deployed on a distributed computing platform. Modules exchange data through a high-speed data bus and a ring buffer. The multi-level data fusion module is connected to the dynamic weight adjustment mechanism through a high-speed data bus. The data transmission rate reaches the Gbps level. A ring buffer is used for data exchange between the dynamic weight adjustment mechanism and the real-time feedback unit to ensure the continuity and stability of data transmission. The distributed computing platform optimizes resource utilization through load balancing strategies, such as allocating computing tasks to multiple computing nodes to improve overall performance.
[0033] When the system is running: First, the basic emotion data acquisition unit obtains the original emotion data from multiple data sources and preprocesses it to generate a preliminary emotion data set. Then, the multi-source information integration layer performs weighted aggregation on the received emotion data set to generate a preliminary emotion feature vector. The feature extraction layer conducts in-depth mining on the preliminary emotion feature vector, extracts a subset of emotion features with discriminability and passes it to the dynamic weight adjustment mechanism. The adaptive weight calculation unit calculates the weight coefficient of each emotion feature based on the market volatility index and trading volume data. The feedback correction loop automatically corrects the weight distribution scheme according to the deviation between the current emotion feature distribution and the historical pattern. Finally, the real-time feedback unit completes the calculation of emotion divergence and anomaly detection through the data cache, fast calculation channel and anomaly detector, and displays the results to the user in real time.
[0034] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0035] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A system for constructing quantitative indicators for investor sentiment divergence, characterized by: It includes a multi-level data fusion module, a dynamic weight adjustment mechanism and a real-time feedback unit, among which: The multi-level data fusion module includes a basic emotion data acquisition unit, a multi-source information integration layer, and a feature extraction layer, which is used to obtain raw emotion data from multiple data sources and generate emotion feature subsets; The dynamic weight adjustment mechanism includes an adaptive weight calculation unit and a feedback correction loop, which is used to dynamically adjust the weight coefficients of the sentiment feature subset based on the market volatility index and trading volume data; The real-time feedback unit includes a data buffer, a fast calculation channel and an anomaly detector, which is used to calculate the emotional divergence and identify abnormal emotional fluctuations.
2. The system for constructing a quantitative indicator of investor sentiment divergence according to claim 1, characterized in that: The basic emotion data collection unit obtains original emotion data from social media, news platforms and financial forums through distributed crawler technology, and uses natural language processing technology to perform word segmentation, denoising and emotion labeling operations on the text.
3. The system for constructing a quantitative indicator of investor sentiment divergence according to claim 1, characterized in that: The multi-source information integration layer adopts a weighted aggregation algorithm based on time series to divide the emotion data from different sources into multiple time periods according to the time dimension and generate preliminary emotion feature vectors.
4. The system for constructing a quantitative indicator of investor sentiment divergence according to claim 1, characterized in that: The multi-source information integration layer adopts a weighted aggregation algorithm based on time series to divide the emotion data from different sources into multiple time periods according to the time dimension and generate preliminary emotion feature vectors.
5. The system for constructing a quantitative indicator of investor sentiment divergence according to claim 1, characterized in that: The adaptive weight calculation unit analyzes the changing trend of historical sentiment data through a sliding window incremental update algorithm, and calculates the weight coefficient of each sentiment feature in combination with the market volatility index and trading volume data.
6. The system for constructing a quantitative indicator of investor sentiment divergence according to claim 1, characterized in that: The feedback correction loop calculates the deviation value by comparing the difference between the current emotional feature distribution and the historical benchmark distribution, and uses it as a correction factor to adjust the weight distribution scheme.
7. The system for constructing a quantitative indicator of investor sentiment divergence according to claim 1, characterized in that: The data buffer area adopts a double buffer structure design, the main buffer area stores the latest emotion data, and the backup buffer area preprocesses the data flow that is about to enter the main buffer area.
8. The system for constructing a quantitative indicator of investor sentiment divergence according to claim 1, characterized in that: The anomaly detector is based on an improved isolation forest algorithm, which can identify abnormal emotional fluctuations within milliseconds and trigger an early warning mechanism.
Citation Information
Patent Citations
A financial data management method and system based on artificial intelligence
CN119004339B
Cited By
Text data preprocessing method suitable for large financial model
CN121561264A