Quantitative strategy analysis method and system based on multi-modal data
By introducing multimodal data processing methods in quantitative strategy analysis, using the combination of text and picture data to independently learn emotional weights, the subjectivity and deviation problems caused by relying on emotional dictionaries in the existing technology are solved, and a more efficient and objective quantitative strategy analysis is achieved.
Patent Information
- Application Number
- CN202510473817.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing quantitative strategy analysis methods rely on emotional dictionaries, and the results are subjective and prone to deviations, and cannot effectively utilize multimodal data.
A quantitative strategy analysis method based on multimodal data is proposed. By obtaining text and picture data of market dynamic news, using text regression model and deep neural network emotion recognition model, we can independently learn word emotional weights, dynamically adapt to different contexts, and conduct strategy analysis based on text and picture data.
It improves the generalization ability of sentiment analysis, can analyze the information conveyed by news data more comprehensively, reduces dependence on emotional dictionaries, and reduces subjective bias.
Smart Images

Figure CN119989161A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text analysis and computer vision technology, and in particular to a quantitative strategy analysis method and system based on multimodal data. Background Art
[0002] With the popularization of computers, the Internet and mobile terminals, more and more netizens browse market dynamics news, corporate annual reports and other information on online information platforms. Market dynamics news media data has application value in network trend analysis and prediction. Due to the excellent performance of machine learning in unstructured data processing, machine learning is widely used in the current research on quantitative strategies in the financial market. Machine learning in the field of unstructured data processing includes computer vision, natural language processing and speech recognition. Thanks to the rapid development of online media, the volume of market dynamics-related data on the Internet has become larger, and there is a huge amount of data available. Relying on the advantages of machine learning algorithms in processing text and image data, the research value of these data in the quantitative field can be explored.
[0003] However, most existing methods are based on sentiment dictionaries, which pre-set sentiment dictionaries and sentiment scores of corresponding words to measure text scores. The results obtained in this way are highly dependent on dictionaries. The results obtained based on human experience are highly subjective and prone to deviations. Summary of the invention
[0004] In view of the above situation, the main purpose of the present invention is to propose a quantitative strategy analysis method and system based on multimodal data to solve the above technical problems.
[0005] The present invention proposes a quantitative strategy analysis method based on multimodal data, which comprises the following steps: Step 1: Obtain text data of market dynamic news, and clean the text data to obtain cleaned text data; Determine learning objectives using the settlement benchmark values before and after the date in the cleaned text data; Step 2: Count the common words in the cleaned text data and build a bag of words for common words; Get the word frequency vector through the bag of common words; Generate a word frequency matrix using word frequency vectors; Take the word frequency matrix as the independent variable and the learning goal as the dependent variable, calculate the Pearson coefficient, and process the word frequency matrix according to the Pearson coefficient to obtain a new word frequency matrix; Reduce the dimension of the new word frequency matrix to obtain the reduced-dimensional word frequency matrix; Step 3: Build a regression model for the text based on the word frequency matrix after dimensionality reduction; Step 4: Obtain image data of market dynamic news, and use the regression model for text to classify the image data to obtain labeled image data; Input the labeled image data into the pre-trained convolutional neural network model to obtain a fine-tuned pre-trained convolutional neural network model; Step 5: Use the image data to train the fine-tuned pre-trained convolutional neural network model to obtain a deep neural network emotion recognition model for images; Step 6: Input the text data and image data of market dynamic news into the regression model for text and the deep neural network emotion recognition model for images respectively to obtain the strategy results.
[0006] The present invention also proposes a quantitative strategy analysis system based on multimodal data, the system comprising: Text data processing module, used for: Acquire text data of market dynamic news, and clean the text data to obtain cleaned text data; Determine learning objectives using the settlement benchmark values before and after the date in the cleaned text data; Matrix building blocks for: Count the common words in the cleaned text data and build a bag of words for common words; Get the word frequency vector through the bag of common words; Generate a word frequency matrix using word frequency vectors; Take the word frequency matrix as the independent variable and the learning goal as the dependent variable, calculate the Pearson coefficient, and process the word frequency matrix according to the Pearson coefficient to obtain a new word frequency matrix; Reduce the dimension of the new word frequency matrix to obtain the reduced-dimensional word frequency matrix; Text model building blocks for: Build a regression model for text based on the reduced word frequency matrix; Image data processing module, used for: Obtain image data of market dynamic news, and use a regression model for text to classify the image data to obtain labeled image data; Input the labeled image data into the pre-trained convolutional neural network model to obtain a fine-tuned pre-trained convolutional neural network model; Image model training module, used for: The fine-tuned pre-trained convolutional neural network model is trained using image data to obtain a deep neural network emotion recognition model for images; Policy Results Module, which is used to: The text data and image data of market dynamic news are input into the regression model for text and the deep neural network emotion recognition model for image respectively to obtain the strategy results.
[0007] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention uses a data-driven approach to allow the model to autonomously learn the sentiment weight values of words, so that the model can dynamically adapt to different contexts and improve the generalization ability of sentiment analysis; 2. The present invention introduces the mining of news picture data. Existing news data mining technologies are mostly based on text. The introduction of pictures can more comprehensively analyze the information conveyed by news data.
[0008] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A flowchart of the quantitative strategy analysis method based on multimodal data proposed by the present invention; Figure 2 A flowchart of a deep neural network emotion recognition model for images based on a quantitative strategy analysis method for multimodal data proposed by the present invention; Figure 3 Schematic diagram of the overall framework of the quantitative strategy analysis system based on multimodal data proposed in the present invention. DETAILED DESCRIPTION
[0010] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0011] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0012] See also Figure 1 The embodiment of the present invention proposes a quantitative strategy analysis method based on multimodal data, which includes the following steps: Step 1: Obtain text data of market dynamic news, and clean the text data to obtain cleaned text data; Determine learning objectives using the settlement benchmark values before and after the date in the cleaned text data; In step 1, the settlement benchmark values before and after the date in the cleaned text data are used to determine the learning target. The relationship between the corresponding process is: ; in, express The target is The daily volatility, express The target is The settlement base value of the day, express The target is The settlement base value of the day.
[0013] Specifically, in this step, the text data of market dynamic news obtained contains a large amount of news content that does not contain the name of the target or contains multiple target names. This data is not easy to train the subsequent model, so the collected data needs to be cleaned; First, delete the risky targets marked with ST, *ST, SST, S*ST and PT from the target code library; Secondly, to prevent the additional impact of industry targets, name-related targets were deleted; Finally, the corresponding label code is matched according to the news content. If there are multiple labels in one news, the news will be deleted.
[0014] Step 2: Count the common words in the cleaned text data and build a bag of words for common words; Get the word frequency vector through the bag of common words; Generate a word frequency matrix using word frequency vectors; Take the word frequency matrix as the independent variable and the learning goal as the dependent variable, calculate the Pearson coefficient, and process the word frequency matrix according to the Pearson coefficient to obtain a new word frequency matrix; Reduce the dimension of the new word frequency matrix to obtain the reduced-dimensional word frequency matrix; In step 2, the common words in the cleaned text data are counted and a bag of common words is constructed. The relationship between the corresponding process is: ; in, represents the bag of common words, represents a list of words, represents the number of articles containing the vocabulary list, Indicates the threshold value; The word frequency vector is obtained through the bag of words of common words. The relationship between the corresponding process is: ; in, represents the word frequency vector, Indicates Words in this article The frequency of occurrence, Indicates the total number of articles; The word frequency matrix is generated using the word frequency vector. The corresponding relationship is: ; in, represents the word frequency matrix, Both represent word frequency vectors, Represents matrix transpose; Taking the word frequency matrix as the independent variable and the learning goal as the dependent variable, the Pearson coefficient is calculated. The relationship between the corresponding process is: ; in, represents the Pearson coefficient, Indicates the attributes in the independent variable No. Observations, The dependent variable Observations, represents the mean of the independent variable, represents the mean of the dependent variable; The Pearson coefficient is judged based on the threshold. When the Pearson coefficient is less than the threshold, the attribute is removed; After removing the relevant attributes, the remaining attributes are combined into a new word frequency matrix; The new word frequency matrix is standardized to obtain the standardized word frequency matrix. The relationship between the corresponding process is: ; in, represents the normalized word frequency matrix, represents the new word frequency matrix, Indicates the matrix The mean of the column, represents standard deviation; The covariance matrix is calculated using the standardized word frequency matrix to obtain the standardized word frequency covariance matrix. The corresponding relationship is: ; in, represents the standardized word frequency covariance matrix; Perform eigenvalue decomposition on the standardized word frequency covariance matrix to obtain eigenvalues and eigenvectors. The corresponding relationship is: ; in, represents the index of the feature, represents the feature vector, represents the eigenvalue; The contribution rate of the factor is calculated by the eigenvalue and eigenvector, and the corresponding process has the following relationship: ; in, Representation characteristics The contribution rate of the corresponding factor, Representation characteristics The characteristic value of Represents the total number of factors corresponding to the feature, represents the eigenvalue, represents the sum index; Based on the contribution rate of the factor, the cumulative contribution rate is obtained, and the relationship between the corresponding process is: ; in, Before The cumulative contribution rate of the factors, The contribution rate of the table factor, The index representing the contribution rate of the factor; Select the eigenvector based on the cumulative contribution rate, and use the selected eigenvector to project the standardized word frequency matrix into the new feature space to generate the reduced-dimensional word frequency matrix. The corresponding process has the following relationship: ; in, represents the word frequency matrix after dimensionality reduction, Before A matrix composed of factors.
[0015] Step 3: Build a regression model for the text based on the word frequency matrix after dimensionality reduction; In step 3, a regression model for the text is constructed based on the reduced word frequency matrix. The specific steps are as follows: The linear regression model is fitted using the reduced-dimensional word frequency matrix. The corresponding relationship is: ; in, represents the intercept to be estimated, represents the slope to be estimated, represents the random error term; The least squares method is used to estimate the regression parameters, and the corresponding process relationship is: ; in, represents the estimated intercept term, represents the estimated slope coefficient, Indicates The independent variable value of each observation; Based on the regression parameters, a regression model for text is constructed. The corresponding process has the following relationship: ; in, represents the predicted value of the dependent variable for a new observation, Represents the value of the independent variable for the new observation.
[0016] Step 4: Obtain image data of market dynamic news, and use the regression model for text to classify the image data to obtain labeled image data; The labeled image data is input into the pre-trained convolutional neural network model to obtain a fine-tuned pre-trained convolutional neural network model.
[0017] Furthermore, in this step, the image data of market dynamic news is collected by using the scrapy crawler framework to collect news data with pictures from major mainstream market dynamic news websites.
[0018] Specifically, in this step, the picture and the corresponding text are input into the regression model for the text, and the emotion classification of the corresponding news picture is output, and the image data is categorized based on the emotion classification.
[0019] For details, please refer to Figure 2 In this step, the Inception V3 model is used as the pre-trained convolutional neural network model. First, the Inception V3 model is imported from PyTorch, and the last fully connected layer is adjusted. The last layer of the original model has 1000 nodes, which is changed to 2 nodes. All network layer parameters except the last fully connected layer are frozen so that they do not participate in back propagation. Next, the last fully connected layer is fine-tuned using the annotated image dataset. In this way, the fine-tuned pre-trained convolutional neural network model is obtained.
[0020] Step 5: Use the image data to train the fine-tuned pre-trained convolutional neural network model to obtain a deep neural network emotion recognition model for images.
[0021] Step 6: Input the text data and image data of market dynamic news into the regression model for text and the deep neural network emotion recognition model for images respectively to obtain the strategy results.
[0022] See also Figure 3 The embodiment of the present invention further provides a quantitative strategy analysis system based on multimodal data, the system comprising: Text data processing module, used for: Acquire text data of market dynamic news, and clean the text data to obtain cleaned text data; Determine learning objectives using the settlement benchmark values before and after the date in the cleaned text data; Matrix building blocks for: Count the common words in the cleaned text data and build a bag of words for common words; Get the word frequency vector through the bag of common words; Generate a word frequency matrix using word frequency vectors; Take the word frequency matrix as the independent variable and the learning goal as the dependent variable, calculate the Pearson coefficient, and process the word frequency matrix according to the Pearson coefficient to obtain a new word frequency matrix; Reduce the dimension of the new word frequency matrix to obtain the reduced-dimensional word frequency matrix; Text model building blocks for: Build a regression model for text based on the reduced word frequency matrix; Image data processing module, used for: Obtain image data of market dynamic news, and use a regression model for text to classify the image data to obtain labeled image data; Input the labeled image data into the pre-trained convolutional neural network model to obtain a fine-tuned pre-trained convolutional neural network model; Image model training module, used for: The fine-tuned pre-trained convolutional neural network model is trained using image data to obtain a deep neural network emotion recognition model for images; Policy Results Module, which is used to: The text data and image data of market dynamic news are input into the regression model for text and the deep neural network emotion recognition model for image respectively to obtain the strategy results.
[0023] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0024] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0025] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A quantitative strategy analysis method based on multimodal data, characterized in that: The method comprises the following steps: Step 1: Obtain text data of market dynamic news, and clean the text data to obtain cleaned text data; Determine learning objectives using the settlement benchmark values before and after the date in the cleaned text data; Step 2: Count the common words in the cleaned text data and build a bag of words for common words; Get the word frequency vector through the bag of common words; Generate a word frequency matrix using word frequency vectors; Take the word frequency matrix as the independent variable and the learning goal as the dependent variable, calculate the Pearson coefficient, and process the word frequency matrix according to the Pearson coefficient to obtain a new word frequency matrix; Reduce the dimension of the new word frequency matrix to obtain the reduced-dimensional word frequency matrix; Step 3: Build a regression model for the text based on the word frequency matrix after dimensionality reduction; Step 4: Obtain image data of market dynamic news, and use the regression model for text to classify the image data to obtain labeled image data; Input the labeled image data into the pre-trained convolutional neural network model to obtain a fine-tuned pre-trained convolutional neural network model; Step 5: Use the image data to train the fine-tuned pre-trained convolutional neural network model to obtain a deep neural network emotion recognition model for images; Step 6: Input the text data and image data of market dynamic news into the regression model for text and the deep neural network emotion recognition model for images respectively to obtain the strategy results.
2. The quantitative strategy analysis method based on multimodal data according to claim 1 is characterized in that: In step 1, the learning target is determined using the settlement benchmark values before and after the date in the cleaned text data. The corresponding relationship in the process is: ; in, express The target is The daily volatility, express The target is The settlement base value of the day, express The target is The settlement base value of the day.
3. The quantitative strategy analysis method based on multimodal data according to claim 2 is characterized in that: In step 2, common words in the cleaned text data are counted and a bag of common words is constructed. The relationship between the corresponding process is: ; in, represents the bag of common words, represents a list of words, represents the number of articles containing the vocabulary list, Indicates the threshold value.
4. The quantitative strategy analysis method based on multimodal data according to claim 3 is characterized in that: In step 2, the word frequency vector is obtained through the common word bag, and the relationship between the corresponding process is: ; in, represents the word frequency vector, Indicates Words in this article The frequency of occurrence, Indicates the total number of articles.
5. The quantitative strategy analysis method based on multimodal data according to claim 4 is characterized in that: In step 2, the word frequency vector is used to generate a word frequency matrix, and the relationship between the corresponding process is: ; in, represents the word frequency matrix, Both represent word frequency vectors, Represents matrix transpose.
6. The quantitative strategy analysis method based on multimodal data according to claim 5 is characterized in that: In step 2, the word frequency matrix is used as the independent variable, the learning target is used as the dependent variable, the Pearson coefficient is calculated, and the word frequency matrix is processed according to the Pearson coefficient to obtain a new word frequency matrix. The specific steps are as follows: Taking the word frequency matrix as the independent variable and the learning goal as the dependent variable, the Pearson coefficient is calculated. The relationship between the corresponding process is: ; in, represents the Pearson coefficient, Indicates the attributes in the independent variable No. Observations, The dependent variable Observations, represents the mean of the independent variable, represents the mean of the dependent variable; The Pearson coefficient is judged based on the threshold. When the Pearson coefficient is less than the threshold, the attribute is removed; After removing the relevant attributes, the remaining attributes are combined into a new word frequency matrix.
7. The quantitative strategy analysis method based on multimodal data according to claim 6 is characterized in that: In step 2, the new word frequency matrix is reduced in dimension to obtain a word frequency matrix after dimension reduction. The specific steps are as follows: The new word frequency matrix is standardized to obtain the standardized word frequency matrix. The relationship between the corresponding process is: ; in, represents the normalized word frequency matrix, represents the new word frequency matrix, Indicates the matrix The mean of the column, represents standard deviation; The covariance matrix is calculated using the standardized word frequency matrix to obtain the standardized word frequency covariance matrix. The corresponding relationship is: ; in, represents the standardized word frequency covariance matrix; Perform eigenvalue decomposition on the standardized word frequency covariance matrix to obtain eigenvalues and eigenvectors. The corresponding relationship is: ; in, represents the index of the feature, represents the feature vector, represents the eigenvalue; The contribution rate of the factor is calculated by the eigenvalue and eigenvector, and the corresponding process has the following relationship: ; in, Representation characteristics The contribution rate of the corresponding factor, Representation characteristics The characteristic value of Represents the total number of factors corresponding to the feature, represents the eigenvalue, represents the sum index; Based on the contribution rate of the factor, the cumulative contribution rate is obtained, and the relationship between the corresponding process is: ; in, Before The cumulative contribution rate of the factors, represents the contribution rate of the factor, The index representing the contribution rate of the factor; Select the eigenvector based on the cumulative contribution rate, and use the selected eigenvector to project the standardized word frequency matrix into the new feature space to generate the reduced-dimensional word frequency matrix. The corresponding process has the following relationship: ; in, represents the word frequency matrix after dimensionality reduction, Before A matrix composed of factors.
8. The quantitative strategy analysis method based on multimodal data according to claim 7 is characterized in that: In step 3, a regression model for the text is constructed based on the word frequency matrix after dimensionality reduction. The specific steps are as follows: The linear regression model is fitted using the reduced-dimensional word frequency matrix. The corresponding relationship is: ; in, represents the intercept to be estimated, represents the slope to be estimated, represents the random error term; The least squares method is used to estimate the regression parameters, and the corresponding process relationship is: ; in, represents the estimated intercept term, represents the estimated slope coefficient, Indicates The independent variable value of each observation; Based on the regression parameters, a regression model for text is constructed. The corresponding process has the following relationship: ; in, represents the predicted value of the dependent variable for a new observation, Represents the value of the independent variable for the new observation.
9. A quantitative strategy analysis system based on multimodal data, characterized in that: The system applies any one of claims 1 to 8 of the quantitative strategy analysis method based on multimodal data, and the system comprises: Text data processing module, used for: Acquire text data of market dynamic news, and clean the text data to obtain cleaned text data; Determine learning objectives using the settlement benchmark values before and after the date in the cleaned text data; Matrix building blocks for: Count the common words in the cleaned text data and build a bag of words for common words; Get the word frequency vector through the bag of common words; Generate a word frequency matrix using word frequency vectors; Take the word frequency matrix as the independent variable and the learning goal as the dependent variable, calculate the Pearson coefficient, and process the word frequency matrix according to the Pearson coefficient to obtain a new word frequency matrix; Reduce the dimension of the new word frequency matrix to obtain the reduced-dimensional word frequency matrix; Text model building blocks for: Build a regression model for text based on the reduced word frequency matrix; Image data processing module, used for: Obtain image data of market dynamic news, and use a regression model for text to classify the image data to obtain labeled image data; Input the labeled image data into the pre-trained convolutional neural network model to obtain a fine-tuned pre-trained convolutional neural network model; Image model training module, used for: The fine-tuned pre-trained convolutional neural network model is trained using image data to obtain a deep neural network emotion recognition model for images; Policy Results Module, which is used to: The text data and image data of market dynamic news are input into the regression model for text and the deep neural network emotion recognition model for image respectively to obtain the strategy results.