Investor sentiment analysis method and system in financial stock market

By conducting word segmentation, semantic dependency matrix construction and attribute feature matrix analysis on investor text data in the financial stock market, combined with clustering algorithms and neural network models, the accuracy and timeliness of investor sentiment analysis are solved, and accurate market sentiment feedback and investment decision support are achieved.

CN120218077AInactive Publication Date: 2025-06-27TIBET DOLPHIN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510275066.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for existing technology to accurately analyze investor emotions in the financial stock market, especially in the face of massive noise data, complex language expressions and rapid changes in market sentiment, resulting in lagging and inaccurate analysis results.

Method used

By obtaining the text data of investor investment communication in the financial stock market, word segmentation processing and semantic dependency matrix construction, word vectors are extracted and text feature vectors are constructed. Then, the differences in attribute feature values ​​between text data are analyzed, the attribute feature matrix is ​​constructed, and the text data is divided through clustering algorithms, and finally sentiment analysis is performed using neural network models.

Benefits of technology

It realizes accurate division of financial stock market data and accurate analysis of investor emotions, can promptly feedback changes in market sentiment, provide investors with scientific and reasonable investment decision support, and improves the accuracy and efficiency of sentiment analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218077A_ABST
    Figure CN120218077A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of financial stock text analysis, in particular to an investor sentiment analysis method and system in a financial stock market, and the method comprises the steps: obtaining investment communication text data of investors in the financial stock market; obtaining a semantic dependency matrix and a dependency relationship type matrix of each piece of text data, and constructing a text feature vector of each piece of text data; constructing attribute feature values between each piece of text data and other pieces of text data, constructing an attribute feature matrix of each piece of text data, and further calculating a judgment coefficient between different pieces of text data so as to perform clustering division on the text data; and performing sentiment analysis by adopting a neural network model according to various types of text data after clustering division. According to the invention, the accuracy and efficiency of investor sentiment analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of financial stock text analysis, and specifically relates to a method and system for analyzing investors' sentiment in the financial stock market. Background Art

[0002] In the research and practice of the financial stock market, the accurate analysis of investors' sentiment is crucial. In addition to traditional channels such as financial news and company announcements, a large amount of investors' speech data has emerged on online platforms such as social media, stock forums, investment platform comment areas, and stock Q&A intelligent AIs. Among them, the Dolphin Stock intelligent AI provides a unique perspective and data source for investors' sentiment analysis. The chat records in its dialogue function cover the communication content between investors and the main bot and vertical intents. These text data reflect the emotional tendencies of investors during the process of consulting stock information, obtaining market information, etc. By analyzing these chat records, it is possible to explore investors' views and feelings on specific stocks, market sectors, or investment strategies.

[0003] At the same time, the extensive application of machine learning and deep learning technologies in the field of sentiment analysis, such as support vector machines, naive Bayes, recurrent neural networks, etc., has brought new opportunities to sentiment analysis. These technologies can automatically learn and identify the sentiment patterns in text, significantly improving the analysis efficiency and accuracy. However, the quality of investors' speech data on current online platforms is uneven, containing a large amount of noise, such as irrelevant information, advertisements, repetitive content, colloquial expressions, and typos. This not only increases the difficulty of data cleaning and preprocessing, but also may interfere with the subsequent sentiment analysis results. The language expressions of investors are complex, with rhetorical devices such as metaphors, irony, and exaggeration, and polysemous words and sentences are also relatively common, making it a difficult task to accurately judge the emotional tendency. Most of the existing sentiment analyses are based on the overall sentiment of investors, ignoring the characteristics of the rapidly changing emotions in the financial market, making it difficult to achieve precise analysis and timely feedback, resulting in the analysis results lagging behind the actual market situation and being unable to provide timely and effective decision-making support for investors. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of this application is to provide a method and system for analyzing investors' sentiment in the financial stock market, and the specific technical solutions adopted are as follows:

[0005] An embodiment of this application provides a method for analyzing investors' sentiment in the financial stock market, including the following steps:

[0006] Obtain the investment communication text data of investors in the financial stock market;

[0007] Perform word segmentation on each text data, obtain the semantic dependency matrix and dependency relationship type matrix of each text data, and construct the text feature vector of each text data;

[0008] Analyze the difference in the number of word segments of each text data from other text data, measure the correlation between each text data and other text data with respect to the semantic dependency relationship type matrix, evaluate the difference between each text data and other text data with respect to the semantic dependency matrix, and the similarity between each text data and other text data with respect to the text feature vector, and construct the attribute feature values between each text data and other text data, including the first, second, third, and fourth attribute feature values;

[0009] Construct the attribute feature matrix of each text data based on the first, second, third, and fourth attribute feature values between each text data and other text data;

[0010] Construct the judgment coefficient between different text data through the average level of each attribute feature value in the attribute feature matrix of each text data and the difference of the same attribute feature value between different text data, and perform clustering division on the text data in combination with the clustering algorithm;

[0011] Perform sentiment analysis on each type of text data after clustering using a neural network model.

[0012] Preferably, the construction process of the text feature vector further includes: extracting all word vectors in each text data, and forming the text feature vector of each text data by all the word vectors in each text data.

[0013] Preferably, the first attribute feature value between each text data and other text data is the absolute value of the difference between the number of word segments of each text data and the number of word segments of other text data.

[0014] Preferably, the second attribute feature value between each text data and other text data is the Jaccard coefficient between each text data and other text data with respect to the replaced and updated semantic dependency relationship type matrix, where the dependency relationship types are numerically marked respectively, and the dependency relationship types in the semantic dependency relationship type matrix are replaced with the marked numerical values to obtain the replaced and updated semantic dependency relationship type matrix.

[0015] Preferably, the third attribute feature value between each text data and other text data is the cumulative sum of the absolute differences between all elements at the same positions in the semantic dependency matrix of each text data and other text data.

[0016] Preferably, the fourth attribute feature value between each piece of text data and other pieces of text data is the cosine similarity between each piece of text data and other pieces of text data with respect to the text feature vector.

[0017] Preferably, the construction of the attribute feature matrix of each piece of text data further includes:

[0018] Taking the vector composed of the first, second, third, and fourth attribute feature values between each piece of text data and other pieces of text data as the attribute feature vector between each piece of text data and other pieces of text data;

[0019] Taking all the attribute feature vectors obtained between each piece of text data and other pieces of text data as column vectors, and the matrix formed is the attribute feature matrix of each piece of text data.

[0020] Preferably, the construction of the judgment coefficient between different pieces of text data further includes:

[0021] Using the objective weighting method to obtain the weights of each row vector in the attribute feature matrix, and calculating the mean value of each row vector in the attribute feature matrix of each piece of text data;

[0022] Calculating the product of the mean value of each row vector and the weight, taking the vector composed of all the products of each row vector in the attribute feature matrix as the judgment vector of each piece of text data, and taking the Euclidean distance between different pieces of text data with respect to the judgment vector as the judgment coefficient between different pieces of text data.

[0023] Preferably, the judgment coefficient between different pieces of text data is used as the measurement distance between different pieces of text data during the clustering and partitioning process.

[0024] The embodiment of the present application also provides an investor sentiment analysis system in the financial stock market, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the method described in any one of the above.

[0025] As can be seen from the above, an investor sentiment analysis method and system provided by the present application at least have the following beneficial effects:

[0026] The investor sentiment analysis method and system under the financial stock market proposed in this application effectively address the interference and influence of complex text data sources, diverse investor language expressions, and rapid market sentiment changes on the accuracy of sentiment analysis through a comprehensive and systematic data collection, data processing, and data analysis process. It can accurately classify financial stock market data and analyze investor sentiment, providing timely and accurate sentiment analysis results for investors and decision-makers, helping them better grasp market dynamics, make scientific and reasonable investment decisions, and improving the accuracy and efficiency of investor sentiment analysis. Brief Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a flowchart of the steps of an investor sentiment analysis method under the financial stock market provided by this application. Detailed Embodiments

[0029] In order to further elaborate on the technical means and effects adopted by the present application to achieve the intended invention purpose, the following, in combination with the drawings and preferred embodiments, details the specific embodiments, structures, features, and effects of an investor sentiment analysis method and system under the financial stock market proposed according to the present application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0030] Unless otherwise specified and limited, terms such as "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element. In addition, the term "and / or" used herein includes any and all combinations of one or more of the related listed items. All technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0031] The following specifically describes the specific solutions of a method and system for analyzing investors' sentiment in the financial stock market provided by this application in conjunction with the accompanying drawings.

[0032] Please refer to Figure 1 , which shows a flowchart of the steps of a method for analyzing investors' sentiment in the financial stock market provided by an embodiment of this application, including the following steps:

[0033] Step 1: Obtain the investment communication text data of investors in the financial stock market.

[0034] First of all, for the analysis of investors' sentiment in the financial stock market, a large amount of relevant text content needs to be obtained. Therefore, in the embodiment, the investment communication text data in the financial stock market will be obtained. Specifically, collect the text data of the Q&A history chat records of the dolphin stock intelligent AI assistant, financial news websites, popular stock bar forums, the financial sections of social media platforms, and professional investment communication communities. Among them, the text data of the Q&A history chat records of the dolphin stock intelligent AI assistant; specifically use the Scrapy framework in Python to obtain the above text content and record the metadata. In this embodiment, the metadata includes the release time, author, and title.

[0035] Step 2: Perform word segmentation processing on each text data, obtain the semantic dependency matrix and dependency relationship type matrix of each text data, extract all word vectors in each text data, and form the text feature vector of each text data.

[0036] Preprocess the obtained text data, specifically including removing advertisement, duplicate content, irrelevant information, and format error noise data in the text data. In this embodiment, the preprocessing of the text data can be achieved by regular expression matching.

[0037] For each text data in all the obtained text data, use the LTP (Language Technology Platform) tool to perform word segmentation on each text data, and split each text data into individual words; at the same time, use the LTP tool to obtain the semantic dependency graph of each text data, and obtain the semantic dependency matrix and dependency relationship type matrix of each text data.

[0038] The dependency relationship type matrix is represented by t=(t i,j ), where t i,jIt represents the dependency relationship type between the i-th word and the j-th word in the text. Among them, the semantic relationship types obtained by the LTP tool in this embodiment include subject-predicate relationship (SBV), verb-object relationship (VOB), indirect object relationship (IOB), and preposed object relationship (FOB). The semantic relationship types in all the obtained text data are encoded, that is, numerically marked. Specifically, in this embodiment, the above four relationships can be marked as 1, 2, 3, and 4 in sequence. In the actual application scenario, the implementer can mark and set them according to the actual situation. According to the encoding results, the semantic relationship type matrix of each text data is replaced and updated.

[0039] The semantic dependency matrix is represented by p = (p i,j ), where p i,j represents the dependency relationship weight between the i-th word and the j-th word in the text. It should be noted that if the word segmentation quantity of a text data is n, the dependency relationship type matrix and the semantic dependency matrix constructed by this text data are both n×n matrices.

[0040] Furthermore, the finbert model oriented to the financial field is used to process each text data; each text data is input into the finbert model, and the word vectors containing the prior knowledge in the field are output for each text data. The vector composed of all the word vectors of each text data is the text feature vector of each text data.

[0041] Step 3, by evaluating the difference in word segmentation quantity between each text data and other text data, analyzing the correlation relationship between each text data and other text data regarding the semantic dependency relationship type matrix, analyzing the difference between each text data and other text data regarding the semantic dependency relationship matrix, and the similarity between each text data and other text data regarding the text feature vector, the attribute feature values between each text data and other text data are constructed, including the first, second, third, and fourth attribute feature values.

[0042] Furthermore, since the text data obtained in this embodiment mainly comes from financial news websites, popular stock bar forums, the financial sections of social media platforms, and professional investment communication communities, its content source is relatively complex, and there are significant differences in the evaluation methods and expressed content of different types of users facing the financial stock market. Therefore, it is easy to have a large deviation in part-of-speech tagging and then analyzing the sentiment tendency for complex text content.

[0043] To solve the above-mentioned deviation analysis existing in the investor sentiment analysis process, in this embodiment, the collected text data is classified, and the sentiment analysis is accurately analyzed based on the characteristics of different groups of text data in the classification results.

[0044] Specifically, for each piece of text data in the obtained text data, compare the relationship of attribute characteristics between each piece of text data and other text data. The attribute characteristics include the number of word segments, semantic dependency relationship characteristics, and text feature vectors. Specifically, the calculation method of the attribute characteristic values in this embodiment is as follows:

[0045] (1) Calculate the absolute value of the difference between the number of word segments of each piece of text data and the number of word segments of other pieces of text data, and record it as the first attribute characteristic value between each piece of text data and other pieces of text data.

[0046] (2) Calculate the correlation coefficient between each piece of text data and other pieces of text data with respect to the matrix of semantic dependency relationship types after replacement and update, and record it as the second attribute characteristic value between each piece of text data and other pieces of text data. There are many existing calculation methods for the correlation coefficient: correlation coefficient method, eigenvector method, Jaccard coefficient, etc. In this embodiment, the Jaccard coefficient is used to calculate the correlation coefficient.

[0047] (3) Calculate the sum of the absolute differences between all elements at the same positions in the semantic dependency matrix of each piece of text data and the semantic dependency matrix of other pieces of text data, and record it as the third attribute characteristic value between each piece of text data and other pieces of text data.

[0048] (4) Calculate the similarity between each piece of text data and other pieces of text data with respect to the text feature vectors, and record the similarity as the fourth attribute characteristic value between each piece of text data and other pieces of text data. There are many existing calculation methods for the similarity: cosine similarity, dot product, etc. In this embodiment, cosine similarity is used for evaluation.

[0049] Step 4, construct an attribute characteristic matrix for each piece of text data according to the first, second, third, and fourth attribute characteristic values between each piece of text data and other pieces of text data.

[0050] Furthermore, the obtained attribute characteristic values reflect the relationship of attribute characteristics in the process of text data acquisition. That is, if the number of word segments of the text data is close, and the differences in semantic dependency relationships and the correlation characteristics of text feature vectors between the text data are more significant, it indicates that the text attribute characteristic relationships between the current text data are close. Then, when facing sentiment analysis, the influence of factors such as the length of semantic dependency relationship texts on the accuracy of sentiment analysis can be reduced.

[0051] Specifically, since the attribute sensitivity differences between different text data and other text data are different, for each piece of text data, the vectors composed of the first, second, third, and fourth attribute characteristic values between each piece of text data and other pieces of text data are respectively used as the attribute characteristic vectors between each piece of text data and other pieces of text data.

[0052] For each piece of text data, construct an attribute feature matrix for each piece of text data. Specifically, all the attribute feature vectors obtained between each piece of text data and other pieces of text data are used as column vectors, and the formed matrix is used as the attribute feature matrix for each piece of text data. Further, taking the attribute feature matrix as the input, an objective weighting method is used to obtain the weights of each row vector in the attribute feature matrix. The weights are used to reflect the influence degree of the sensitive attributes of the current piece of text data among all the text data for sentiment analysis.

[0053] Step 5: Construct a judgment coefficient between different pieces of text data based on the average level of each attribute feature value in the attribute feature matrix of each piece of text data and the difference of the same attribute feature value between different pieces of text data, and combine the clustering algorithm to perform clustering division on the text data.

[0054] Based on the attribute features between the text data and the influence degree of the sensitive attributes, obtain the judgment coefficient of the attribute features between the texts. For each piece of text data, the calculation method of the judgment coefficient in this embodiment is as follows:

[0055] (1) First, calculate the mean value of each row vector in the attribute feature matrix of each piece of text data. This mean value reflects the average level of the same attribute feature relationship between each piece of text data and other pieces of text data.

[0056] (2) Calculate the product of the mean value of each row vector and the weight of the corresponding row vector in the attribute feature matrix of each piece of text data, and form a vector with all the products of the row vectors in the attribute feature matrix as the judgment vector for each piece of text data.

[0057] (3) Calculate the difference between the judgment vectors of different pieces of text data as the judgment coefficient between the text data, which reflects the associated difference characteristics of the two pieces of text data for investor sentiment analysis; among them, there are many existing methods for measuring the difference between judgment vectors: Euclidean distance, Manhattan distance, and Mahalanobis distance. In this embodiment, the Euclidean distance is used for measurement.

[0058] Further, after the above calculations, in the process of investor sentiment analysis in the financial stock market, due to the complex data sources of the collected texts, and the large differences in the expression methods and semantic dependency relationships of the text data, etc., it may have a great impact on the sentiment analysis results of the text data. Therefore, in this embodiment, the attribute feature relationships between different pieces of text data are analyzed to accurately judge the change influence characteristics of the sensitive attributes of different pieces of text data compared with the overall text data, and then calculate the judgment coefficient between different pieces of text data; through the judgment coefficient, accurately divide the text data with similar influence relationships of attribute sensitive characteristics, and avoid affecting the accuracy of sentiment analysis due to large differences in semantic dependency relationships, text sources, and semantic characteristics, etc.

[0059] Therefore, in this embodiment, all text data is used as input, and a clustering method is adopted to obtain the clustering division results of all text data. The calculation method of the distance between different text data during the clustering process is the judgment coefficient between different text data, and its purpose is to accurately divide the text data based on the influence relationship of the attribute-sensitive features of the text data, thereby improving the accuracy of text data sentiment analysis for complex sources and semantic dependence features; the clustering method can be agglomerative hierarchical clustering and divisive hierarchical clustering. In this embodiment, the agglomerative hierarchical clustering method is adopted, and the specific clustering process is a well-known prior art and will not be elaborated in this embodiment.

[0060] Step 6: Perform sentiment analysis on each type of text data after clustering using a neural network model.

[0061] In this embodiment, a sentiment analysis model is used to obtain the results of investor sentiment analysis; the range of available sentiment analysis models is extensive, covering various types. Common methods include sentiment analysis methods based on dictionaries, models using machine learning algorithms such as support vector machines, naive Bayes, and decision trees, and models based on deep learning architectures such as long short-term memory networks (LSTM) and convolutional neural networks (CNN).

[0062] Preferably, the specific process of sentiment analysis in this embodiment is as follows:

[0063] First, each piece of text data and its corresponding judgment vector are used as a training sample. All the training samples corresponding to each group of text data after division are marked as a training set and a test set according to a ratio of 7:3, and part-of-speech tagging is performed on the text data in each training sample. Subsequently, the long short-term memory network (LSTM) model is trained using the marked training set. During the training process, by inputting the data of the training set into the LSTM model, it learns the internal relationship between the text data and the corresponding sentiment tendency to achieve accurate analysis of the sentiment tendency; the optimizer of the LSTM model is the Adam optimizer (Adaptive Moment Estimation), and the loss function is the cross-entropy loss function. The specific training process of the LSTM model is well-known to those skilled in the art and will not be elaborated further.

[0064] It should be noted that the specific training process of the LSTM model is a well-known prior art, and no special restrictions are imposed on this in this embodiment and will not be elaborated one by one. At the same time, in actual application scenarios, implementers can use other neural network models and machine learning models for sentiment analysis and recognition.

[0065] Further, input the text data to be analyzed and its corresponding judgment vector into the trained sentiment analysis model, so that the model conducts sentiment tendency analysis on these data and outputs the sentiment tendency results of each text. Specifically, the sentiment tendencies involved in this embodiment will be specifically divided into three categories: positive, negative, and neutral, in order to clearly reflect the emotional attitudes expressed by the texts.

[0066] Further, to facilitate the understanding and interpretation of the sentiment analysis results by investors and decision-makers, present the sentiment analysis results in an intuitive chart form, such as bar charts, line charts, and word clouds. Bar charts can visually display the distribution of the number of texts with different sentiment tendencies; line charts can reflect the changing trends of sentiment tendencies over time or other factors; word clouds can highlight the keywords in texts with different sentiment tendencies, helping users quickly grasp the key information and providing a clear and intuitive basis for subsequent decision-making and analysis. In this embodiment, bar charts are used for the visualization of sentiment result analysis. Specifically, the construction of bar charts, line charts, and word clouds is a well-known technology for statistical analysis of texts with different sentiment tendencies in this field and will not be elaborated here.

[0067] Based on the same inventive concept as the above method, the embodiment of the present application also provides an investor sentiment analysis system in the financial stock market, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods for investor sentiment analysis in the financial stock market.

[0068] It can be understood that: the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0069] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.

[0070] The above content is only the implementation manner of the present application and is not used to limit the scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, shall be included in the protection scope of the present application by the same token.

Claims

1. A method for analyzing investor sentiment in a financial stock market, characterized in that: The following steps are involved: Obtain investment communication text data of investors in the financial stock market; Perform word segmentation on each piece of text data, obtain the semantic dependency matrix and dependency relationship type matrix of each piece of text data, and construct a text feature vector for each piece of text data; Analyze the difference between each text data and other text data in terms of the number of segmented words, measure the correlation between each text data and other text data in terms of the semantic dependency type matrix, evaluate the difference between each text data and other text data in terms of the semantic dependency matrix, and the similarity between each text data and other text data in terms of the text feature vector, and construct attribute feature values ​​between each text data and other text data, including the first, second, third, and fourth attribute feature values; Constructing an attribute feature matrix for each piece of text data according to the first, second, third and fourth attribute feature values ​​between each piece of text data and other pieces of text data; Based on the average level of each attribute feature value in the attribute feature matrix of each text data and the difference of the same attribute feature value between different text data, the judgment coefficient between different text data is constructed, and the text data is clustered and divided in combination with the clustering algorithm; According to the various types of text data after clustering, the neural network model is used for sentiment analysis.

2. The investor sentiment analysis method in a financial stock market as claimed in claim 1, characterized in that: The process of constructing the text feature vector further includes: extracting all word vectors in each piece of text data, and combining all word vectors in each piece of text data into a text feature vector for each piece of text data.

3. The investor sentiment analysis method in a financial stock market as claimed in claim 1, characterized in that: The first attribute feature value between each piece of text data and other pieces of text data is the absolute value of the difference between the number of segmented words in each piece of text data and the number of segmented words in other pieces of text data.

4. The investor sentiment analysis method in a financial stock market as claimed in claim 1, characterized in that: The second attribute characteristic value between each text data and other text data is the Jaccard coefficient between each text data and other text data with respect to the replaced updated semantic dependency type matrix, wherein the dependency types are numerically labeled respectively, and the labeled numerical values ​​are used to replace the dependency types in the semantic dependency type matrix to obtain the replaced updated semantic dependency type matrix.

5. The investor sentiment analysis method in a financial stock market as claimed in claim 1, characterized in that: The third attribute feature value between each piece of text data and other pieces of text data is the cumulative sum of the absolute differences between all elements at the same position in the semantic dependency matrix of each piece of text data and other pieces of text data.

6. The investor sentiment analysis method in a financial stock market as claimed in claim 1, characterized in that: The fourth attribute feature value between each piece of text data and other pieces of text data is the cosine similarity between each piece of text data and other pieces of text data with respect to the text feature vector.

7. The investor sentiment analysis method in a financial stock market as claimed in claim 1, characterized in that: The construction of the attribute feature matrix of each text data further includes: The vector formed by the first, second, third and fourth attribute feature values ​​between each piece of text data and other pieces of text data is used as the attribute feature vector between each piece of text data and other pieces of text data; All attribute feature vectors obtained between each text data and other text data are taken as column vectors, and the matrix formed is the attribute feature matrix of each text data.

8. The investor sentiment analysis method in a financial stock market as claimed in claim 1, characterized in that: The construction of the judgment coefficient between the different text data further includes: The objective weighting method is used to obtain the weight of each row vector in the attribute feature matrix, and the mean of each row vector in the attribute feature matrix of each text data is calculated; The product of the mean value of each row vector and the weight is calculated, and the vector formed by the product of all row vectors in the attribute feature matrix is ​​used as the judgment vector of each text data, and the Euclidean distance of different text data with respect to the judgment vector is used as the judgment coefficient between different text data.

9. The investor sentiment analysis method in a financial stock market as claimed in claim 1, characterized in that: In the clustering process, the judgment coefficient between different text data is used as a distance measurement between different text data.

10. An investor sentiment analysis system in a financial stock market, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.