Real-time public opinion analysis and investment opportunity mining method and system for stock staring at disk

By capturing and analyzing the Internet public opinion corpus and its associated corpus, determining the text correlation characteristics between corpus and calculating the weighted word frequency, the problem of inaccurate keyword statistics caused by ignoring the corpus propagation structure in the existing technology is solved, and more accurate public opinion analysis and investment opportunity mining is achieved.

CN120146033AInactive Publication Date: 2025-06-13TIBET DOLPHIN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510223405.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art ignores the communication structure and communication methods of the corpus in the keyword statistics of Internet public opinion corpus, resulting in inaccurate keyword statistics and affecting the accuracy of public opinion analysis.

Method used

By capturing corpus data and its associated corpus, determining the text correlation characteristics between corpus, obtaining strongly associated text and supplementary corpus, calculating weighted word frequency, and using the associated TF value as the TF value of the TF-IDF algorithm, performing word frequency extraction to achieve more accurate public opinion analysis.

Benefits of technology

By considering the corpus propagation structure, the accuracy of keyword statistics is improved, the accuracy of public opinion analysis is enhanced, and keyword extraction errors caused by too short text are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146033A_ABST
    Figure CN120146033A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of word frequency statistics, and provides a real-time public opinion analysis and investment opportunity mining method and system for stock staring in a staring disk, and the method comprises the steps: capturing all corpus data which comprise text data and related corpora of the text data; marking the target corpus data, and determining all strong association texts of the target corpus data; determining supplementary corpora of the target corpus data, and obtaining weighted word frequencies of all different segmented words contained in the target corpus data and a word segmentation result of the supplementary corpora of the target corpus data; according to the target corpus data, the supplementary corpus of the target corpus data, all the strong association texts of the target corpus data and the weighted word frequency of all the segmented words contained in the word segmentation result of the supplementary corpus of all the strong association texts of the target corpus data, calculating an association TF value of each same segmented word, and obtaining a real-time public opinion analysis and investment opportunity mining result in the stock staring disk according to the associated TF value. The accuracy of keyword statistics and public opinion analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of word frequency statistics, and particularly to a method and system for real-time public opinion analysis and investment opportunity mining in stock monitoring. Background Art

[0002] In the financial market, real-time public opinion analysis and investment opportunity mining are key links in stock monitoring. By analyzing market public opinion, the basis for investment decisions can be obtained to assist investors in understanding the dynamics of the financial market and quickly extracting market information. For example, by analyzing news reports, discussions on social media, etc., to understand investors' views and expectations on a certain stock or the entire market, and then judge the rise and fall trend of the stock. Usually, the TF-IDF algorithm is used as the keyword extraction algorithm for public opinion analysis in public opinion analysis to realize the word frequency statistics of Internet corpus.

[0003] The TF-IDF algorithm takes all the extracted corpus as a whole and finds the keywords in the corpus through the difference in the distribution probability of words. However, Internet public opinion corpus often has obvious user attributes, and the TF-IDF algorithm cannot capture user attribute characteristics, often resulting in inaccurate statistics of corpus keywords and affecting the accuracy of public opinion analysis. Summary of the Invention

[0004] The present invention provides a method and system for real-time public opinion analysis and investment opportunity mining in stock monitoring to solve the problem that inaccurate keyword statistics and inaccurate public opinion analysis are caused by ignoring the dissemination structure and dissemination method of Internet public opinion corpus. The specific technical solutions adopted are as follows:

[0005] In a first aspect, an embodiment of the present invention provides a method for real-time public opinion analysis and investment opportunity mining in stock monitoring, and the method includes the following steps:

[0006] Grab all corpus data, and each piece of the corpus data includes text data and associated corpus of the text data;

[0007] Denote any piece of the corpus data as target corpus data, and denote any associated corpus of the target corpus data as candidate corpus data. Determine the text association characteristics of the target corpus data and the candidate corpus data according to the text similarity degree between the target corpus data and the candidate corpus data, and determine all strongly associated texts of the target corpus data according to all text association characteristics corresponding to the target corpus data and the corpus data corresponding to all associated corpora.

[0008] Determine the supplementary corpus of the target corpus data according to the target corpus data and other corpus data cited in the associated corpus included in the target corpus data. Obtain the weighted word frequencies of all different words included in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data according to the number of associated corpora included in the supplementary corpus of the target corpus data, the word segmentation result of the supplementary corpus of the target corpus data, and the word segmentation result of the target corpus data. Obtain the weighted word frequencies of all different words included in the word segmentation results of each grabbed corpus data and the supplementary corpus of the corpus data.

[0009] Calculate the associated TF value of each identical word according to the weighted word frequencies of all words included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data. Take the associated TF value of the word as the TF value of the TF-IDF algorithm, and use the TF-IDF algorithm to obtain the word frequency extraction result of the target corpus data, and obtain the real-time public opinion analysis and investment opportunity mining result in stock monitoring.

[0010] Furthermore, the specific method for determining the text association characteristics of the target corpus data and the candidate corpus data according to the text similarity degree between the target corpus data and the candidate corpus data is as follows:

[0011] Perform word segmentation extraction on the target corpus data and the candidate corpus data respectively to obtain the word segmentation result of the target corpus data and the word segmentation result of the candidate corpus data; record the hash value of the word segmentation result of the target corpus data as the text hash value of the target corpus data; record the hash value of the word segmentation result of the candidate corpus data as the text hash value of the candidate corpus data.

[0012] Determine the text association characteristics of the target corpus data and the candidate corpus data according to the difference between the text hash values of the target corpus data and the candidate corpus data.

[0013] Furthermore, the specific method for determining the text association characteristics of the target corpus data and the candidate corpus data according to the difference between the text hash values of the target corpus data and the candidate corpus data is as follows:

[0014] Take the reciprocal of the Mahalanobis distance of the text hash values of the target corpus data and the candidate corpus data as the text association characteristics of the target corpus data and the candidate corpus data.

[0015] Furthermore, the determination method for all strongly associated texts of the target corpus data is as follows:

[0016] Record the associated corpus corresponding to the maximum value in the text association characteristics of the target corpus data as the strongly associated text of the target corpus data.

[0017] Obtain the corpus data corresponding to the strongly associated text of the target corpus data. According to the method of obtaining the strongly associated text of the target corpus data, determine the strongly associated text of the corpus data corresponding to the strongly associated text, and also record the strongly associated text of the corpus data corresponding to the strongly associated text as the strongly associated text of the target corpus data;

[0018] According to the same method, continue to obtain the strongly associated text of the target corpus data based on the strongly associated text of the new target corpus data until H strongly associated texts of the target corpus data are obtained, where H represents the first preset threshold.

[0019] Furthermore, the method for determining the supplementary corpus of the target corpus data is as follows:

[0020] Record the other corpus data cited in the associated corpus contained in the target corpus data as the indirectly cited corpus of the target corpus data;

[0021] Record both the indirectly cited corpus of the target corpus data and the associated corpus contained in the target corpus data as the supplementary corpus of the target corpus data.

[0022] Furthermore, the specific method for obtaining the weighted word frequencies of all different words contained in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data according to the number of associated corpora contained in the supplementary corpus of the target corpus data, the word segmentation result of the supplementary corpus of the target corpus data, and the word segmentation result of the target corpus data includes:

[0023] Record the number of associated corpora contained in the supplementary corpus of the target corpus data as the citation times of the supplementary corpus of the target corpus data; record the cumulative sum of the citation times of all supplementary corpora of the target corpus data as the potential citation times of the target corpus data; record the ratio of the citation times of the supplementary corpus of the target corpus data to the potential citation times of the target corpus data as the citation word frequency weight of the supplementary corpus of the target corpus data;

[0024] Obtain the word segmentation result of the supplementary corpus of the target corpus data by word segmentation extraction, and record the number of times the same word appears in the word segmentation result of the supplementary corpus as the frequency of the same word contained in the word segmentation result of the supplementary corpus;

[0025] Record the product of the cumulative sum of the frequencies of the same words contained in the word segmentation result of the supplementary corpus of the target corpus data and the citation word frequency weight of the supplementary corpus of the target corpus data as the weighted word frequency of the same word contained in the word segmentation result of the supplementary corpus of the target corpus data;

[0026] Based on the word segmentation results of the target corpus data, determine the weighted word frequency of each different word segment in the word segmentation results of the target corpus data, and update the weighted word frequency of the same word segments included in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data.

[0027] Further, the method for determining the weighted word frequency of each different word segment in the word segmentation results of the target corpus data, and updating the weighted word frequency of the same word segments included in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data, includes the following specific methods:

[0028] Record the cumulative sum of the occurrences of any same word segment included in the word segmentation results of the target corpus data as the weighted word frequency of the any same word segment;

[0029] Record the sum of the weighted word frequencies of the same word segments in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data as the weighted word frequency of the same word segments in the word segmentation results.

[0030] Further, the method for calculating the associated TF value of each same word segment based on the weighted word frequencies of all word segments included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data, includes the following specific methods:

[0031] Take the cumulative sum of the weighted word frequencies of all same word segments among the weighted word frequencies of all different word segments included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data as the weighted word frequency of the same word segments;

[0032] Record the cumulative sum of the weighted word frequencies of all different word segments included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data as the total word frequency;

[0033] Determine the associated TF value of each word segment respectively according to the weighted word frequency and the total word frequency of the word segments included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data.

[0034] Further, the method for respectively determining the associated TF value of each word segment, based on the weighted word frequency and the total word frequency of the word segments included in the word segment results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data, includes the following specific method:

[0035] The ratio of the weighted word frequency to the total word frequency of any word segment included in the word segment results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data is recorded as the associated TF value of the any word segment.

[0036] In a second aspect, an embodiment of the present invention further provides a real-time public opinion analysis and investment opportunity mining system in stock monitoring, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0037] The beneficial effects of the present invention are:

[0038] This application analyzes the user corpus propagation structure that is propagated multiple times in the form of forwarding and re-commenting. Considering that the less the associated relationship between two corpus data and the more similar the text data of the two corpus data, the more similar the two corpus data are, the text association characteristics between different corpus data are obtained. Furthermore, considering that the corpus data with more citation times is more likely to be the center of public opinion, the weighted word frequencies of all different word segments included in the word segment results of each grabbed corpus data and the supplementary corpus of the corpus data are obtained. The weighted word frequency is the comprehensive calculation result of the word segment occurrence times based on the corpus data with indirect and direct citation relationships. Among them, the same keywords in the associated corpus with more citation times and more times of being cited are more suitable as the extracted keywords. Selecting these same keywords as the extracted keywords can expand the text length of the text data in the corpus data, which is beneficial to improving the accuracy of extracting short network texts and avoiding the problem of keyword extraction errors caused by too short texts. Finally, according to the weighted word frequency, the associated TF value of each same word segment is calculated, and the associated TF value of the word segment is used as the TF value of the TF-IDF algorithm. The TF-IDF algorithm is used to obtain the word frequency extraction result of the target corpus data, and the real-time public opinion analysis and investment opportunity mining result in stock monitoring is obtained, so as to solve the problem that due to ignoring the propagation structure and propagation mode of Internet public opinion corpus, the keyword statistics are inaccurate, resulting in inaccurate public opinion analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0040] Figure 1 It is a schematic flowchart of a method for real-time public opinion analysis and investment opportunity mining in stock monitoring provided by an embodiment of the present invention;

[0041] Figure 2 It is a flowchart for obtaining text association characteristics provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0043] Please refer to Figure 1 , which shows a flowchart of a method for real-time public opinion analysis and investment opportunity mining in stock monitoring provided by an embodiment of the present invention. The method includes the following steps:

[0044] Step S001, capture all corpus data, and each piece of the corpus data includes text data and associated corpus of the text data.

[0045] Capture corpus data using corpus capture software in a financial information exchange forum. Among them, using corpus capture software to capture corpus data is a well-known technology and will not be elaborated.

[0046] In a financial information exchange forum, users can post comments, and other users can comment on or forward the comments again. Therefore, events and comments can be secondarily spread in the way of forwarding and re-commenting. Therefore, there is often a high degree of correlation between different corpus data of forwarding and re-commenting. This correlation is very important in the process of extracting word frequency statistics, and the accuracy of word frequency statistics can be improved according to the high degree of correlation between different corpus data. However, the TF-IDF algorithm cannot count corpus keywords according to the correlation between different corpus data.

[0047] Therefore, it can be understood that the corpus data includes text data and associated corpus of the text data. The text data is the main body of the corpus data, and the associated corpus of the text data is other corpus data that forwards and re - comments on the corpus data, as well as other corpus data cited in the corpus data.

[0048] It can be understood that the comments published by users can be commented and forwarded by different users. Therefore, the associated corpus included in the corpus data may be one or more, and the multiple associated corpora included in the corpus data respectively correspond to different other corpus data.

[0049] Preferably, as an embodiment of the present application, a total of 50,000 corpus data are collected, and each corpus data includes two parts: text data and associated corpus of the text data.

[0050] Thus, the corpus data is obtained.

[0051] Step S002: Denote any one of the corpus data as the target corpus data, and denote any one of the associated corpora of the target corpus data as the candidate corpus data. Determine the text association characteristics of the target corpus data and the candidate corpus data according to the text similarity between the target corpus data and the candidate corpus data. Determine all strongly associated texts of the target corpus data according to all the text association characteristics corresponding to the target corpus data and the corpus data corresponding to all the associated corpora.

[0052] Comments of all users on events in the financial information exchange forum can be spread multiple times through forwarding and re - commenting, forming a user corpus propagation structure of the corpus data in the financial information exchange forum. Each corpus data contains associated corpora, and there is an association relationship between the associated corpora and other corpus data. Therefore, a complex corpus network can be constructed with any one of the corpus data as the center and the association situation as the connection method. There are complex association relationships between different corpus data in the corpus network. And when the number of association relationships between two corpus data is less and the text data of the two corpus data is more similar, the two corpus data are more similar. Based on this specific user corpus propagation structure, more accurate word frequency statistics can be performed, thereby improving the accuracy of keyword statistics and further improving the accuracy of public opinion analysis.

[0053] Denote any corpus data as the target corpus data, and denote any associated corpus of the target corpus data as the candidate corpus data. Perform word segmentation extraction on the target corpus data and the candidate corpus data respectively using the bidirectional matching method to obtain the word segmentation results of the target corpus data and the candidate corpus data; process the word segmentation results of the target corpus data using the Simhash algorithm to obtain the hash value of the word segmentation results of the target corpus data, and denote the hash value of the word segmentation results of the target corpus data as the text hash value of the target corpus data; process the word segmentation results of the candidate corpus data using the Simhash algorithm to obtain the hash value of the word segmentation results of the candidate corpus data, and denote the hash value of the word segmentation results of the candidate corpus data as the text hash value of the candidate corpus data. Take the reciprocal of the Mahalanobis distance between the text hash values of the target corpus data and the candidate corpus data as the text association characteristic between the target corpus data and the candidate corpus data.

[0054] Among them, performing word segmentation extraction using the bidirectional matching method, obtaining the hash value using the Simhash algorithm, and calculating the Mahalanobis distance of different text hash values are all well-known technologies and will not be elaborated here.

[0055] According to the same method, the text association characteristics between any corpus data in all corpus data and each of its associated corpora can be obtained. The flowchart for obtaining the text association characteristics is as Figure 2 shown.

[0056] Denote the associated corpus corresponding to the maximum value in the text association characteristics of the target corpus data as the strong associated text of the target corpus data.

[0057] Obtain the corpus data corresponding to the strong associated text of the target corpus data. According to the method of obtaining the strong associated text of the target corpus data, determine the strong associated text of the corpus data corresponding to the strong associated text, and also denote the strong associated text of the corpus data corresponding to the strong associated text as the strong associated text of the target corpus data, that is, obtain the new strong associated text of the target corpus data. According to the same method, continue to obtain the strong associated text of the target corpus data based on the new strong associated text of the target corpus data until H strong associated texts of the target corpus data are obtained.

[0058] Among them, H represents the first preset threshold, and in this embodiment, the value of the first preset threshold is 10; the H strong associated texts of the slogan corpus data are H corpus data determined according to the public opinion dissemination relationship and having a high similarity degree with the target text.

[0059] According to the same method, all the strong associated texts of each corpus data can be obtained.

[0060] Thus, all the strong associated texts of the corpus data are obtained.

[0061] Step S003: Determine the supplementary corpus of the target corpus data based on the target corpus data and other corpus data cited in the associated corpus contained in the target corpus data. Obtain the weighted word frequencies of all different words contained in the word segmentation results of the target corpus data and its supplementary corpus according to the number of associated corpus contained in the supplementary corpus of the target corpus data, the word segmentation result of the supplementary corpus of the target corpus data, and the word segmentation result of the target corpus data. Obtain the weighted word frequencies of all different words contained in the word segmentation results of each grabbed corpus data and its supplementary corpus.

[0062] In a financial information exchange forum, users usually express opinions based on the same event and spread the opinions expressed by users through re - forwarding or re - commenting, gradually forming an opinion spread. Therefore, the associated corpus of text data is also part of the corpus data. However, the TF - IDF algorithm cannot perform word frequency statistics according to the spread order and relevance between different corpus data, and inaccurate word frequency statistics may occur, which may lead to incorrect keyword extraction.

[0063] Among all the extracted corpus data, the corpus data with more citation times is more likely to be the center of public opinion. Therefore, when extracting keywords from the target corpus data, the corpus data that is cited multiple times should be found from the associated corpus cited by the target corpus data, and the corpus data corresponding to the associated corpus with more citation times and being cited times should be used as the text supplement of the target corpus data. Further, the keywords that are the same as the target corpus data in these associated corpus with more citation times and being cited times are more suitable as the extracted keywords. Selecting these same keywords as the extracted keywords can expand the text length of the text data in the corpus data, which is beneficial to improving the accuracy of extracting short network texts and avoiding the problem of incorrect keyword extraction due to overly short texts.

[0064] Record the other corpus data cited in the associated corpus contained in the target corpus data as the indirect citation corpus of the target corpus data. Record both the indirect citation corpus of the target corpus data and the associated corpus contained in the target corpus data as the supplementary corpus of the target corpus data.

[0065] It can be understood that the associated corpus contained in the corpus data is other corpus data of the re - forwarded and re - commented corpus data, as well as other corpus data cited in the corpus data. Therefore, there is a direct citation relationship between the associated corpus contained in the corpus data and the corpus data. And the indirect citation corpus of the target corpus data is the corpus data that has an indirect citation relationship with the target corpus data. Therefore, the supplementary corpus of the target corpus data includes all corpus data that presents a direct citation relationship and an indirect citation relationship with the target corpus data.

[0066] Determine the citation word frequency weight of the supplementary corpus of the target corpus data according to the number of associated corpora included in all supplementary corpora of the target corpus data.

[0067] Count the number of associated corpora in each corpus data among all corpus data. Denote the number of associated corpora included in the supplementary corpus of the target corpus data as the citation times of the supplementary corpus of the target corpus data; Denote the cumulative sum of the citation times of all supplementary corpora of the target corpus data as the potential citation times of the target corpus data; Denote the ratio of the citation times of the supplementary corpus of the target corpus data to the potential citation times of the target corpus data as the citation word frequency weight of the supplementary corpus of the target corpus data.

[0068] When the citation word frequency weight of the supplementary corpus of the target corpus data is larger, the citation times of the supplementary corpus are more, the supplementary corpus is more likely to be the center of public opinion, and the word frequency of the target corpus data calculated based on the supplementary corpus is more accurate. Therefore, more reference should be made to the supplementary corpus of the target corpus data with a larger citation word frequency weight to determine the word frequency of the target corpus data. At the same time, during the process of determining the word frequency of the target corpus data, less reference should be made to the supplementary corpus of the target corpus data with a smaller citation word frequency weight.

[0069] Furthermore, perform word segmentation extraction on the supplementary corpus of the target corpus data using the bidirectional matching method to obtain the word segmentation result of the supplementary corpus of the target corpus data. Denote the number of occurrences of each word segmentation included in the word segmentation result of the supplementary corpus as the frequency of the word segmentation.

[0070] Denote the product of the cumulative sum of the frequencies of the same word segmentations included in the word segmentation result of the supplementary corpus of the target corpus data and the citation word frequency weight of the supplementary corpus of the target corpus data as the weighted word frequency of the same word segmentation.

[0071] Thus, obtain the weighted word frequency of each different word segmentation included in the word segmentation result of the supplementary corpus of the target corpus data.

[0072] According to the word segmentation result of the target corpus data, determine the weighted word frequency of each different word segmentation in the word segmentation result of the target corpus data.

[0073] Denote the cumulative sum of the number of occurrences of any same word segmentation included in the word segmentation result of the target corpus data as the weighted word frequency of the any same word segmentation.

[0074] Thus, obtain the weighted word frequency of each different word segmentation included in the word segmentation result of the target corpus data.

[0075] Since the word segmentation results of the supplementary corpus of the target corpus data and the word segmentation results of the target corpus data may contain the same word segmentations, update the weighted word frequencies of the same word segmentations included in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data. Specifically: Denote the sum of the weighted word frequencies of the same word segmentations in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data as the weighted word frequency of the same word segmentation in the word segmentation results.

[0076] Thus, obtain the weighted word frequencies of all different word segmentations included in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data.

[0077] The weighted word frequency is the result of word frequency statistics based on the propagation order and relevance between different corpus data, and can more accurately reflect the true semantics expressed by the corpus by combining the characteristics of the public opinion propagation structure of the Internet corpus in the financial information exchange forum, improving the accuracy of word frequency statistics.

[0078] According to the same method, the weighted word frequencies of all different word segmentations included in the word segmentation results of each crawled corpus data and the supplementary corpus of the corpus data can be obtained.

[0079] It can be understood that each strongly associated text of the corpus data corresponds to a corpus data. Therefore, the weighted word frequencies of all different word segmentations included in the word segmentation results of each strongly associated text of the corpus data and the supplementary corpus of the strongly associated text can be obtained.

[0080] Thus, the weighted word frequencies of all different word segmentations included in the word segmentation results of each crawled corpus data and the supplementary corpus of the corpus data are obtained.

[0081] Step S004, according to the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the weighted word frequencies of all word segmentations included in the word segmentation results of the supplementary corpus of all strongly associated texts of the target corpus data, calculate the associated TF value of each same word segmentation, use the associated TF value of the word segmentation as the TF value of the TF-IDF algorithm, and use the TF-IDF algorithm to obtain the word frequency extraction result of the target corpus data, and obtain the real-time public opinion analysis and investment opportunity mining result in stock monitoring.

[0082] Since the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the word segmentation results of the supplementary corpus of all strongly associated texts of the target corpus data may contain the same word segmentations, update the weighted word frequencies of the same word segmentations included in these word segmentation results.

[0083] Specifically, the cumulative sum of the weighted word frequencies of all the same word segments among the weighted word frequencies of all different word segments included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data is used as the weighted word frequency of the same word segment. At the same time, the weighted word frequencies of the same word segments before calculating the cumulative sum in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data are deleted.

[0084] The cumulative sum of the weighted word frequencies of all different word segments included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data is denoted as the total word frequency of word segmentation.

[0085] The ratio of the weighted word frequency of any one word segment included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data to the total word frequency of word segmentation is denoted as the associated TF value of the any one word segment.

[0086] The associated TF values of the word segments included in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data are used as the TF values of the TF-IDF algorithm to obtain the word frequency extraction result of the target corpus data.

[0087] Using the TF-IDF algorithm to obtain the word frequency extraction result of the target corpus data is a well-known technique and will not be elaborated here. Among them, the word frequency extraction result of the target corpus data includes the TF-IDF value of each word segment and a keyword table. In the keyword table, the word segments are arranged in descending order of their TF-IDF values. The greater the TF-IDF value of a word segment, the greater the reliability of the word segment as a keyword of the current Internet public opinion in the financial information exchange forum, and the more likely the meaning expressed by the word segment with a greater TF-IDF value corresponds to information on stock investment opportunities.

[0088] The associated TF value of a word segment is a word frequency value that is more in line with the actual public opinion semantics obtained according to the public opinion dissemination structure. Using the associated TF value of a word segment as the TF value of the TF-IDF algorithm to obtain the word frequency extraction result of the target corpus data can avoid the problem that the TF-IDF algorithm ignores the public opinion dissemination structure information in Internet information, resulting in inaccurate keyword statistics.

[0089] The word frequency extraction results of all crawled corpus data can be obtained in the same way, that is, the real-time public opinion analysis and investment opportunity mining results in stock monitoring.

[0090] Thus, the real-time public opinion analysis and investment opportunity mining results in stock monitoring are obtained.

[0091] Based on the same inventive concept as the above method, an embodiment of the present invention further provides a real-time public opinion analysis and investment opportunity mining system in stock monitoring, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above methods for real-time public opinion analysis and investment opportunity mining in stock monitoring are implemented.

[0092] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for real-time public opinion analysis and investment opportunity mining in stock market monitoring, characterized in that: The method comprises the following steps: Capture all corpus data, each of which includes text data and associated corpus of the text data; Record any corpus data as target corpus data, record any associated corpus of the target corpus data as candidate corpus data, determine text association characteristics of the target corpus data and the candidate corpus data according to the text similarity between the target corpus data and the candidate corpus data, and determine all strongly associated texts of the target corpus data according to all text association characteristics corresponding to the target corpus data and the corpus data corresponding to all associated corpora; Determine the supplementary corpus of the target corpus data according to the target corpus data and other corpus data cited in the associated corpus contained in the target corpus data, obtain the weighted word frequencies of all different segmentations contained in the target corpus data and the segmentation results of the supplementary corpus of the target corpus data according to the number of associated corpora contained in the supplementary corpus of the target corpus data, the segmentation results of the supplementary corpus of the target corpus data, and the segmentation results of the target corpus data, and obtain the weighted word frequencies of all different segmentations contained in the segmentation results of each captured corpus data and the supplementary corpus of the corpus data; According to the weighted word frequencies of all segmented words contained in the segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data, the associated TF value of each identical segmented word is calculated, and the associated TF value of the segmented word is used as the TF value of the TF-IDF algorithm. The TF-IDF algorithm is used to obtain the word frequency extraction results of the target corpus data, and obtain the real-time public opinion analysis and investment opportunity mining results in stock monitoring.

2. The method for real-time public opinion analysis and investment opportunity mining in stock market monitoring according to claim 1 is characterized in that: The specific method of determining the text correlation characteristics of the target corpus data and the to-be-selected corpus data according to the text similarity between the target corpus data and the to-be-selected corpus data is as follows: Perform word segmentation extraction on the target corpus data and the corpus data to be selected, respectively, to obtain the word segmentation results of the target corpus data and the word segmentation results of the corpus data to be selected; record the hash value of the word segmentation result of the target corpus data as the text hash value of the target corpus data; The hash value of the word segmentation result of the corpus data to be selected is recorded as the text hash value of the corpus data to be selected; According to the difference between the text hash values ​​of the target corpus data and the corpus data to be selected, the text correlation characteristics of the target corpus data and the corpus data to be selected are determined.

3. The method for real-time public opinion analysis and investment opportunity mining in stock market monitoring according to claim 2 is characterized in that: The text association characteristics of the target corpus data and the corpus data to be selected are determined according to the difference between the text hash values ​​of the target corpus data and the corpus data to be selected, including the specific method of: The inverse of the Mahalanobis distance between the text hash values ​​of the target corpus data and the selected corpus data is used as the text association feature between the target corpus data and the selected corpus data.

4. The method for real-time public opinion analysis and investment opportunity mining in stock market monitoring according to claim 1, It is characterized in that The method for determining all strongly associated texts of the target corpus data is: The associated corpus corresponding to the maximum value in the text association characteristic of the target corpus data is recorded as the strongly associated text of the target corpus data; Acquire corpus data corresponding to the strongly associated text of the target corpus data, determine the strongly associated text of the corpus data corresponding to the strongly associated text according to the method for acquiring the strongly associated text of the target corpus data, and record the strongly associated text of the corpus data corresponding to the strongly associated text as the strongly associated text of the target corpus data; According to the same method, the strongly associated texts of the target corpus data are continuously obtained according to the strongly associated texts of the new target corpus data, until H strongly associated texts of the target corpus data are obtained, where H represents a first preset threshold.

5. The method for real-time public opinion analysis and investment opportunity mining in stock market monitoring according to claim 1 is characterized in that: The method for determining the supplementary corpus of the target corpus data is as follows: Record other corpus data cited in the associated corpus contained in the target corpus data as indirect reference corpus of the target corpus data; The indirect reference corpus of the target corpus data and the associated corpus contained in the target corpus data are both recorded as the supplementary corpus of the target corpus data.

6. The method for real-time public opinion analysis and investment opportunity mining in stock market monitoring according to claim 1 is characterized in that: The method of obtaining the weighted word frequencies of all different segmentations contained in the target corpus data and the segmentation results of the supplementary corpus of the target corpus data according to the number of associated corpora contained in the supplementary corpus of the target corpus data, the segmentation results of the supplementary corpus of the target corpus data, and the segmentation results of the target corpus data includes the following specific methods: The number of related corpora contained in the supplementary corpus of the target corpus data is recorded as the number of citations of the supplementary corpus of the target corpus data; the cumulative sum of the number of citations of all the supplementary corpora of the target corpus data is recorded as the potential number of citations of the target corpus data; the ratio of the number of citations of the supplementary corpus of the target corpus data to the potential number of citations of the target corpus data is recorded as the citation frequency weight of the supplementary corpus of the target corpus data; Obtaining the word segmentation results of the supplementary corpus of the target corpus data through word segmentation extraction, and recording the number of occurrences of the same word segmentation contained in the word segmentation results of the supplementary corpus as the frequency of the same word segmentation contained in the word segmentation results of the supplementary corpus; The product of the cumulative sum of the frequencies of the same participles contained in the segmentation results of the supplementary corpus of the target corpus data and the weight of the referenced word frequency of the supplementary corpus of the target corpus data is recorded as the weighted word frequency of the same participles contained in the segmentation results of the supplementary corpus of the target corpus data; According to the word segmentation results of the target corpus data, the weighted word frequency of each different word segmentation in the word segmentation results of the target corpus data is determined, and the weighted word frequency of the same word segmentation included in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data is updated.

7. The method for real-time public opinion analysis and investment opportunity mining in stock market monitoring according to claim 6 is characterized in that: The method of determining the weighted word frequency of each different word segmentation in the word segmentation result of the target corpus data according to the word segmentation result of the target corpus data, and updating the weighted word frequency of the same word segmentation contained in the word segmentation results of the target corpus data and the supplementary corpus of the target corpus data includes the following specific methods: The cumulative sum of the number of occurrences of any identical participle contained in the participle results of the target corpus data is recorded as the weighted word frequency of any identical participle; The sum of the weighted word frequencies of the same segmented words in the segmentation results of the target corpus data and the supplementary corpus of the target corpus data is recorded as the weighted word frequency of the same segmented words in the segmentation results.

8. The method for real-time public opinion analysis and investment opportunity mining in stock market monitoring according to claim 1 is characterized in that: The method of calculating the associated TF value of each identical word according to the weighted word frequency of all word segments contained in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data includes the following specific methods: The weighted word frequency of the same participle is the cumulative sum of the weighted word frequencies of all different participles contained in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data, as the weighted word frequency of the same participle; The total frequency of the segmented words is the cumulative sum of the weighted word frequencies of all different segmented words contained in the segmented results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data; The associated TF value of each segmented word is determined based on the weighted word frequency and the total frequency of the segmented words contained in the segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data.

9. The method for real-time public opinion analysis and investment opportunity mining in stock market monitoring according to claim 8 is characterized in that: The method of determining the associated TF value of each segmentation respectively according to the weighted word frequency and the total frequency of the segmentation contained in the segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data includes the following specific methods: The ratio of the weighted term frequency of any word segmented in the word segmentation results of the target corpus data, the supplementary corpus of the target corpus data, all strongly associated texts of the target corpus data, and the supplementary corpus of all strongly associated texts of the target corpus data to the total word segmentation frequency is recorded as the associated TF value of the any word segmented.

10. A real-time public opinion analysis and investment opportunity mining system in stock monitoring, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.