A method, device and equipment for identifying negative public opinion situational awareness risks

By obtaining news and public opinion data, using risk sentiment prediction models to construct negative public opinion indicators and conducting regression analysis, it solves the problem that investors find it difficult to quickly understand risks in the Internet era, and achieves rapid and accurate public opinion risk prediction and defense.

CN119740874BActive Publication Date: 2025-08-19SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411937240.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-08-19
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

In the Internet era, how to provide investors with accurate and readable risk information to help them quickly understand and make decisions, and prevent negative public opinion from causing systemic financial risks.

Method used

By obtaining news and public opinion data, using risk sentiment prediction models to determine risk type probability and emotional information, construct negative public opinion indicators and conduct regression analysis, determine emotional risk indicators and information risk indicators, and ultimately provide public opinion risk description information.

Benefits of technology

It realizes rapid and accurate public opinion risk prediction, helps users to conduct risk warnings and defenses, and reduces market fluctuations caused by investors' irrational behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119740874B_ABST
    Figure CN119740874B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and equipment for identifying the risk of negative public opinion situation perception. The method obtains news public opinion data by responding to an identification request sent by a requesting end; adopts a risk emotion prediction model to predict the news public opinion data and determine the corresponding risk type probability and emotion information; constructs a negative public opinion index based on the risk type probability and emotion information; performs regression analysis on the negative public opinion index to determine the emotion risk index and the information risk index; determines the public opinion risk description information corresponding to the requesting end based on the negative public opinion index, the emotion risk index and the information risk index, thereby quickly and accurately predicting and describing the public opinion risk, which helps users to carry out risk warning and risk defense for the company where the requesting end is located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of risk identification technology, and in particular to a method, device and equipment for identifying risks of negative public opinion situation awareness. Background Art

[0002] In recent years, the rise of internet media and self-media has transformed the traditional landscape of financial news being reported by a few mainstream outlets, significantly transforming the financial media ecosystem. Internet media's rapid dissemination, high frequency of publication, and unpredictability have led to both positive and negative impacts on the securities market in an era of information explosion. On the one hand, it can increase the proportion of informed traders, alleviate information asymmetry, and enhance market efficiency. On the other hand, information disarray can interfere with investor judgment, and the media's subjective sentiment can easily trigger irrational investor behavior, leading to significant stock market volatility.

[0003] Internet media is subject to fewer constraints, often lacking objectivity and displaying more subjective emotions. In the internet age, retail investors primarily access news online, and media sentiment can easily influence investor sentiment and reflect this in stock prices. Negative news can have a greater impact on stock prices, and negative media sentiment can easily cause stock price fluctuations.

[0004] Existing methods for aggregating individual stock sentiment based on media news data provide information support for investors and quantitative fund managers through steps such as news crawling, popularity calculation, preprocessing, and comprehensive sentiment and theme analysis. However, in the internet age, media reprints information rapidly. Beyond aggregating public opinion news, providing investors with accurate and readable risk information to help them quickly understand and make decisions, explore the risks of negative news, and prevent negative public opinion from triggering systemic financial risks has become a key issue. Summary of the Invention

[0005] The present invention provides a method, device and equipment for identifying the risks of negative public opinion situation awareness, which solves the technical problem of how to provide investors with accurate and readable risk information to help them quickly understand and make decisions.

[0006] The first aspect of the present invention provides a method for identifying negative public opinion situation awareness risks, comprising:

[0007] Respond to the identification request sent by the requesting end and obtain news and public opinion data;

[0008] Using a risk sentiment prediction model to predict the news public opinion data, and determine the corresponding risk type probability and sentiment information;

[0009] Constructing a negative public opinion index by weighting the risk type probability and the sentiment information;

[0010] Conducting regression analysis on the negative public opinion indicators to determine sentiment risk indicators and information risk indicators;

[0011] Determine the public opinion risk description information corresponding to the requesting end according to the negative public opinion index, the emotional risk index and the information risk index.

[0012] Optionally, the method further includes:

[0013] Obtain historical news and public opinion data;

[0014] Preprocessing the historical news and public opinion data to obtain multiple key words and form a bag-of-words model;

[0015] Perform one-hot encoding on the bag-of-words model to generate document vectors and label them with risk labels and sentiment labels;

[0016] The document vector is used to train multiple preset machine learning models and calculate prediction effect indicators respectively, and a risk sentiment prediction model is selected according to the prediction effect indicators.

[0017] Optionally, the adopting of the document vector to train multiple preset machine learning models and respectively calculating prediction effect indicators, and selecting a risk sentiment prediction model according to the prediction effect indicators, includes:

[0018] Inputting the document vectors into multiple preset machine learning models respectively to obtain prediction results;

[0019] Compare the prediction results with the risk label and the emotion label respectively, and calculate the loss function value;

[0020] Optimizing parameters of each of the machine learning models according to the loss function value until the loss function value is less than a preset loss threshold, thereby obtaining a plurality of intermediate machine learning models;

[0021] Calculate the prediction effect indicators corresponding to each of the intermediate machine learning models, and select the intermediate machine learning model whose prediction effect indicators meet the selection conditions as the risk sentiment prediction model.

[0022] Optionally, the method further includes:

[0023] When multiple negative public opinion indicators are obtained, the requesting terminals are sorted according to the negative public opinion indicators to obtain a requesting terminal sequence;

[0024] Divide the request end sequence according to a preset number of groups to obtain multiple request end combinations;

[0025] A resource configuration strategy is created to allocate resources to each of the request end combinations according to preset weights.

[0026] Optionally, the use of a risk sentiment prediction model to predict the news public opinion data to determine the corresponding risk type probability and sentiment information includes:

[0027] Convert the news and public opinion data into a vector to be input;

[0028] Inputting the vector to be input into a risk sentiment prediction model, and determining the probability of the risk type corresponding to the vector to be input through the risk prediction branch of the risk sentiment prediction model;

[0029] The emotion probability corresponding to the to-be-input vector is determined as emotion information through the emotion prediction branch of the risk emotion prediction model.

[0030] Optionally, weighting the risk type probability and the sentiment information to construct a negative public opinion index includes:

[0031] Use the analytic hierarchy process to determine the risk weight corresponding to the probability of each risk type;

[0032] Constructing a negative public opinion index according to the risk weight, the risk type probability, and the sentiment information;

[0033] Negative public opinion indicators are:

[0034]

[0035] Among them, risk k,t,m is the probability value of risk type m for requester k on day t; n is the total number of news and public opinion data of requester k on day t; The sentiment information of the a-th news public opinion data of the requesting end k on the t-th day; is the risk type probability of the a-th news public opinion data of the requesting end k on the t-th day belonging to risk type m; θ m is the risk weight of risk type m; b is the total number of risk types; MRI k,t is the total public opinion risk value of the requesting end k on day t.

[0036] Optionally, performing regression analysis on the negative public opinion indicator to determine the emotional risk indicator and the information risk indicator includes:

[0037] Conduct regression analysis using the negative public opinion indicator as the independent variable and the number of daily comments as the dependent variable to construct a regression analysis formula;

[0038] Transforming the regression analysis formula to determine the emotional risk index and the information risk index;

[0039] The regression analysis formula is:

[0040] MRI k,t =a k +b k ×Tpostk,t +e k,t

[0041] The sentiment risk indicators are:

[0042]

[0043] The information risk indicator MRSI k,t for:

[0044] MRSI k,t =MRI k,t -MRII k,t

[0045] Among them, MRII k,t is the information risk index of requester k on day t, MRSI k,t is the sentiment risk index of requester k on day t, a k and b k are the regression coefficients of the request end k, Tpost k,t is the number of daily comments related to the request end k on day t, e k,t is the error term of requester k on day t, and is the estimated value of the regression coefficient.

[0046] Optionally, determining the public opinion risk description information corresponding to the requesting end according to the negative public opinion index, the emotional risk index, and the information risk index includes:

[0047] Calculate the first lagged rate of return according to the negative public opinion indicator and the preset request end parameters;

[0048] Calculating a second lagged rate of return according to the sentiment risk indicator and the information risk indicator in combination with the requesting end parameters;

[0049] The first lagged rate of return and the second lagged rate of return are used to determine the public opinion risk description information corresponding to the requesting end.

[0050] A second aspect of the present invention provides a device for identifying negative public opinion situational awareness risks, comprising:

[0051] The data acquisition module is used to respond to the identification request sent by the requesting end and obtain news and public opinion data;

[0052] A risk sentiment determination module is used to predict the news public opinion data using a risk sentiment prediction model to determine the corresponding risk type probability and sentiment information;

[0053] A negative public opinion index determination module is used to construct a negative public opinion index by weighting the risk type probability and the sentiment information;

[0054] An indicator regression analysis module is used to perform regression analysis on the negative public opinion indicators to determine the emotional risk indicators and information risk indicators;

[0055] The risk identification module is used to determine the public opinion risk description information corresponding to the requesting end based on the negative public opinion index, the emotional risk index and the information risk index.

[0056] The third aspect of the present invention provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor performs the steps of the negative public opinion situation perception risk identification method as described in any one of the first aspects of the present invention.

[0057] It can be seen from the above technical solutions that the present invention has the following advantages:

[0058] The present invention obtains news public opinion data by responding to the identification request sent by the requesting end; adopts the risk emotion prediction model to predict the news public opinion data, and determines the corresponding risk type probability and emotional information; constructs a negative public opinion index according to the risk type probability and emotional information; performs regression analysis on the negative public opinion index to determine the emotional risk index and the information risk index; determines the public opinion risk description information corresponding to the requesting end based on the negative public opinion index, the emotional risk index and the information risk index, thereby quickly and accurately predicting and describing the public opinion risk, which helps users to conduct risk warning and risk defense for the company where the requesting end is located. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0060] Figure 1 A flowchart of the steps of a method for identifying the risk of negative public opinion situation awareness provided by an embodiment of the present invention;

[0061] Figure 2 An ROC curve diagram drawn for the Naive Bayes method in an embodiment of the present invention;

[0062] Figure 3 The average monthly MRI and discounted rate of return graph provided by the embodiment of the present invention;

[0063] Figure 4 A monthly MRI and yield line chart provided by an embodiment of the present invention;

[0064] Figure 5 A schematic diagram of the impact of MRI on stock prices over subsequent days provided by an embodiment of the present invention;

[0065] Figure 6 A schematic diagram of an MRI autocorrelation coefficient provided by an embodiment of the present invention;

[0066] Figure 7 This is a structural block diagram of a negative public opinion situation awareness risk identification device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0067] A prior art method for aggregating sentiment about individual stocks based on media news data has been disclosed. The method includes the following steps: first, crawling news information, generating news documents, and storing them in a document storage database; then, calculating the popularity of each article and removing duplicate documents; preprocessing the content items in the news documents to form a text collection; performing sentiment and topic analysis on each text collection to form a set of two-tuples; then performing text topic clustering and grouping; integrating all relevant financial news to form a set of three-tuples based on individual stocks; and finally, aggregating the above results around individual stocks and presenting them to users. This solution can provide investors in the financial market with accurate, highly readable, and concise topic sentiment information, helping them to understand the information more quickly and make better investment decisions. It also provides important auxiliary information for quantitative fund companies' forecasting models. According to agenda-setting theory in communication studies, media outlets increase the public's importance of news data by reporting it in a biased and repetitive manner. In the internet age, the speed at which information is republished across various media outlets is extremely high. Therefore, in addition to aggregating and classifying public opinion news, it is crucial to provide investors with more accurate and readable risk information, help investors understand and make investment decisions in a shorter time, explore the risks contained in negative news, and prevent the spread of negative public opinion risks from causing systemic financial risks.

[0068] To this end, embodiments of the present invention provide a method, device, and apparatus for identifying negative public opinion situational awareness risks, designed to address the technical problem of providing investors with accurate and readable risk information to help them quickly understand and make decisions. This method strictly monitors the risks inherent in negative public opinion, providing investors with more accurate and readable risk information, helping them understand and make investment decisions more quickly and preventing the spread of negative public opinion risks from causing systemic financial risks.

[0069] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0070] See also Figure 1 , Figure 1 A flowchart of the steps of a method for identifying the risk of negative public opinion situation awareness provided by an embodiment of the present invention.

[0071] The present invention provides a method for identifying negative public opinion situational awareness risks, comprising:

[0072] Step 101: Respond to the identification request sent by the requesting end and obtain news and public opinion data;

[0073] The requesting end refers to a client that has a demand for risk identification and has a user identifier, which can be used to identify an enterprise or user.

[0074] News and public opinion data refers to relevant data involving the enterprise or user, including but not limited to news headlines, full text of news, news release time, news release source, and company stock price. Among them, the source of news release is not limited to a few news platforms, but should cover the entire network, including official reports, financial media platforms, news public accounts, etc.

[0075] In this embodiment, after receiving the identification request sent by the requesting end, news and public opinion data for the day or historical time period is obtained to obtain the data basis for subsequent risk identification.

[0076] In one example of the present invention, the method further includes S11-S14:

[0077] S11. Obtain historical news and public opinion data;

[0078] S12. Preprocess historical news and public opinion data to obtain multiple key words and form a bag-of-words model;

[0079] In this embodiment, the data type of historical news and public opinion data is essentially the same as that of news and public opinion data, except that its data source is historical time periods. After obtaining the relevant historical news and public opinion data from the requesting end, it is preprocessed, such as through data cleaning, text segmentation, part-of-speech tagging, and stop word removal, to obtain multiple key segmentations. Taking each article in the historical news and public opinion data as a unit, the multiple key segmentations associated with it are assembled into a bag-of-words model for subsequent processing.

[0080] In the preprocessing of text segmentation, relevant proper nouns, including company names and financial terms, are imported to improve segmentation accuracy. In the preprocessing of stop word removal, articles and conjunctions that do not contain information in the text are removed. This removal of stop words can further improve the accuracy of text mining.

[0081] For example, after classifying historical news and public opinion data into risk categories, we can use TF-IDF natural language processing to extract the top 20 keywords for each risk category, as shown in Table 1 below:

[0082] Table 1

[0083]

[0084] While a few risk keywords overlap, overall, the keywords for each risk category clearly reflect the underlying risks. For operational risk, terms like "product," "project," and "business" primarily refer to problems with a company's product projects during its operations. Terms like "repayment," "overdue," "debtor," and "investor" indicate problems with a company's financing and investment operations. Broader terms like "finance," "supply chain," "equity," and "capital" are also categorized as operational risk. For accounting and financial risk, the "%" indicates the frequent use of comparisons with previous years in news reports, such as growth in net profit, assets, and liabilities. Terms like "liabilities," "debt-to-asset ratio," "assets," "net profit," "goodwill," and "revenue" are all accounting items and indicators found in financial statements, directly reflecting a company's underlying accounting and financial risks. It's also worth noting that the inclusion of terms like "tax," "deferred," "income tax," "taxes," and "tax bureau" within the accounting and financial risk categories indicates that taxation is also a significant financial concern. Regarding market risks, media reports primarily focus on equity risks, including shareholder increases and decreases, equity buybacks and transfers, and stock pledges. They also cover mergers and acquisitions, such as backdoor listings, acquisitions, control, and restructuring. Regarding legal risks, media reports focus on verdicts, lawsuits, and cases, clearly representing legal issues. These disputes primarily fall into categories such as violations, contract disputes, and bond and debt repayment.

[0085] S13. Perform one-hot encoding on the bag-of-words model to generate document vectors and label them with risk labels and sentiment labels.

[0086] After obtaining the bag-of-words model corresponding to each article, one-hot encoding is used to generate a document vector.

[0087] In addition, the encoding method of document vectors can also refer to natural language processing methods, such as the term frequency-inverse document frequency method TF-IDF and the word2vec method.

[0088] At the same time, the document vectors are annotated with risk labels and sentiment labels.

[0089] Specifically, when labeling public opinion news data with risk tags, the risks involved in public opinion news data can be divided into: business risk, accounting and financial risk, stock market risk, legal and policy risk, other risks, and no risk. The definition of each risk category is shown in Table 2 below:

[0090] Table 2

[0091]

[0092] As for the annotation of sentiment tags, it involves the judgment of positive and negative sentiments of public opinion news data. This embodiment uses the title of the document for modeling. This is because the title of the document is usually a one-sentence summary of the news report, which can condense the most important information of public opinion and usually has more obvious emotional characteristics. In order to extract the key information of the title, the term frequency-inverse document frequency method TF-IDF is used to extract the first 10 keywords of each title (which can be less than ten words), and the extracted parts of speech include four categories: place names, nouns, gerunds, and verbs. All keywords are combined into a bag-of-words model, and one-hot encoding is used to represent the spatial vector of each article.

[0093] This example uses article content as the data source for public opinion risk classification. This is because articles often provide detailed coverage and analysis of company news events, thus containing the most information. To better summarize the meaning of each article, we use the TF-IDF method to obtain the top 50 keywords for each article. These keywords are then combined into a bag-of-words model, and one-hot encoding is used to represent the spatial vector of each article.

[0094] S14. Use document vectors to train multiple preset machine learning models and calculate prediction effect indicators respectively, and select a risk sentiment prediction model according to the prediction effect indicators.

[0095] Furthermore, S14 may include the following sub-steps:

[0096] Input document vectors into multiple preset machine learning models to obtain prediction results;

[0097] Compare the predicted results with the risk label and sentiment label respectively, and calculate the loss function value;

[0098] Optimize the parameters of each machine learning model according to the loss function value until the loss function value is less than the preset loss threshold, and obtain multiple intermediate machine learning models;

[0099] Calculate the prediction effect indicators corresponding to each intermediate machine learning model, and select the intermediate machine learning model whose prediction effect indicators meet the selection criteria as the risk sentiment prediction model.

[0100] Machine learning models include but are not limited to naive Bayesian models, decision trees, random forests, Adaboost, and support vector machines. Prediction performance metrics include but are not limited to precision, recall, F1 score, ROC curve, and AUC value.

[0101] In this embodiment, after completing the initialization of the machine learning models of various model architectures, the labeled document vectors are input into multiple machine learning models to obtain prediction results under different model architectures, which include risk prediction results and sentiment prediction results.

[0102] After obtaining the prediction results, the risk prediction results are compared with the annotated risk labels, and the sentiment prediction results are compared with the sentiment labels. The loss function value is calculated based on the comparison results. If the loss function value is greater than or equal to the preset loss threshold, gradient descent or other parameter adjustment methods are used to adjust the process parameters of each machine learning model until the loss function value is less than the preset loss threshold. At this point, multiple intermediate machine learning models are obtained.

[0103] Furthermore, according to the above-mentioned prediction effect indicator types, the prediction effect indicators of each intermediate machine learning model are calculated respectively, and the intermediate machine learning model whose prediction effect indicator reflects the best model performance is selected as the risk sentiment prediction model.

[0104] For example, the following table 3 shows examples of using precision, recall, and F1 scores as prediction performance indicators:

[0105] Table 3

[0106] Accuracy Precision Recall F1-score Naive Bayes 0.8527 0.8180 0.8482 0.8256 Decision Tree 0.8409 0.8017 0.8205 0.8090 Random Forest 0.8802 0.8624 0.8745 0.8675 Adaboost 0.7466 0.7862 0.7081 0.7153 Support Vector Machine 0.6189 0.8264 0.4806 0.5315

[0107] Among all machine learning models, Naive Bayes, Decision Tree, Random Forest, and Adaboost performed well. Random Forest achieved the highest accuracy, reaching 88.02%. The worst performer was the Support Vector Machine (SVM), with a mere 61.89%. Adaboost and SVM performed poorly, with the SVM's recall rate being only 48.06%, indicating that the SVM model produced a high number of false negatives.

[0108] This example also needs to predict the sentiment value. The precision, recall, F1 score, and AUC value are used to evaluate the prediction effectiveness of the machine learning model. The prediction results of the sentiment are shown in Table 4. Figure 2 Plot the ROC curve for the Naive Bayes method.

[0109] Table 4

[0110] Accuracy Precision Recall F1-score AUC Naive Bayes 0.9273 0.9073 0.9278 0.9164 0.9750 Decision Tree 0.8821 0.8880 0.8282 0.8502 0.9313 Random Forest 0.9037 0.8993 0.8691 0.8821 0.9575 Adaboost 0.8389 0.8106 0.8479 0.8220 0.9244 Support Vector Machine 0.7878 0.8647 0.6552 0.6714 0.9275

[0111] The Naive Bayes model performed best among all machine learning models, achieving over 90% accuracy across all metrics, 92.73% accuracy, and an AUC of 97.5%, demonstrating its practicality in text analysis. The Support Vector Machine model performed the worst, with an accuracy of only 78.78%. This may be because Support Vector Machines are not well suited for text classification with sparse vectors. The decision tree, a basic machine learning algorithm, also delivered very robust results, achieving an accuracy of 88.21%. The Random Forest, a classic ensemble learning algorithm based on bagging, performed better than the decision tree overall, achieving an accuracy of 90.37%. This is due to the bootstrap sampling and voting of multiple decision trees. Adaboost, another classic ensemble learning algorithm based on boosting, was not as effective as the decision tree, achieving an accuracy of only 83.89%. This may be because Adaboost over-focuses on accuracy on the training set, leading to overfitting on the test set.

[0112] Based on the above results, the naive Bayes model can be selected as the emotion prediction branch in the risk emotion prediction model, and the random forest can be selected as the risk prediction branch in the risk emotion prediction model.

[0113] Step 102: Use the risk sentiment prediction model to predict news public opinion data and determine the corresponding risk type probability and sentiment information;

[0114] In one example of the present invention, step 102 may include the following sub-steps:

[0115] Convert news and public opinion data into vectors to be input;

[0116] Input the input vector to the risk sentiment prediction model, and determine the risk type probability corresponding to the input vector through the risk prediction branch of the risk sentiment prediction model;

[0117] The emotion prediction branch of the risk emotion prediction model is used to determine the emotion probability corresponding to the input vector as emotion information.

[0118] In this embodiment, the news and public opinion data is preprocessed and converted into a vector to be input in vector form. The specific process can be seen in S12. The vector to be input is input into the risk sentiment prediction model. The risk prediction branch of the risk sentiment prediction model determines the probability of the risk type corresponding to the vector to be input based on the above-mentioned learned parameters and patterns, such as the market risk probability of 0.4, the reputation risk probability of 0.3, and the other risk probability of 0.3. At the same time, the emotion prediction branch also processes the input vector and outputs the corresponding emotion probability as emotion information, such as the positive emotion probability of 0.2, the negative emotion probability of 0.7, and the neutral emotion probability of 0.1.

[0119] Step 103: weighting the risk type probability and sentiment information to construct a negative public opinion index;

[0120] In one example of the present invention, step 103 may include the following sub-steps:

[0121] The analytic hierarchy process is used to determine the risk weight corresponding to the probability of each risk type;

[0122] Construct negative public opinion indicators based on risk weight, risk type probability, and sentiment information;

[0123] Negative public opinion indicators are:

[0124]

[0125] Among them, risk k,t,m is the probability value of risk type m for requester k on day t; n is the total number of news and public opinion data of requester k on day t; The sentiment information of the a-th news public opinion data of the requesting end k on the t-th day; is the risk type probability of the a-th news public opinion data of the requesting end k on the t-th day belonging to risk type m; θ m is the risk weight of risk type m; b is the total number of risk types; MRI k,t is the total public opinion risk value of the requesting end k on day t.

[0126] In this embodiment, the risk weight corresponding to the probability of each risk type can be determined by the hierarchical analysis method. Specifically, based on the classification of risks, the impact caused by each type of risk is different. For example, accounting and financial risks or judicial policy risks have a greater impact on the company, and these indicators should have a greater weight, while other risks should have a smaller weight. In order to scientifically allocate the weight of each risk category, the hierarchical analysis method is used for weighting. Experts in the field are used to weight different risk types. After the judgment matrix passes the consistency test, the specific weight of each risk is obtained, and the average is used to obtain the overall risk weight. The specific weights are shown in Table 5 below.

[0127] Table 5

[0128] Risk Type Risk Weight Financial Market Risks: 0.235046237210725 Accounting and financial risks: 0.310145369396954 Operational business risks: 0.160831663858335 Legal compliance risks: 0.24377521732951 Other risks: 0.0502015122044757

[0129] After obtaining the risk type probability, sentiment information and risk weight of each article, a negative public opinion index is constructed based on this data.

[0130] Step 104: Perform regression analysis on the negative public opinion indicators to determine the emotional risk indicators and information risk indicators;

[0131] In one example of the present invention, step 104 may include the following sub-steps:

[0132] Using negative public opinion indicators as independent variables and the number of daily comments as dependent variables, we conducted regression analysis and constructed a regression analysis formula.

[0133] Transform the regression analysis formula to determine the emotional risk index and information risk index.

[0134] The Media Risk Index (MRI) can be used in public opinion analysis, communication studies, and other related research or business contexts to measure the media's response and influence to specific events or entities (such as companies or topics). For example, the index is constructed by analyzing multiple factors, including the volume of media coverage, the reach of coverage, and the sentiment of coverage. This numerical representation of the media's overall response helps assess the public opinion risk and social attention faced by the event or entity. It combines information theory and sentiment theory within a unified framework, separating information from sentiment through investor attention. Information in the MRI refers to all negative company-related information on a given day, including the degree of negativity in the news coverage, the weight of the information regarding the risk type, and even the importance of information derived from repeated coverage. Sentiment in the MRI refers to investor overreaction and irrational behavior caused by media hype and repeated coverage. Media buzz leads to heated discussion on social media and heightened investor attention, which in turn drives overreaction and irrational emotions. Investor emotions are divided into careful emotions and focused emotions through investor attention. In this embodiment, MRI can be divided into emotional risk indicators (Media-Risk Sentiment Indicator, MRSI) that reflect emotions and information risk indicators (Media-Risk Information Index, MRII) that reflect information through investor attention.

[0135] In actual implementation, investor attention is usually measured mostly through search indexes, such as the Google Index and the Baidu Index. However, although such measurements can reflect the search popularity of specific keywords over a period of time, they do not intuitively represent the actual behavior of investors.

[0136] To this end, in this embodiment, the number of daily comments on comment communities, such as Tieba, is typically used to measure investor attention. Taking stock investment as an example, comment communities, as social media focused on stock investment, are written by actual investors. This directly reflects the interest and level of discussion surrounding a particular stock or market, reducing the noise associated with irrelevant searches. Furthermore, comment communities have a community effect, forming a network for information dissemination and exchange of opinions among users. Comments are typically published instantly. When a topic or stock is discussed more widely, it may attract more potential investors to join the discussion, further increasing the stock's exposure and attention.

[0137] In this embodiment, for each stock associated with the request end, MRI (ie, the above MRI k,t ) Performed a regression analysis on the number of daily comments in the stock review community and obtained the following analysis results:

[0138] The regression analysis formula is:

[0139] MRI k,t =a k +b k ×Tpost k,t +e k,t

[0140] The sentiment risk indicators are:

[0141]

[0142] Information Risk Index MRI k,t for:

[0143] MRI k,t =MRI k,t -MRII k,t

[0144] Among them, MRII k,t is the information risk index of requester k on day t, MRSI k,t is the sentiment risk index of requester k on day t, a k and b k are the regression coefficients of the request end k, Tpost k,t is the number of daily comments related to the request end k on day t, e k,t is the error term of requester k on day t, and is the estimated value of the regression coefficient

[0145] For example, news and public opinion data can cover the period from year A, month B, and day C to year D, month E, and day F. For example, 3,285,120 media reports collected from 453 media platforms cover 250 listed companies, all of which may face the risk of a stock price crash due to the outbreak of negative news. News sources include all well-known self-media official account platforms and major financial media platforms. The data includes article titles, text content, publication dates, and corresponding platform information. For each news report, the MRI construction calculation shown in step 103 above is performed to obtain daily MRI indicators for these 250 companies, totaling 795,750 sample data. To separate information and sentiment in the MRI, data from stock comment communities on specific websites is also collected, including the daily number of comments (Tpostnum), which includes the number of positive comments (Pospostnum) and the number of negative comments (Negpostnum). Using the daily number of comments as a proxy for investor attention (Dong et al.), the aforementioned regression analysis formula is used to decompose the MRI into MRSI and MRII.

[0146] The explained variable is the company's daily return rate. The daily stock opening prices are collected from the CSMAR database. The stock's daily return rate (r) is calculated by calculating the opening prices of two consecutive days. In order to intuitively show the relationship between return rate and MRI, Figure 3 The chart shows the average monthly MRI (including MRSI and MRII) and yield discount chart of 250 stocks. At the same time, it finds the companies that have the most news among the 250 stocks. Figure 4 Plot its monthly MRI and yield line chart. Comparison reveals that the yield and MRI of a single stock exhibit greater volatility and extreme values compared to the market average, particularly during 2016-2017, when this stock was suspended for rectification, during which time it experienced significant fluctuations in media risk and yield. There is a clear negative correlation between MRI and yield, and a clear positive correlation between MRI and MRII and MRSI. Table 6 shows the correlation between its yield r, MRI, MRII, MRSI, and investor attention (IA):

[0147] Table 6

[0148] r MRI MRII MRSI IA r 1.00 MRI -0.03 1.00 MRII -0.02 0.76 1.00 MRSI -0.02 0.65 0.00 1.00 IA -0.03 0.33 0.00 0.51 1.00

[0149] Table 6 shows that the correlation between return and investor attention is significantly negative, indicating that falling stock prices are more likely to attract strong investor attention than rising stock prices. The correlation between MRI and return is significantly negative, while the correlation between MRI and IA is significantly positive. This suggests that falling stock prices can lead to negative media coverage, which in turn attracts investor attention. Furthermore, since MRII is the portion of MRI that cannot be explained by IA, the correlation between MRII, IA, and MRSI is zero.

[0150] Step 105: Determine the public opinion risk description information corresponding to the requesting end based on the negative public opinion index, the emotional risk index, and the information risk index.

[0151] In one example of the present application, step 105 may include the following sub-steps:

[0152] Calculate the first lag rate of return based on the negative public opinion index and the preset request-side parameters;

[0153] Calculate the second lagged rate of return based on the sentiment risk index and information risk index, combined with the requester parameters;

[0154] The first lagged rate of return and the second lagged rate of return are used to determine the public opinion risk description information corresponding to the requesting end.

[0155] In this embodiment, the percentile values and maximum values corresponding to the negative public opinion indicators over a period of time can be calculated. According to the percentile values and the maximum values, the public opinion characteristic types are matched respectively. For example, as shown in Table 7 below, the 25th percentile and the 50th percentile are both 0, indicating that negative news has the characteristics of concentrated outbreak and rapid spread. And because MRI is composed of MRII and MRSI, and MRII is the residual term after MRI is regressed on media coverage intensity, the average value of MRII is 0, and the mean value of MRSI is equal to the mean value of MRI. At the 25th, 50th and 75th percentiles, each media risk is 0, which shows that the risks of various media are relatively sparse. As shown in Table 7 below:

[0156] Table 7

[0157] Obs. Mean Std.Dev. Min Max 25% 50% 75% r(%) 795,750 0.025 0.037 -80.732 47.058 -1.615 0.040 1.567 MRI 795,750 0.294 1.784 0 204.789 0 0 0.014 MRII 795,750 0.000 1.362 -123.806 202.366 -0.156 -0.044 0 MRSI 795,750 0.294 1.151 -2.705 180.978 0.016 0.076 0.232 Operational risk 795,750 0.030 0.152 0 19.040 0 0 0 Accounting risk 795,750 0.075 0.341 0 64.473 0 0 0 Stock market risk 795,750 0.116 0.425 0 48.667 0 0 0 Legal policy risk 795,750 0.125 0.476 0 57.128 0 0 0 Other risks 795,750 0.041 0.194 0 23.14 0 0 0 MKT_RF(%) 2,148 0.040 0.016 -9.478 9.165 -0.626 0.118 0.802 SMB (%) 2,148 0.028 0.008 -5.848 4.051 -0.365 0.091 0.482 HML (%) 2,148 0.011 0.007 -4.267 3.615 -0.461 -0.019 0.420 BIG4 1,972 0.048 0.214 0 0 0 0 0 SOE 1,972 0.254 0.436 0 1 0 0 1 TOP5(%) 1,972 49.006 15.892 6.907 97.456 37.519 48.633 59.369 POI (%) 1,972 40.846 23.034 0.001 98.125 22.685 40.098 57.745 SIZE 1,972 22.104 1.533 16.649 29.218 21.183 21.880 22.783 BM 1,972 0.391 0.646 -2.582 14.022 0.159 0.295 0.498 ROA 1,972 -0.069 1.302 -48.316 1.408 -0.012 0.017 0.045 LEV 1,972 2.155 9.116 -56.445 275.340 0.386 0.883 2.020 Analyst 1,972 0.951 1.147 0 4 0 0 2 Tpost 795,750 42.837 109.922 1 18,855 6 17 43 Pospost 795,750 10.989 26.532 0 3,919 2 5 11 Negpost 795,750 9.596 27.210 0 4,839 1 3 10

[0158] As shown in Table 4, in addition to the MRI, it is necessary to control for market and company-level factors. Therefore, the control variables are divided into two categories: the first category comprises daily time series data, such as the Fama-French three-factor model; the second category comprises annual company panel data, including whether the company is audited by one of the Big Four audit firms (BIG4), whether it is a specific type of enterprise (SOE), the shareholding ratio of the top five shareholders (TOP5), the shareholding ratio of institutional investors (POI), the logarithm of total assets (SIZE), the price-to-book ratio (MTB), the return on assets (ROA), the leverage ratio (LEV), and the level of analyst attention (Analyst). The minimum value of the company's stock return is -80.732%, and the maximum value is 47.058%, indicating that penalized listed companies are more likely to experience extreme market conditions. The maximum value of the MRI (Media Reaction Index) is 204.789, and its 25th and 50th percentiles are both 0, indicating that negative news is characterized by concentrated outbreaks and rapid dissemination. Because MRI is composed of MRII and MRSI, and MRII is the residual after regressing MRI on media coverage intensity, the mean of MRII is 0, and the mean of MRSI is equal to the mean of MRI. At the 25th, 50th, and 75th percentiles, each media risk is 0, indicating that the risk of each media type is relatively sparse.

[0159] Overreaction and underreaction are defined as the correlation between a company's initial daily returns after a news event and its daily returns after the event. A positive correlation indicates underreaction, while a negative correlation indicates overreaction. Existing technologies mainly focus on the impact of individual news events released by a few authoritative media on stock prices. Therefore, event study methods are usually used to assess the impact of these specific events. However, in actual implementation, a stock may be affected by multiple related news articles and multiple events in a single day. In this case, the event study method has difficulty capturing the overall impact of all these concurrent events. In addition, the event study method often relies on a predetermined time window, which is not flexible enough for frequently updated information flows and cannot be appropriately adjusted according to needs. MRI is a daily frequency indicator obtained by summing the risks of all relevant news each day. It can capture the overall trend of media coverage and its long-term impact on a company's stock price. This supports the use of a more flexible linear regression to examine the impact of media risk on stock prices. The first lag return is:

[0160]

[0161] The second lagged rate of return is:

[0162]

[0163] Among them, r is the return rate of the requesting end k at the current time t lagged w days, α is the preset parameter, is a request-side parameter. Factor is the Fama-French three-factor model, which controls for the impact of market factors. Control is other company-level control variables, including BIG4, SOE, TOP5, POI, SIZE, BM, ROA, LEV, and Analyst. Furthermore, we control for year fixed effects (Year) to prevent endogeneity caused by omitted variables. The above formula can be used to determine the impact of MRI, MRII, and MRSI on stock prices, yielding the following Table 8:

[0164] Table 8

[0165]

[0166] The t-statistics are in brackets. *: p < 0.10, **: p < 0.05, ***: p < 0.01

[0167] Note: This table describes the impact of MRI (Media Reaction Index), MRII (Media Risk Information Index), and MRSI (Media Risk Sentiment Index) on stock returns. Fama-French factors and control variables are also included in the explanatory variables. Columns (1) and (2) test the impact of MRI on the stock price on the same day, where the dependent variable is rk,t Columns (3) and (4) investigate the effect of MRI on the next day's stock price, where the dependent variable is r k,t+1 e-2 represents 10 of the reported figures -2 , for example, -0.044 in column (1) means -0.044%.

[0168] As shown in Table 8, first, the stock price does not react enough to MRI: the first column shows that the impact of MRI on the stock return rate on the same day is significantly negative (coefficient = -0.044, t-stat = -26.68, and the third column shows that the impact of MRI on the stock return rate on the next day is also significantly negative (coefficient = -0.021, t-stat = -12.00), and the two have the same signs, so the stock price does not react enough to MRI. In addition, for the description of public opinion risk information, the duration of the emotional effect and the duration of the information effect can be further analyzed through the emotional risk index and the information risk index. For example, the impact of emotion on stock price is greater than the impact of information: the second column shows that in the media reports on the same day, MRSI has a significant impact on stock price. The effect of MRSI on the stock price is -0.053 (t-stat = -20.78), which is greater than the coefficient of the MRII's effect on the stock price, -0.038 (t-stat = -17.38). Finally, the impact of sentiment lasts longer, while information is absorbed more quickly: In the fourth column, the impact of MRSI on the stock price the next day decreases from -0.053 to -0.040 (t-stat = -14.67), and the impact of MRII on the stock price the next day decreases from -0.038 to -0.007 (t-stat = -3.32). This shows that the influence of sentiment (MRSI) lasts until the next day, while the influence of information (MRII) is significantly weakened the next day, indicating that information is absorbed more quickly by the stock price.

[0169] To better demonstrate the impact of MRI on stock prices in subsequent days, Figure 5 This figure shows the coefficients of the MRI (Media Reaction Index), MRSI (Media Risk Sentiment Index), and MRII (Media Risk Information Index) on stock returns, lagged 20 trading days, along with their 95% confidence intervals. The x-axis represents the number of lagged days, corresponding to w in the above formula. Each point in the figure represents the coefficient of MRI / MRSI / MRII derived from the regression analysis, and the vertical line at each point represents the 95% confidence interval. Figure 5The MRI, MRII, and MRSI coefficients (95% confidence intervals) for the subsequent 20 trading days are presented. This chart demonstrates that stock prices underreact to MRI. Over time, the impact of MRI on subsequent stock prices gradually converges to zero. Meanwhile, the impact of the information component of MRI on stock prices has already converged to zero two days after the news report. Even when the coefficient exceeds 0 after period t+10, the confidence intervals indicate that the hypothesis of the coefficient being equal to 0 cannot be rejected. Compared to MRII, MRSI has a persistent impact on stock prices. The emotional impact of repeated media coverage continues to ferment, and the impact on stock price returns after period t+10 is still -0.02.

[0170] like Figure 6 As shown in Figure 1, this figure shows the autocorrelation coefficients of MRI (Media Reaction Index), MRSI (Media Risk Sentiment Index), and MRII (Media Risk Information Index). The x-axis represents the number of lagged days, corresponding to w in the above formula. The y-axis shows the correlation coefficient between the current indicator and the same indicator lagged w days, from Figure 6 As can be seen in the figure, the autocorrelation coefficient of MRSI remains high over longer time lags, especially after lag 10, where it remains significantly above 0.5. In contrast, the autocorrelation coefficients of MRI and MRII decrease rapidly over shorter time lags, especially for MRII, where the autocorrelation coefficient almost drops below 0.2 after lag 4. This further indicates that most information in the media is independent and novel, while emotions in the media are contagious and last longer.

[0171] Taking into account the autocorrelation of the media, a panel vector autoregression model is used to further determine the impact of MRI on stock prices. Compared with the direct linear regression model, PVAR can not only handle the dynamic relationship between multiple variables, but also effectively capture the lag effect of time series, control cross-sectional heterogeneity, and identify causal relationships, thereby providing more accurate and comprehensive results. Figure 3 and Figure 4 As can be seen in the figure, MRI (including MRSI and MRII) enter a relatively stable state after a four-day lag. This is likely because the market is only open on weekdays, and the four-day lag encompasses a full work week. Therefore, a fourth-order lag operator, L4, can be defined: L4(x_t) = [x_(t-1), x_(t-2), x_(t-3), x_(t-4)]. The regression formula is shown above, and the regression results are reported in Table 7. It should also be noted that because the PVAR model deals with the dynamic interactions between multiple endogenous variables, the presence of the lag term makes the traditional R2 concept no longer applicable. Therefore, the chi-square statistic is used to assess the overall goodness of fit and significance of the model and to identify causal relationships:

[0172] rk,t =α1+β1L4(r k,t )+β2L4(MRI k,t )+ε 1,t

[0173] MRI k,t =α1+β1L4(r k,t )+β2L4(MRI k,t )+ε 2,t

[0174] The calculation results are shown in Table 9:

[0175] Table 9

[0176]

[0177]

[0178] T-statistics are in parentheses. *p<0.10, **p<0.05, ***p<0.01

[0179] Note: This table describes a vector autoregression model of four-day lagged returns and the MRI (Media Reaction Index). It uses the current day's returns and the MRI as dependent variables, and the previous four days' returns and the MRI as independent variables, respectively. Where e-2 represents numbers reported in units of 10-2, for example, an MRI of -0.016 in column (2) means -0.016%.

[0180] As can be seen in the regression formula, the first lag of MRI is significantly negative (coeffcient = -0.016, t-stat = -4.70), while the second, third, and fourth lags are insignificant. Considering that their signs are consistent with the effect of MRI on same-day stock prices in Table 3, this further confirms that stock prices are underreacting to MRI, and this underreaction only occurs in media coverage of the previous day. In the MRI regression formula, the first lag of return is significantly negative at first order (coeffcient = -0.196, t-stat = -2.02). This finding suggests that despite the MRI's underreaction to stock prices, the news media also reported on the risk of a stock price decline the previous day. The chi-square statistic shows that MRI (especially MRIt-1) has a significant Granger causal effect on stock price returns, while stock price returns (mainly rt-1) have a weaker causal effect on MRI.

[0181] Table 10

[0182]

[0183]

[0184] This table describes vector autoregression models of four-day lagged returns and the MRI (Media Reaction Index) / MRSI (Media Risk Sentiment Index). For information effects, the MRII (Media Risk Information Index) and the PVAR (Vector Autoregression Model) of returns are used. For sentiment effects, the MRSI and the PVAR of returns are used. e-2 represents numbers reported in units of 10-2.

[0185] In addition, we examined the information effect (MRII) and the sentiment effect (MRSI) in the MRI model. The results are shown in Table 10. For the information effect, a vector autoregression model of the MRII and the return rate was used; for the sentiment effect, the MRSI was substituted. In the regression of the return rate, the first lag of the MRII was negative (coefficient = -0.006, t-stat = -1.96), consistent with the sign in Table 6, indicating that stock prices underreacted to the information component of media risk. This underreaction was small and lagged only one day, indicating that the information was quickly absorbed by the stock price. On the other hand, the first lag of the MRSI was negative (coefficient = -0.046, t-stat = -4.61), while the fourth lag was positive (coefficient = 0.015, t-stat = 1.96), indicating that stock prices initially underreacted to the sentiment component of media risk but subsequently overreacted. In the regression of the MRII, the first lagged return variable is -0.151 (t-stat = -1.75), and the third lagged return variable is -0.155 (s-stat = -2.04), indicating that media coverage also considers previous stock price declines as negative information. In the regression of the MRSI, the second lagged return variable is significantly positive (coefficient = 0.146, t-stat = 3.95), indicating that stock prices overreact to the MRSI. The chi-square statistic shows that the MRSI has a more significant causal relationship with the return rate than the MRII.

[0186] In one example of the present invention, the method further comprises the following steps:

[0187] When multiple negative public opinion indicators are obtained, the requesting ends are sorted according to the negative public opinion indicators to obtain a requesting end sequence;

[0188] Divide the request end sequence according to a preset number of groups to obtain multiple request end combinations;

[0189] Create a resource allocation policy to allocate resources to each request end combination according to preset weights.

[0190] In this embodiment, the following method is used to illustrate how the resource configuration policy can effectively allocate resources.

[0191] Taking stock investment as an example, as shown in Table 7, the mean proportion of companies audited by the Big Four (BIG4) is 0.048, indicating that 4.8% of companies are audited by them. The mean proportion of SOEs (Specialized Enterprises) is 0.254, indicating that approximately 25% of companies are SOEs. The mean return on assets (ROA) is -0.069, suggesting that the profitability of penalized companies is relatively weak. Furthermore, variables such as the shareholding ratio of the top five shareholders (TOP5), the shareholding ratio of institutional investors (POI), the logarithm of total assets (SIZE), the price-to-book ratio (MTB), the leverage ratio (LEV), and the level of analyst attention (Analyst) provide information on the governance structure, capital structure, and market attention of these companies. These statistics reflect the characteristics of the sample companies across various dimensions and are of great significance for further analysis of the relationship between media risk and stock price volatility.

[0192] The above analysis confirmed that stock prices under-react to the MRI (Market Reaction Index), which can be used to predict the stock price the following day. To test the predictive power of the MRI, a long-short portfolio was constructed based on the MRI and its effectiveness in asset pricing was examined. 250 companies were sorted from high to low based on their MRI the previous day and divided into five portfolios. Table 11 reports the Jensen alpha and risk loading of the MRI quintile portfolios under the CAPM and Fama-French three-factor models. It is worth noting that because the MRI has a negative predictive effect on stock prices, the high group is subtracted from the low group when constructing the long-short portfolio, resulting in a high probability of achieving positive excess returns:

[0193] Table 11

[0194]

[0195]

[0196] T-statistics are in parentheses. *p < 0.10, **p < 0.05, ***p < 0.01. This table divides all stocks into five groups based on the previous day's MRI (Media Reaction Index) and examines the abnormal returns of these five stock portfolios. Panel A directly examines the abnormal returns of different stock portfolios and long-short portfolios, while Panels B and C consider the CAPM model and Fama's three-factor model, respectively.

[0197] Table 11 demonstrates the empirical results of the MRI's predictive power. Although only the high-order groups in Panels B and C exhibit significant alpha, all three regression analyses generate substantial positive alpha for the long-short portfolio. The daily excess return for the long-short portfolio exceeds 0.05% per day, or 12.5% per year (assuming a 250-trading-day year). Descriptive statistics reveal that the MRI occurs at the 50th percentile, meaning that a company's MRI is zero more than half the time. Consequently, no significant decreasing alpha trend is observed in the third, fourth, and lowest quintiles. However, significant negative alpha is observed in the top and second quintiles, with the magnitude of the negative alpha being greater in the top quintile. These findings suggest that common factors cannot fully explain the return differences between portfolios with different MRIs. Investors can utilize MRI to construct appropriate resource allocation strategies, such as investment strategies, to achieve returns.

[0198] In an embodiment of the present invention, news public opinion data is obtained by responding to an identification request sent by a requesting end; a risk sentiment prediction model is used to predict the news public opinion data to determine the corresponding risk type probability and sentiment information; weighting is performed according to the risk type probability and sentiment information to construct a negative public opinion index; regression analysis is performed on the negative public opinion index to determine the sentiment risk index and the information risk index; based on the negative public opinion index, the sentiment risk index and the information risk index, the public opinion risk description information corresponding to the requesting end is determined, thereby quickly and accurately predicting and describing the public opinion risk, which helps users to conduct risk warnings and risk defense for the company where the requesting end is located.

[0199] See also Figure 7 , Figure 7 The present invention provides a block diagram of the structure of a device for identifying the risk of negative public opinion situation awareness.

[0200] An embodiment of the present invention provides a device for identifying negative public opinion situational awareness risks, comprising:

[0201] The data acquisition module 701 is used to respond to the identification request sent by the requesting end and obtain news and public opinion data;

[0202] The risk sentiment determination module 702 is used to predict news public opinion data using a risk sentiment prediction model to determine the corresponding risk type probability and sentiment information;

[0203] Negative public opinion index determination module 703, used to construct a negative public opinion index by weighting the risk type probability and sentiment information;

[0204] The indicator regression analysis module 704 is used to perform regression analysis on negative public opinion indicators to determine emotional risk indicators and information risk indicators;

[0205] The risk identification module 705 is used to determine the public opinion risk description information corresponding to the requesting end based on the negative public opinion index, the emotional risk index and the information risk index.

[0206] Optionally, the device further comprises:

[0207] Historical data acquisition module, used to obtain historical news and public opinion data;

[0208] The data preprocessing module is used to preprocess historical news and public opinion data, obtain multiple key words and form a bag-of-words model;

[0209] The vector generation module is used to perform one-hot encoding on the bag-of-words model, generate document vectors, and annotate risk and sentiment labels;

[0210] The model determination module is used to determine the public opinion risk description information corresponding to the requesting end based on negative public opinion indicators, emotional risk indicators and information risk indicators.

[0211] Optionally, the model determination module is specifically configured to:

[0212] Input document vectors into multiple preset machine learning models to obtain prediction results;

[0213] Compare the predicted results with the risk label and sentiment label respectively, and calculate the loss function value;

[0214] Optimize the parameters of each machine learning model according to the loss function value until the loss function value is less than the preset loss threshold, and obtain multiple intermediate machine learning models;

[0215] Calculate the prediction effect indicators corresponding to each intermediate machine learning model, and select the intermediate machine learning model whose prediction effect indicators meet the selection conditions as the risk sentiment prediction model.

[0216] Optionally, the device further comprises:

[0217] A sorting module is used to sort the requesting terminals according to the negative public opinion indicators when multiple negative public opinion indicators are obtained to obtain a requesting terminal sequence;

[0218] A combination division module is used to divide the request end sequence according to a preset number of groups to obtain multiple request end combinations;

[0219] The resource configuration module is used to create resource configuration strategies to allocate resources to each request end combination according to preset weights.

[0220] Optionally, the risk sentiment determination module 702 is specifically configured to:

[0221] Convert news and public opinion data into vectors to be input;

[0222] Input the input vector to the risk sentiment prediction model, and determine the risk type probability corresponding to the input vector through the risk prediction branch of the risk sentiment prediction model;

[0223] The emotion prediction branch of the risk emotion prediction model is used to determine the emotion probability corresponding to the input vector as emotion information.

[0224] Optionally, the negative public opinion index determination module 703 is specifically configured to:

[0225] The analytic hierarchy process is used to determine the risk weight corresponding to the probability of each risk type;

[0226] Construct negative public opinion indicators based on risk weight, risk type probability, and sentiment information;

[0227] Negative public opinion indicators are:

[0228]

[0229] Among them, risk k,t,m is the probability value of risk type m for requester k on day t; n is the total number of news and public opinion data of requester k on day t; The sentiment information of the a-th news public opinion data of the requesting end k on the t-th day; is the risk type probability of the a-th news public opinion data of the requesting end k on the t-th day belonging to risk type m; θ m is the risk weight of risk type m; b is the total number of risk types; MRI k,t is the total public opinion risk value of the requesting end k on day t.

[0230] Optionally, the indicator regression analysis module 704 is specifically used to:

[0231] Using negative public opinion indicators as independent variables and the number of daily comments as dependent variables, we conducted regression analysis and constructed a regression analysis formula.

[0232] Transform the regression analysis formula to determine the emotional risk index and information risk index;

[0233] The regression analysis formula is:

[0234] MRI k,t =a k +b k ×Tpost k,t +e k,t

[0235] The sentiment risk indicators are:

[0236]

[0237] Information Risk Index (MRSI) k,t for:

[0238] MPSI k,t =MRI k,t -MRII k,t

[0239] Among them, MRII k,t is the information risk index of requester k on day t, MRSI k,t is the sentiment risk index of requester k on day t, a k and b k are the regression coefficients of the request end k, Tpost k,t is the number of daily comments related to the request end k on day t, e k,t is the error term of requester k on day t, and is the estimated value of the regression coefficient.

[0240] Optionally, the risk identification module 705 is specifically configured to:

[0241] Calculate the first lag rate of return based on the negative public opinion index and the preset request-side parameters;

[0242] Calculate the second lagged rate of return based on the sentiment risk index and information risk index, combined with the requester parameters;

[0243] The first lagged rate of return and the second lagged rate of return are used to determine the public opinion risk description information corresponding to the requesting end.

[0244] An embodiment of the present invention provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor performs the steps of the negative public opinion situation perception risk identification method as described in any embodiment of the present invention.

[0245] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0246] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0247] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.

[0248] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules.

[0249] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying negative public opinion situational awareness risks, characterized in that: include: Respond to the identification request sent by the requesting end and obtain news and public opinion data; Using a risk sentiment prediction model to predict the news public opinion data, and determining the corresponding risk type probability and sentiment information; Constructing a negative public opinion index by weighting the risk type probability and the sentiment information; Conducting regression analysis on the negative public opinion indicators to determine sentiment risk indicators and information risk indicators; Determining public opinion risk description information corresponding to the requesting end according to the negative public opinion index, the emotional risk index, and the information risk index; The weighting according to the risk type probability and the sentiment information to construct a negative public opinion index includes: Use the analytic hierarchy process to determine the risk weight corresponding to the probability of each risk type; Constructing a negative public opinion index according to the risk weight, the risk type probability, and the sentiment information; Negative public opinion indicators are: ; ; in, is the probability value of risk type m for requester k on day t; n is the total number of news and public opinion data of requester k on day t; The sentiment information of the a-th news public opinion data of the requesting end k on the t-th day; The risk type probability of the a-th news and public opinion data of the requesting end k on the t-th day belongs to the risk type m; is the risk weight of risk type m; b is the total number of risk types; is the total public opinion risk value of requester k on day t; The performing of regression analysis on the negative public opinion indicators to determine the emotional risk indicators and information risk indicators includes: Conduct regression analysis using the negative public opinion indicator as the independent variable and the number of daily comments as the dependent variable to construct a regression analysis formula; Transforming the regression analysis formula to determine the emotional risk index and the information risk index; The regression analysis formula is: ; The sentiment risk indicators are: ; The information risk indicators for: ; in, is the information risk index of requester k on day t, is the sentiment risk index of requester k on day t, and are the regression coefficients of the request end k, is the number of daily comments related to requester k on day t, is the error term of requester k on day t, and is the estimated value of the regression coefficient.

2. The method according to claim 1, characterized in that The method further comprises: Obtain historical news and public opinion data; Preprocessing the historical news and public opinion data to obtain multiple key words and form a bag-of-words model; Perform one-hot encoding on the bag-of-words model to generate document vectors and label them with risk labels and sentiment labels; The document vector is used to train multiple preset machine learning models and calculate prediction effect indicators respectively, and a risk sentiment prediction model is selected according to the prediction effect indicators.

3. The method according to claim 2, characterized in that The document vector is used to train multiple preset machine learning models and respectively calculate prediction effect indicators, and the risk sentiment prediction model is selected according to the prediction effect indicators, including: Inputting the document vectors into multiple preset machine learning models respectively to obtain prediction results; Compare the prediction results with the risk label and the emotion label respectively, and calculate the loss function value; Optimizing parameters of each of the machine learning models according to the loss function value until the loss function value is less than a preset loss threshold, thereby obtaining a plurality of intermediate machine learning models; Calculate the prediction effect indicators corresponding to each of the intermediate machine learning models, and select the intermediate machine learning model whose prediction effect indicators meet the selection conditions as the risk sentiment prediction model.

4. The method according to claim 1, wherein The method further comprises: When multiple negative public opinion indicators are obtained, the requesting terminals are sorted according to the negative public opinion indicators to obtain a requesting terminal sequence; Divide the request end sequence according to a preset number of groups to obtain multiple request end combinations; A resource configuration strategy is created to allocate resources to each of the request end combinations according to preset weights.

5. The method according to claim 1, wherein The risk sentiment prediction model is used to predict the news public opinion data to determine the corresponding risk type probability and sentiment information, including: Convert the news and public opinion data into a vector to be input; Inputting the vector to be input into a risk sentiment prediction model, and determining the probability of the risk type corresponding to the vector to be input through the risk prediction branch of the risk sentiment prediction model; The emotion probability corresponding to the to-be-input vector is determined as emotion information through the emotion prediction branch of the risk emotion prediction model.

6. The method according to claim 1, characterized in that The determining, based on the negative public opinion index, the emotional risk index, and the information risk index, the public opinion risk description information corresponding to the requesting end includes: Calculate the first lagged rate of return according to the negative public opinion indicator and the preset request end parameters; Calculating a second lagged rate of return according to the sentiment risk indicator and the information risk indicator in combination with the requesting end parameters; The first lagged rate of return and the second lagged rate of return are used to determine the public opinion risk description information corresponding to the requesting end.

7. A device for identifying negative public opinion situational awareness risks, characterized in that: include: The data acquisition module is used to respond to the identification request sent by the requesting end and obtain news and public opinion data; A risk sentiment determination module is used to predict the news public opinion data using a risk sentiment prediction model to determine the corresponding risk type probability and sentiment information; A negative public opinion index determination module is used to construct a negative public opinion index by weighting the risk type probability and the sentiment information; An indicator regression analysis module is used to perform regression analysis on the negative public opinion indicators to determine the emotional risk indicators and information risk indicators; A risk identification module, configured to determine the public opinion risk description information corresponding to the requesting end based on the negative public opinion index, the emotional risk index, and the information risk index; The negative public opinion index determination module is specifically used to: Use the analytic hierarchy process to determine the risk weight corresponding to the probability of each risk type; Constructing a negative public opinion index according to the risk weight, the risk type probability, and the sentiment information; Negative public opinion indicators are: ; ; in, is the probability value of risk type m for requester k on day t; n is the total number of news and public opinion data of requester k on day t; The sentiment information of the a-th news public opinion data of the requesting end k on the t-th day; The risk type probability of the a-th news and public opinion data of the requesting end k on the t-th day belongs to the risk type m; is the risk weight of risk type m; b is the total number of risk types; is the total public opinion risk value of requester k on day t; The indicator regression analysis module is specifically used for: Conduct regression analysis using the negative public opinion indicator as the independent variable and the number of daily comments as the dependent variable to construct a regression analysis formula; Transforming the regression analysis formula to determine the emotional risk index and the information risk index; The regression analysis formula is: ; The sentiment risk indicators are: ; The information risk indicators for: ; in, is the information risk index of requester k on day t, is the sentiment risk index of requester k on day t, and are the regression coefficients of the request end k, is the number of daily comments related to requester k on day t, is the error term of requester k on day t, and is the estimated value of the regression coefficient.

8. An electronic device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the negative public opinion situation perception risk identification method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Enterprise negative public opinion intelligent risk identification and index construction method

    CN116227909A

  • Stock card stopping risk early warning method and system

    CN118071493A