Network risk comprehensive assessment index calculation method

By combining a large language model and a dynamic weighting mechanism, the problems of data utilization and weight allocation in network risk assessment are solved, enabling accurate assessment of online public opinion and improving the accuracy and scientific nature of the assessment.

CN121125332APending Publication Date: 2025-12-12中共重庆市委网络安全和信息化委员会办公室
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511560407.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing technologies for comprehensive network risk assessment are inadequate in terms of data utilization, weight allocation, and information analysis. They are unable to accurately determine the degree of risk of online public opinion, especially since they neglect implicit information and fixed weights cannot adapt to the dynamic and uncertain nature of online public opinion.

Method used

A large language model is used to classify network data. A Markov decision process and Q-learning algorithm are combined to generate dynamic weights. The network data is subdivided into eight categories and risk index scores are calculated. The network risk comprehensive assessment index is calculated by weighting the data with dynamic weights.

Benefits of technology

It improves the accuracy and reliability of comprehensive network risk assessment, enabling more precise identification of implicit semantics and relationships in network data. The dynamic weighting mechanism adapts to the risk characteristics of different event types, enhancing the comprehensiveness and scientific rigor of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125332A_ABST
    Figure CN121125332A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network risk management, in particular to a network risk comprehensive assessment index calculation method, which comprises the following steps: S1, acquiring a plurality of pieces of network data associated with a target network hotspot event; s2, classifying each piece of network data through a large language model to obtain a prediction category of each piece of network data; s3, calculating a corresponding risk index score for the network data of each prediction category; s4, generating a corresponding dynamic weight for each prediction category based on the event type of the target network hotspot event by combining a Markov decision process with a Q-learning algorithm; and S5, based on the dynamic weight of each prediction category, carrying out weighted calculation on the risk index score of the network data of each prediction category to obtain a network risk comprehensive assessment index of the target network hotspot event. According to the invention, the accuracy and reliability of network risk comprehensive assessment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network risk technology, specifically to a method for calculating a comprehensive network risk assessment index. Background Technology

[0002] With the rapid development of internet technology, the internet has become a core platform for information dissemination and public communication. Massive amounts of information spread and disseminate at an extremely fast pace in cyberspace, affecting all aspects of social life, from politics and economics to culture and entertainment. Online public opinion, as the sum of attitudes, opinions, and emotions expressed by the public online regarding various events and topics, is increasingly influential. It not only reflects public opinion but also, to a certain extent, influences the development of events and policy-making.

[0003] Online public opinion monitoring aims to promptly grasp public views and attitudes towards specific events or topics through real-time collection, organization, and analysis of online information. While online public opinion monitoring can obtain a wealth of information about public opinion, simple information collection and preliminary analysis are often insufficient to accurately assess the degree of risk inherent in online public opinion. Therefore, comprehensive online risk assessment, as a crucial component of online public opinion monitoring, is undeniably important. It transforms fragmented public opinion data into a risk index with practical guiding significance by comprehensively considering and analyzing information from multiple dimensions, thereby providing a scientific basis for relevant decision-making. Comprehensive online risk assessment helps to understand the potential impact of online public opinion in a timely manner, formulate response strategies in advance, and effectively prevent and resolve crises.

[0004] While existing technologies for comprehensive online risk assessment have made some progress, numerous problems remain. Regarding data utilization, current technologies often focus only on explicit information, such as the number of trending searches and keyword frequency, while neglecting a large amount of implicit information. This implicit information is crucial for accurately assessing online risks; ignoring it may lead to biased or inaccurate assessment results. In terms of weight allocation, most existing methods adopt a fixed-weight model, assigning constant weight values ​​to indicators across different dimensions. However, online public opinion is dynamic and uncertain; the importance of indicators across different dimensions changes at different stages of different events. Fixed weights cannot adapt to these changes, easily leading to assessment results that are out of sync with reality. Furthermore, existing technologies also have limitations in information analysis capabilities, struggling to accurately understand the complex semantics, metaphors, and irony in online language, thus affecting the accurate judgment of public opinion risks.

[0005] In summary, existing comprehensive network risk assessment technologies have significant shortcomings in data utilization, weight allocation, and information analysis, and cannot meet the increasingly complex and ever-changing needs of network public opinion monitoring. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, the technical problem this invention aims to solve is: how to provide a method for calculating a comprehensive network risk assessment index, which utilizes the powerful language parsing capabilities of a large language model to deeply analyze raw network data, filter out key target information in eight dimensions, and design a dynamic weight generation mechanism for each dimension of target information, flexibly adjusting the weights according to the types of network hot events, thereby improving the accuracy and reliability of comprehensive network risk assessment.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0008] A method for calculating a comprehensive network risk assessment index, comprising:

[0009] S1: Obtain several pieces of network data related to the target network hotspot event;

[0010] S2: Classify each piece of network data using a large language model to obtain the predicted category of each piece of network data;

[0011] Among them, the prediction categories of network data include data on the influence of key opinion leaders (KOLs), data on netizens' sentiment, data on trending topics, data on events and topics, data on the level of participating media, data on the nature of participating media, data on key participating accounts, and / or data on platform participation.

[0012] S3: Calculate the corresponding risk indicator score for network data of each prediction category;

[0013] S4: By combining Markov decision process with Q-learning algorithm, dynamic weights are generated for each prediction category based on the event type of hot events in the target network;

[0014] S5: The risk index of network data for various prediction categories is weighted based on the dynamic weight of each prediction category to obtain the comprehensive network risk assessment index of the target network hotspot event.

[0015] Preferably, in step S2, the process of classifying the network data using a large language model includes:

[0016] S201: Obtain the network data to be classified, as well as preset examples of network data for each category;

[0017] S202: Encode the network data and each network data example separately to obtain the corresponding data embedding;

[0018] S203: Calculate the similarity between network data and each network data example based on data embeddings of network data and each network data example;

[0019] S204: Select the top k most similar network data examples for each category and construct a set of hint examples for the network data;

[0020] S205: Input the network data and its corresponding set of prompt examples into the large language model for classification to obtain the predicted category of the network data.

[0021] Preferably, in step S3, the formula for calculating the risk index score of the influencer data is as follows:

[0022] Risk index score of influencer data = ∑[initial influence × e] -γt [×Interaction Rate]×Fan Quality Coefficient;

[0023] in:

[0024] Interaction rate = Number of effective interactions per unit time / Total number of fans;

[0025] Fan quality index = percentage of active fans × percentage of verified fans;

[0026] In the formula: initial influence is the set historical influence benchmark value of the big V; γ represents the time decay constant; t represents the time difference between the current time and the last time the big V was active.

[0027] Preferably, in step S3, the formula for calculating the risk index score of netizens' sentiment data is expressed as follows:

[0028] The risk index score of netizens' sentiment data = |Negative sentiment percentage - Benchmark negative threshold| × sentiment intensity coefficient + positive sentiment percentage × 0.3;

[0029] in:

[0030] Emotional intensity coefficient = Frequency of extreme emotional words / Total number of words;

[0031] In the formula: the proportion of negative emotions is the proportion of text with negative emotions identified by the sentiment analysis model; the baseline negative threshold is the negative emotion warning value set based on industry consensus; and the proportion of positive emotions is the proportion of text with positive emotions identified by the sentiment analysis model.

[0032] Preferably, in step S3, the processing steps of the sentiment analysis model in the calculation of the risk index score of netizens' sentiment data include:

[0033] S301: Obtain the text of the online sentiment data to be identified;

[0034] S302: By encoding the text of netizens' sentiment data, a text embedding representation is obtained;

[0035] S303: Generate word syntax information for each sentence by using the syntactic dependencies within each sentence in the text of netizens' sentiment data;

[0036] S304: The text embedding representation and the word syntax information of all sentences are fused to obtain the fused text embedding;

[0037] S305: Sentiment classification is performed based on fused text embeddings using a classifier to obtain the sentiment classification prediction result.

[0038] Preferably, in step S302, the text embedding representation is obtained through the following steps:

[0039] S3021: Map each word of each sentence in the text of netizens' sentiment data to a vector representation;

[0040] S3022: Generate the hidden state of each word through a bidirectional gated loop unit to obtain the sentence vector sequence of each sentence;

[0041] S3023: Perform average pooling on all word representations in the statement vector sequence of each statement to obtain the overall statement representation of each statement;

[0042] S3024: Obtain the text embedding representation of the netizen sentiment data text.

[0043] Preferably, in step S303, word syntax information is obtained through the following steps:

[0044] S3031: Map each word in each sentence of the online sentiment data text to a low-dimensional dense vector to obtain the vector space embedding of each sentence;

[0045] S3032: Embed the vector space of each statement into the bidirectional LSTM model and output the corresponding statement context representation;

[0046] S3033: Input the sentence context representation into a graph convolutional neural network, combine it with a directed graph structure to propagate information between nodes to capture the sentence representation combined with the syntactic structure; after passing through an L-layer graph convolutional neural network, the sentence augmentation representation is obtained;

[0047] S3034: Max pooling is performed on the statement augmentation representation of each statement to obtain the corresponding word syntax information;

[0048] S3035: Obtain the set of word and syntax information for all sentences in the text of netizens' sentiment data.

[0049] Preferably, in step S3, the calculation formula for the risk indicator scores of trending search data, event topic data, participating media level data, participating media nature data, key participating account data, and platform participation data is expressed as follows:

[0050] 1) Risk index score of trending search data

[0051] The formula is expressed as:

[0052] Risk index score of trending search data = ∑(real-time popularity value of a single trending search × duration weight) / total number of trending searches × trending search type coefficient;

[0053] In the formula: the real-time popularity value of a single trending search represents the real-time popularity value displayed on the platform; the duration weight is the logarithmic transformation value of the ratio of the duration of the trending search to the baseline duration; the trending search type coefficient is a correction coefficient assigned based on the credibility of the platform from which the trending search originates.

[0054] 2) Risk indicator score of event topic data

[0055] The formula is expressed as:

[0056] Risk index score for event topic data = topic spread breadth index × topic diffusion speed index × topic homogenization coefficient;

[0057] in:

[0058] Topic reach index = number of unique users reached by the topic / base number of users;

[0059] In the formula: the topic diffusion speed index is the hierarchical growth of topic forwarding per unit time; the topic homogenization coefficient is the normalized value of the ratio of topic content duplication to originality.

[0060] 3) Risk indicator scores for participating in media-level data

[0061] The formula is expressed as:

[0062] Risk index score for participating in media-level data = ∑(media-level weight × media report length) / ∑ media report length × media-level correction coefficient;

[0063] In the formula: media level weight is the weight value assigned according to the administrative level of the media; media report length is the number of words in a single report; media level correction coefficient is the entropy value of the proportion of media reports of different levels;

[0064] 4) Risk index score for media-related data.

[0065] The formula is expressed as:

[0066] Risk index score for participating media nature data = official media share × 1.2 + commercial media share × 0.9 + self-media share × 0.7 - media nature diversity index;

[0067] in:

[0068] Official media share = Number of official media reports / Total number of reports;

[0069] In the formula: the media type diversity index is a media type richness index improved based on the Simpson index;

[0070] 5) Risk indicator scores of key participating account data

[0071] The formula is expressed as:

[0072] Risk indicator score for key participating accounts = ∑(Account type coefficient × Activity index × Content sensitivity index);

[0073] In the formula: Account type coefficient is the weight assigned according to the account authentication type; Activity index is the normalized value of the product of the number of posts and the number of interactions per unit time; Content sensitivity index is the density value of the content involving sensitive keywords;

[0074] 6) Risk indicator score of platform participation data

[0075] The formula is expressed as:

[0076] Risk indicator score for platform participation data = ∑(platform weight × platform user penetration rate) × platform type correction factor;

[0077] in:

[0078] Platform user penetration rate = User participation rate of the event on the platform / Total number of platform users;

[0079] In the formula: the platform weight is the weight value assigned based on the number of monthly active users on the platform; the platform type correction factor is the correction value assigned based on the characteristics of the platform.

[0080] Preferably, in step S4, the processing step of generating dynamic weights for each prediction category using a Markov decision process combined with a Q-learning algorithm includes:

[0081] S401: Define the state of a Markov decision process as: the event type of a network hotspot event;

[0082] S402: Define the action of a Markov decision process as: the dynamic weights assigned to each prediction category;

[0083] S403: Define the reward function of a Markov decision process as: minimizing the difference between the network risk comprehensive assessment index calculated based on the assigned dynamic weights and the true assessment index determined by experts;

[0084] S404: Select the next action with the highest Q value based on the current state; execute the next action and observe the immediate reward and the next state;

[0085] S405: Update the Q-value of the Q-learning algorithm based on the current action and current state, the next action and next state, and the immediate reward;

[0086] S406: Repeat steps S204 to S205 to iteratively update the Q value of the Q-learning algorithm until the Q function converges; generate the optimal policy using the converged Q function.

[0087] S407: Query the optimal strategy and select an action based on the current state of the target network hotspot event, and generate corresponding dynamic weights for each prediction category based on the selected action.

[0088] Preferably, in step S5, the formula for calculating the comprehensive network risk assessment index is as follows:

[0089] The comprehensive network risk assessment index is calculated as follows: a × risk index score of influence data of major KOLs + b × risk index score of netizen sentiment data + c × risk index score of trending search data + d × risk index score of event topic data + e × risk index score of participating media level data + f × risk index score of participating media nature data + g × risk index score of key participating account data + h × risk index score of platform participation data.

[0090] In the formula: a, b, c, d, e, f, g, and h are the weights of the corresponding risk indicator scores.

[0091] Compared with existing technologies, the network risk comprehensive assessment index calculation method in this invention has the following advantages:

[0092] This invention employs a large language model for classifying online data. Through deep semantic understanding technology, the large language model can automatically parse implicit semantics, sentiment, and relationships within online data, enabling rapid classification of batches of data. Compared to traditional keyword- or rule-based classification methods, the large language model avoids the limitations of manually pre-defined rules and can handle unstructured, multimodal online data, such as long texts and short video comments, significantly improving classification efficiency. Simultaneously, by performing contextual analysis and multi-dimensional feature extraction on online data, the large language model can more accurately identify the boundaries of data categories, avoiding misclassification problems caused by semantic ambiguity in traditional methods. This provides a reliable data foundation for subsequent risk indicator calculations, thereby improving the overall reliability of comprehensive online risk assessment.

[0093] This invention subdivides network data into eight categories and calculates risk indicator scores for each category. A comprehensive index is then obtained through dynamic weighting, offering significant advantages in accuracy and comprehensiveness. The eight categories cover the core dimensions of network risk: influencer data reflects the dissemination effect of opinion leaders; netizen sentiment data reflects public emotional tendencies; trending search data and event topic data capture the popularity of events and the diffusion path of issues; data on the level and nature of participating media reveals the credibility of information sources; and data on key participating accounts and platforms monitors the linkage effect between key nodes and platforms. This multi-dimensional classification system avoids the one-sidedness of single-indicator assessments, ensuring the comprehensiveness of risk assessment. Simultaneously, the independent calculation of risk indicator scores for each category, combined with dynamic weighting, quantifies the contribution of different dimensions of risk to the overall event, avoiding the weighting imbalance problem caused by fixed weights in traditional weighting methods. This makes the comprehensive assessment index more closely reflect the actual risk situation, thereby improving the credibility and accuracy of the comprehensive network risk assessment.

[0094] This invention generates dynamic weights using a Markov decision process combined with a Q-learning algorithm. Traditional weight allocation methods often employ fixed weights or manual experience, making it difficult to adapt to the varying risk characteristics of different event types. This invention, however, uses a Markov decision process to model the dynamic relationship between event types and weight allocation. Through state transitions and reward function design, the weight allocation strategy can automatically adjust according to changes in event types. Simultaneously, the Q-learning algorithm, through a trial-and-error learning mechanism, mines the correlation patterns of various risk categories under different event types from historical data, generating weight coefficients that better reflect actual risk distribution. This dynamic weight mechanism makes weight allocation more closely aligned with event type characteristics, improving the relevance and scientific rigor of risk assessment, avoiding assessment biases caused by rigid weights, and thus enhancing the practical application value of the network risk comprehensive assessment index. Attached Figure Description

[0095] To make the objectives, technical solutions, and advantages of the invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:

[0096] Figure 1 This is a logic diagram of the calculation method for the comprehensive network risk assessment index. Detailed Implementation

[0097] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but only to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0098] The following detailed explanation illustrates the specific implementation methods:

[0099] Example:

[0100] This embodiment discloses a method for calculating a comprehensive network risk assessment index.

[0101] A method for calculating a comprehensive network risk assessment index, comprising:

[0102] S1: Obtain several pieces of network data related to the target network hotspot event;

[0103] S2: Classify each piece of network data using a large language model to obtain the predicted category of each piece of network data;

[0104] Among them, the prediction categories of network data include data on the influence of key opinion leaders (KOLs), data on netizens' sentiment, data on trending topics, data on events and topics, data on the level of participating media, data on the nature of participating media, data on key participating accounts, and / or data on platform participation.

[0105] S3: Calculate the corresponding risk indicator score for network data of each prediction category;

[0106] S4: By combining Markov decision process with Q-learning algorithm, dynamic weights are generated for each prediction category based on the event type of hot events in the target network;

[0107] S5: The risk index of network data for various prediction categories is weighted based on the dynamic weight of each prediction category to obtain the comprehensive network risk assessment index of the target network hotspot event.

[0108] This invention employs a large language model for classifying online data. Through deep semantic understanding technology, the large language model can automatically parse implicit semantics, sentiment, and relationships within online data, enabling rapid classification of batches of data. Compared to traditional keyword- or rule-based classification methods, the large language model avoids the limitations of manually pre-defined rules and can handle unstructured, multimodal online data, such as long texts and short video comments, significantly improving classification efficiency. Simultaneously, by performing contextual analysis and multi-dimensional feature extraction on online data, the large language model can more accurately identify the boundaries of data categories, avoiding misclassification problems caused by semantic ambiguity in traditional methods. This provides a reliable data foundation for subsequent risk indicator calculations, thereby improving the overall reliability of comprehensive online risk assessment.

[0109] This invention subdivides network data into eight categories and calculates risk indicator scores for each category. A comprehensive index is then obtained through dynamic weighting, offering significant advantages in accuracy and comprehensiveness. The eight categories cover the core dimensions of network risk: influencer data reflects the dissemination effect of opinion leaders; netizen sentiment data reflects public emotional tendencies; trending search data and event topic data capture the popularity of events and the diffusion path of issues; data on the level and nature of participating media reveals the credibility of information sources; and data on key participating accounts and platforms monitors the linkage effect between key nodes and platforms. This multi-dimensional classification system avoids the one-sidedness of single-indicator assessments, ensuring the comprehensiveness of risk assessment. Simultaneously, the independent calculation of risk indicator scores for each category, combined with dynamic weighting, quantifies the contribution of different dimensions of risk to the overall event, avoiding the weighting imbalance problem caused by fixed weights in traditional weighting methods. This makes the comprehensive assessment index more closely reflect the actual risk situation, thereby improving the credibility and accuracy of the comprehensive network risk assessment.

[0110] This invention generates dynamic weights using a Markov decision process combined with a Q-learning algorithm. Traditional weight allocation methods often employ fixed weights or manual experience, making it difficult to adapt to the varying risk characteristics of different event types. This invention, however, uses a Markov decision process to model the dynamic relationship between event types and weight allocation. Through state transitions and reward function design, the weight allocation strategy can automatically adjust according to changes in event types. Simultaneously, the Q-learning algorithm, through a trial-and-error learning mechanism, mines the correlation patterns of various risk categories under different event types from historical data, generating weight coefficients that better reflect actual risk distribution. This dynamic weight mechanism makes weight allocation more closely aligned with event type characteristics, improving the relevance and scientific rigor of risk assessment, avoiding assessment biases caused by rigid weights, and thus enhancing the practical application value of the network risk comprehensive assessment index.

[0111] To better illustrate the technical solution of the present invention, this embodiment is described in the following parts.

[0112] I. Classification of Large Language Models

[0113] In the specific implementation process, the steps for classifying various network data using a large language model include:

[0114] S201: Obtain the network data to be classified, as well as preset examples of network data for each category;

[0115] S202: Encode the network data and each network data example separately to obtain the corresponding data embedding;

[0116] S203: Calculate the similarity between network data and each network data example based on data embeddings of network data and each network data example;

[0117] The formula is expressed as:

[0118] sim(W,S)=Cosine(V W V S );

[0119] In the formula: sim(W,S) represents the similarity between network data W and network data example S; Cosine(·) represents the calculation of cosine similarity; V W V S These represent the data embeddings of network data file W and network data example S, respectively.

[0120] S204: Select the top k most similar network data examples for each category and construct a set of hint examples for the network data;

[0121] S205: Input the network data and its corresponding set of prompt examples into the large language model for classification to obtain the predicted category of the network data.

[0122] Traditional large-scale model classification requires independent processing of the entire dataset. In contrast, this invention pre-selects network data examples most relevant to the data to be classified by similarity calculation to construct a hint set. This reduces the length of context that the large language model needs to process simultaneously, thus lowering computational complexity. At the same time, the hint example set serves as a classification guide for the large language model, allowing the model to make rapid decisions based on pre-selected similar examples instead of learning category features from scratch. This shortens the inference path, increases the classification throughput per unit time, and thereby improves the efficiency of large language models in classifying network data. Meanwhile, the set of suggested examples strengthens the model's classification decision through the semantic anchoring effect. Web data often has semantic ambiguity issues such as polysemy and implicit expressions, while the pre-set examples clarify the semantic boundaries of categories through specific cases, enabling large language models to more accurately identify subtle semantic differences by comparing the similarity between the data to be classified and the examples. At the same time, dynamically selecting the most similar examples to construct the suggested example set avoids problems such as outdated examples or domain mismatch that may exist in a fixed example library, ensuring that the suggested content is always highly relevant to the data to be classified. This allows the model to refer to typical cases in the domain while focusing on the semantic features of the data itself when classifying, reducing misclassification caused by domain bias or example quality, thereby improving the reliability of the classification results.

[0123] II. Risk Indicator Scores

[0124] In the specific implementation process, the risk indicator scores for data such as influencer data, online sentiment data, trending search data, event topic data, media level data, media nature data, key participating account data, and platform participation data are assessed as follows:

[0125] 1) Risk indicator score of influencer data

[0126] The formula for calculating the risk index score of influential figures' influence data is as follows:

[0127] Risk index score of influencer data = ∑[initial influence × e] -γt [×Interaction Rate]×Fan Quality Coefficient;

[0128] in:

[0129] Interaction rate = Number of effective interactions per unit time / Total number of fans;

[0130] Fan quality index = percentage of active fans × percentage of verified fans;

[0131] In the formula: the initial influence is the set historical influence benchmark value of the big V; γ represents the time decay constant (0.05-0.3 days / unit time in this invention); t represents the time difference between the current time and the last active time of the big V.

[0132] This invention introduces e-γt The exponential decay factor simulates the decay pattern of influencer influence over time and, combined with the fan quality coefficient, achieves accurate assessment. Compared with the traditional static influence calculation, it is more in line with the changing patterns of the online environment, thereby improving the accuracy of risk indicator score assessment of influencer influence data.

[0133] 2) Risk indicator score of netizens' sentiment data

[0134] The formula for calculating the risk index score of online sentiment data is as follows:

[0135] The risk index score of netizens' sentiment data = |Negative sentiment percentage - Benchmark negative threshold| × sentiment intensity coefficient + positive sentiment percentage × 0.3;

[0136] in:

[0137] Emotional intensity coefficient = Frequency of extreme emotional words / Total number of words;

[0138] In the formula: the proportion of negative emotions is the proportion of negative emotional text identified by the sentiment analysis model; the benchmark negative threshold is the negative emotion warning value set based on industry knowledge (this invention sets it to 30%); the proportion of positive emotions is the proportion of positive emotional text identified by the sentiment analysis model.

[0139] This invention captures the risk of extreme emotions by introducing an emotion intensity coefficient, quantifies abnormal emotional fluctuations by using a benchmark threshold deviation value, and forms a dynamic balance by combining a positive emotion offsetting mechanism, thereby improving the accuracy of risk indicator score assessment of netizens' emotional data.

[0140] 3) Risk index score of trending search data

[0141] The formula is expressed as:

[0142] Risk index score of trending search data = ∑(real-time popularity value of a single trending search × duration weight) / total number of trending searches × trending search type coefficient;

[0143] In the formula: the real-time popularity value of a single hot search represents the real-time popularity value displayed on the platform; the duration weight is the logarithmic transformation value of the ratio of the duration of the hot search to the baseline duration; the hot search type coefficient is a correction coefficient assigned based on the credibility of the hot search source platform (e.g., 1.2 for official platforms, 0.9 for other commercial platforms).

[0144] This invention avoids numerical distortion caused by linear superposition by introducing a duration weight of logarithmic transformation, and combines it with the hot search type coefficient to reflect platform difference risks, thereby improving the accuracy of risk indicator score assessment for hot search data.

[0145] 4) Risk indicator score of event topic data

[0146] The formula is expressed as:

[0147] Risk index score for event topic data = topic spread breadth index × topic diffusion speed index × topic homogenization coefficient;

[0148] in:

[0149] Topic reach index = number of unique users reached by the topic / base number of users;

[0150] In the formula: the topic diffusion speed index is the hierarchical growth of topic forwarding per unit time; the topic homogenization coefficient is the normalized value of the ratio of topic content duplication to originality.

[0151] This invention uses a three-dimensional product model to replace the traditional weighted sum, highlighting the three-dimensional risk characteristics of topic dissemination, and reflects the information cocoon effect through a homogenization coefficient, thereby improving the accuracy of risk indicator score assessment for event topic data.

[0152] 5) Risk indicator scores for participating in media-level data

[0153] The formula is expressed as:

[0154] Risk index score for participating in media-level data = ∑(media-level weight × media report length) / ∑ media report length × media-level correction coefficient;

[0155] In the formula: the media level weight is the weight value assigned according to the administrative level of the media (in this invention: the weight value of central-level media is 1.5, provincial-level media is 1.2, and municipal-level media is 1.0); the media report length is the number of words in a single report; the media level correction coefficient is the entropy value of the proportion of media reports of different levels;

[0156] This invention uses an entropy correction coefficient to reflect the concentration risk of media reports, avoiding excessive dominance of assessment results by a single level of media, thereby improving the accuracy of risk indicator score assessment for media-level data.

[0157] 6) Participation in media-related data risk indicator scores

[0158] The formula is expressed as:

[0159] Risk index score for participating media nature data = official media share × 1.2 + commercial media share × 0.9 + self-media share × 0.7 - media nature diversity index;

[0160] in:

[0161] Official media share = Number of official media reports / Total number of reports;

[0162] In the formula: the media type diversity index is a media type richness index improved based on the Simpson index;

[0163] This invention employs a linear combination of diversity indices to reflect the risk differences among media of different natures, while also suppressing the overestimation of risk caused by the accumulation of homogeneous media through the diversity index, thereby improving the accuracy of risk indicator scores for participating media nature data.

[0164] 7) Risk indicator scores of key participating account data

[0165] The formula is expressed as:

[0166] Risk indicator score for key participating accounts = ∑(Account type coefficient × Activity index × Content sensitivity index);

[0167] In the formula: the account type coefficient is the weight assigned according to the account authentication type (in this invention: the weight of government accounts is 1.3, enterprise accounts is 1.1, and personal accounts is 0.9); the activity index is the normalized value of the product of the number of posts and the number of interactions per unit time; the content sensitivity index is the density value of the content involving sensitive keywords.

[0168] This invention employs a three-dimensional product model to comprehensively assess account risk. The content sensitivity index quantifies risk gradients by using keyword density rather than simply judging their presence or absence, thereby improving the accuracy of risk indicator score assessment for key participating accounts.

[0169] 8) Risk indicator score of platform participation data

[0170] The formula is expressed as:

[0171] Risk indicator score for platform participation data = ∑(platform weight × platform user penetration rate) × platform type correction factor;

[0172] in:

[0173] Platform user penetration rate = User participation rate of the event on the platform / Total number of platform users;

[0174] In the formula: the platform weight is the weight value assigned based on the number of monthly active users of the platform; the platform type correction factor is the correction value assigned based on the characteristics of the platform (in this invention: the correction value for social platforms is 1.2, for news platforms it is 1.0, and for video platforms it is 0.8);

[0175] This invention introduces a platform type correction factor to reflect the risk differences in the propagation characteristics of different platforms, and combines it with user penetration rate to reflect the real penetration risk of events within the platform, thereby improving the accuracy of risk indicator score assessment of platform participation data.

[0176] III. Sentiment Analysis Model

[0177] In the process of calculating the risk index score of netizens' sentiment data in this invention, the processing steps of the sentiment analysis model include:

[0178] S301: Obtain the text of the online sentiment data to be identified;

[0179] S302: By encoding the text of netizens' sentiment data, a text embedding representation is obtained;

[0180] S303: Generate word syntax information for each sentence by using the syntactic dependencies within each sentence in the text of netizens' sentiment data;

[0181] S304: The text embedding representation and the word syntax information of all sentences are fused to obtain the fused text embedding;

[0182] The formula is expressed as:

[0183] H = [E||D];

[0184] In the formula: H represents fused text embedding; E represents text embedding representation; D represents the set of word syntax information for all statements;

[0185] S305: Sentiment classification is performed based on fused text embedding using a classifier to obtain the sentiment classification prediction result;

[0186] The formula is expressed as:

[0187] y = Softmax(H);

[0188] In the formula: y represents the sentiment classification prediction result; H represents the fused text embedding.

[0189] This invention improves the accuracy of sentiment recognition in three dimensions by extracting and fusing text embeddings (coarse-grained information) and syntactic information (fine-grained information): 1) Semantic-structural complementarity: Text embeddings excel at capturing overall semantic tendencies, while syntactic information accurately analyzes local structural features. The fusion of these two approaches avoids misjudgments caused by pure semantic models neglecting structural features such as negation words and degree adverbs; 2) Fine-grained sentiment element extraction: Syntactic dependency analysis can locate the dependency relationship between sentiment words and evaluation objects. Combined with the semantic strength of text embeddings, it enables accurate matching between evaluation objects and sentiment tendencies, solving the problem of ambiguous sentiment objects in traditional methods; 3) Enhanced robustness: In complex sentences, syntactic structures provide clear semantic boundaries. Combined with the contextual understanding of text embeddings, it avoids misjudging neutral expressions as negative sentiments. This dual-granularity fusion makes the model more generalizable in multimodal and cross-domain sentiment recognition, aligning with the evolution of sentiment analysis from coarse-grained tendency judgment to fine-grained element extraction. This invention improves the accuracy of emotion recognition in three aspects—semantic understanding depth, structural parsing accuracy, and emotional element localization—by complementing coarse and fine granular information, which aligns with the current cutting-edge technology of "semantic-structural dual-drive" in the field of emotion analysis.

[0190] Specifically, the text embedding representation is obtained through the following steps:

[0191] S3021: Map each word of each sentence in the text of netizens' sentiment data to a vector representation;

[0192] S3022: Generate the hidden state h of each word through a bidirectional gated loop unit. i,j This yields the statement vector sequence {h} for each statement. i,1 ,…,h i,n};

[0193] The formula is expressed as:

[0194]

[0195] Where: φ emb (·) indicates an embedded function; || indicates a concatenation operation; and These represent the j-th word w in the i-th sentence. i,j Forward and backward representations;

[0196] S3023: The statement vector sequence {h} for each statement i,1 ,…,h i,n All words in the expression are represented by average pooling to obtain the overall expression e of each statement. i ;

[0197] S3024: Obtain the text embedding representation of the netizen sentiment data text E={e1,…,e N}

[0198] Specifically, word syntax information is obtained through the following steps:

[0199] S3031: Map each word in each sentence of the online sentiment data text to a low-dimensional dense vector to obtain the vector space embedding of each sentence;

[0200] S3032: Embed the vector space of each statement into the bidirectional LSTM model and output the corresponding statement context representation;

[0201] The formula is expressed as:

[0202] Vector space embedding

[0203] Statement context representation

[0204] in,

[0205]

[0206] In the formula: Indicates learnable parameters; Vector space data embedding representing statements; The statement representing the bidirectional LSTM model;

[0207] S3033: The Spacy tool is used to extract the syntactic structure of each statement to form a corresponding directed graph structure; then the statement context representation is input into the graph convolutional neural network, and information propagation between nodes is performed in combination with the directed graph structure to capture the statement representation combined with the syntactic structure; after passing through the L layers of graph convolutional neural network, the statement augmentation representation is obtained.

[0208] The formula is expressed as:

[0209] Statement enhancement representation

[0210] in,

[0211] In the formula: All represent learnable parameters; l = [1,2,…,L], l∈L represents the l-th layer graphical convolutional neural network (GCN); σ represents the nonlinear activation function ReLU; Represents the structure of a directed graph; This represents the representation of the j-th statement in the i-th session within the l-1 level of a directed graph structure.

[0212] S3034: Enhanced statement representation for each statement Perform max pooling to obtain the corresponding word syntactic information d. i ;

[0213] S3035: Obtain the set D = {d1, ..., d...} of word and syntax information for all sentences in the online sentiment data text. N}

[0214] IV. Dynamic Weights

[0215] In the specific implementation process, the steps for generating dynamic weights for each prediction category using Markov decision processes combined with Q-learning algorithms include:

[0216] S401: Define the state of a Markov decision process as: the event type of a network hotspot event;

[0217] S402: Define the action of a Markov decision process as: the dynamic weights assigned to each prediction category;

[0218] S403: Define the reward function of a Markov decision process as: minimizing the difference between the network risk comprehensive assessment index calculated based on the assigned dynamic weights and the true assessment index determined by experts;

[0219] S404: Select the next action with the highest Q value based on the current state; execute the next action and observe the immediate reward and the next state;

[0220] S405: Update the Q-value of the Q-learning algorithm based on the current action and current state, the next action and next state, and the immediate reward;

[0221] S406: Repeat steps S204 to S205 to iteratively update the Q value of the Q-learning algorithm until the Q function converges; generate the optimal policy using the converged Q function.

[0222] S407: Based on the current state (event type) of the target network hotspot event, query the optimal strategy, select the action, and generate corresponding dynamic weights for each prediction category based on the selected action.

[0223] Traditional fixed-weight methods struggle to adapt to the varying risk characteristics of different event types. This solution, however, achieves precise weight matching between event types through a dynamic mapping of state (event type) and action (weight allocation). For example, in public health emergencies, the algorithm automatically increases the weight of "media participation data" and "netizen sentiment data," while decreasing the weight of "trending search data." In routine social public opinion events, it emphasizes the weight of "event topic data" and "platform participation data." This dynamic adjustment mechanism avoids the biases caused by fixed weights, making the weight allocation more aligned with the actual risk distribution. Simultaneously, Q-learning, through iterative learning of weight-effect correlation patterns in historical data, can uncover the contribution of each risk category to the comprehensive assessment under different event types, forming an experience-driven weight allocation strategy. The reward function directly correlates the deviation between the assessment results and the true values. By minimizing the deviation, the optimization objective drives the weight allocation strategy towards greater accuracy. This dynamic weight generation possesses both adaptive capabilities and the potential for continuous optimization, fundamentally improving the rationality and scientific nature of weight allocation and enhancing the credibility and decision-making reference value of the comprehensive network risk assessment index.

[0224] V. Comprehensive Network Risk Assessment Index

[0225] In practice, the formula for calculating the comprehensive network risk assessment index is as follows:

[0226] The comprehensive network risk assessment index is calculated as follows: a × risk index score of influence data of major KOLs + b × risk index score of netizen sentiment data + c × risk index score of trending search data + d × risk index score of event topic data + e × risk index score of participating media level data + f × risk index score of participating media nature data + g × risk index score of key participating account data + h × risk index score of platform participation data.

[0227] In the formula: a, b, c, d, e, f, g, and h are the weights of the corresponding risk indicator scores.

[0228] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Those skilled in the art should understand that any modifications or equivalent substitutions to the technical solutions of the present invention without departing from the spirit and scope of the present invention should be covered within the scope of the claims of the present invention.

Claims

1. A method for calculating a comprehensive network risk assessment index, characterized in that, include: S1: Obtain several pieces of network data related to the target network hotspot event; S2: Classify each piece of network data using a large language model to obtain the predicted category of each piece of network data; Among them, the prediction categories of network data include data on the influence of key opinion leaders (KOLs), data on netizens' sentiment, data on trending topics, data on events and topics, data on the level of participating media, data on the nature of participating media, data on key participating accounts, and / or data on platform participation. S3: Calculate the corresponding risk indicator score for network data of each prediction category; S4: By combining Markov decision process with Q-learning algorithm, dynamic weights are generated for each prediction category based on the event type of hot events in the target network; S5: The risk index of network data for various prediction categories is weighted based on the dynamic weight of each prediction category to obtain the comprehensive network risk assessment index of the target network hotspot event.

2. The method for calculating the comprehensive network risk assessment index as described in claim 1, characterized in that: Step S2, the process of classifying the network data using a large language model, includes: S201: Obtain the network data to be classified, as well as preset examples of network data for each category; S202: Encode the network data and each network data example separately to obtain the corresponding data embedding; S203: Calculate the similarity between network data and each network data example based on data embeddings of network data and each network data example; S204: Select the top k most similar network data examples for each category and construct a set of hint examples for the network data; S205: Input the network data and its corresponding set of prompt examples into the large language model for classification to obtain the predicted category of the network data.

3. The method for calculating the comprehensive network risk assessment index as described in claim 1, characterized in that: In step S3, the formula for calculating the risk index score of the influencer data is as follows: Risk index score of influencer data = ∑[initial influence × e] -γt [×Interaction Rate]×Fan Quality Coefficient; in: Interaction rate = Number of effective interactions per unit time / Total number of fans; Fan quality index = percentage of active fans × percentage of verified fans; In the formula: initial influence is the set historical influence benchmark value of the big V; γ represents the time decay constant; t represents the time difference between the current time and the last time the big V was active.

4. The method for calculating the comprehensive network risk assessment index as described in claim 1, characterized in that: In step S3, the formula for calculating the risk index score of netizens' sentiment data is as follows: The risk index score of netizens' sentiment data = |Negative sentiment percentage - Benchmark negative threshold| × sentiment intensity coefficient + positive sentiment percentage × 0.3; in: Emotional intensity coefficient = Frequency of extreme emotional words / Total number of words; In the formula: the proportion of negative emotions is the proportion of text with negative emotions identified by the sentiment analysis model; the baseline negative threshold is the negative emotion warning value set based on industry consensus; and the proportion of positive emotions is the proportion of text with positive emotions identified by the sentiment analysis model.

5. The method for calculating the comprehensive network risk assessment index as described in claim 4, characterized in that: In step S3, during the calculation of the risk index score for netizens' sentiment data, the processing steps of the sentiment analysis model include: S301: Obtain the text of the online sentiment data to be identified; S302: By encoding the text of netizens' sentiment data, a text embedding representation is obtained; S303: Generate word syntax information for each sentence by using the syntactic dependencies within each sentence in the text of netizens' sentiment data; S304: The text embedding representation and the word syntax information of all sentences are fused to obtain the fused text embedding; S305: Sentiment classification is performed based on fused text embeddings using a classifier to obtain the sentiment classification prediction result.

6. The method for calculating the comprehensive network risk assessment index as described in claim 5, characterized in that: In step S302, the text embedding representation is obtained through the following steps: S3021: Map each word of each sentence in the text of netizens' sentiment data to a vector representation; S3022: Generate the hidden state of each word through a bidirectional gated loop unit to obtain the sentence vector sequence of each sentence; S3023: Perform average pooling on all word representations in the statement vector sequence of each statement to obtain the overall statement representation of each statement; S3024: Obtain the text embedding representation of the netizen sentiment data text.

7. The method for calculating the comprehensive network risk assessment index as described in claim 5, characterized in that: In step S303, word syntax information is obtained through the following steps: S3031: Map each word in each sentence of the online sentiment data text to a low-dimensional dense vector to obtain the vector space embedding of each sentence; S3032: Embed the vector space of each statement into the bidirectional LSTM model and output the corresponding statement context representation; S3033: Input the sentence context representation into a graph convolutional neural network, and combine it with a directed graph structure to perform information propagation between nodes in order to capture the sentence representation combined with the syntactic structure; After passing through an L-layer graph convolutional neural network, an enhanced representation of the statement is obtained. S3034: Max pooling is performed on the statement augmentation representation of each statement to obtain the corresponding word syntax information; S3035: Obtain the set of word and syntax information for all sentences in the text of netizens' sentiment data.

8. The method for calculating the comprehensive network risk assessment index as described in claim 1, characterized in that: In step S3, the calculation formula for the risk indicator scores of trending search data, event topic data, participating media level data, participating media nature data, key participating account data, and platform participation data is expressed as follows: 1) Risk index score of trending search data The formula is expressed as: Risk index score of trending search data = ∑(real-time popularity value of a single trending search × duration weight) / total number of trending searches × trending search type coefficient; In the formula: the real-time popularity value of a single trending search represents the real-time popularity value displayed on the platform; the duration weight is the logarithmic transformation value of the ratio of the duration of the trending search to the baseline duration; the trending search type coefficient is a correction coefficient assigned based on the credibility of the platform from which the trending search originates. 2) Risk indicator score of event topic data The formula is expressed as: Risk index score for event topic data = topic spread breadth index × topic diffusion speed index × topic homogenization coefficient; in: Topic reach index = number of unique users reached by the topic / base number of users; In the formula: the topic diffusion speed index is the hierarchical growth of topic forwarding per unit time; the topic homogenization coefficient is the normalized value of the ratio of topic content duplication to originality. 3) Risk indicator scores for participating in media-level data The formula is expressed as: Risk index score for participating in media-level data = ∑(media-level weight × media report length) / ∑ media report length × media-level correction coefficient; In the formula: media level weight is the weight value assigned according to the administrative level of the media; media report length is the number of words in a single report; media level correction coefficient is the entropy value of the proportion of media reports of different levels; 4) Risk index score for media-related data. The formula is expressed as: Risk index score for participating media nature data = official media share × 1.2 + commercial media share × 0.9 + self-media share × 0.7 - media nature diversity index; in: Official media share = Number of official media reports / Total number of reports; In the formula: the media type diversity index is a media type richness index improved based on the Simpson index; 5) Risk indicator scores of key participating account data The formula is expressed as: Risk indicator score for key participating accounts = ∑(Account type coefficient × Activity index × Content sensitivity index); In the formula: Account type coefficient is the weight assigned according to the account authentication type; Activity index is the normalized value of the product of the number of posts and the number of interactions per unit time; Content sensitivity index is the density value of the content involving sensitive keywords; 6) Risk indicator score of platform participation data The formula is expressed as: Risk indicator score for platform participation data = ∑(platform weight × platform user penetration rate) × platform type correction factor; in: Platform user penetration rate = User participation rate of the event on the platform / Total number of platform users; In the formula: the platform weight is the weight value assigned based on the number of monthly active users on the platform; the platform type correction factor is the correction value assigned based on the characteristics of the platform.

9. The method for calculating the comprehensive network risk assessment index as described in claim 1, characterized in that: Step S4, which involves generating dynamic weights for each predicted category using a Markov decision process combined with a Q-learning algorithm, includes the following steps: S401: Define the state of a Markov decision process as: the event type of a network hotspot event; S402: Define the action of a Markov decision process as: the dynamic weights assigned to each prediction category; S403: Define the reward function of a Markov decision process as: minimizing the difference between the network risk comprehensive assessment index calculated based on the assigned dynamic weights and the true assessment index determined by experts; S404: Select the next action with the highest Q value based on the current state; execute the next action and observe the immediate reward and the next state; S405: Update the Q-value of the Q-learning algorithm based on the current action and current state, the next action and next state, and the immediate reward; S406: Repeat steps S204 to S205 to iteratively update the Q value of the Q-learning algorithm until the Q function converges; generate the optimal policy using the converged Q function. S407: Query the optimal strategy and select an action based on the current state of the target network hotspot event, and generate corresponding dynamic weights for each prediction category based on the selected action.

10. The method for calculating the comprehensive network risk assessment index as described in claim 1, characterized in that: In step S5, the formula for calculating the comprehensive network risk assessment index is as follows: The comprehensive network risk assessment index is calculated as follows: a × risk index score of influence data of major KOLs + b × risk index score of netizen sentiment data + c × risk index score of trending search data + d × risk index score of event topic data + e × risk index score of participating media level data + f × risk index score of participating media nature data + g × risk index score of key participating account data + h × risk index score of platform participation data. In the formula: a, b, c, d, e, f, g, and h are the weights of the corresponding risk indicator scores.