Network information security protection method based on big data

Through social media emotion mining and user behavior analysis, combined with multi-source data fusion model, the problem of inaccurate network threat identification in traditional methods is solved, and comprehensive, precise protection and timely response to network information security is achieved.

CN120263442AInactive Publication Date: 2025-07-04SHANDONG WEIPING INFORMATION SECURITY EVALUATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510251547.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional network information security protection methods are difficult to timely and effectively identify new types of network attacks and online rumors, ignoring user emotional and behavioral analysis, resulting in inaccurate identification of network threats.

Method used

Through social media emotion mining, user behavior and emotional association analysis and multi-source data fusion, user behavior and emotional state are monitored in real time, and a multi-source data fusion model is built to conduct threat warning. Combined with natural language processing technology and exclusive emotional dictionary, negative emotions and abnormal behaviors are identified.

Benefits of technology

It realizes accurate identification and timely protection of network information security threats, improves the accuracy of threat identification and comprehensive warning, and provides a scientific basis for security protection measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263442A_ABST
    Figure CN120263442A_ABST
Patent Text Reader

Abstract

The invention discloses a network information security protection method based on big data, and relates to the technical field of network information security, and the protection method comprises the following specific steps: S100, data collection: through a web crawler technology and a data interface, carrying out data collection; network data of texts, user behaviors and network traffic are collected in real time from a social media platform, a network service provider and a user terminal, and accurate recognition and early warning of network information security threats are achieved through social media emotion mining and user behavior emotion association analysis. By applying a natural language processing technology and constructing an exclusive emotion dictionary, potential threat speech which contains negative emotions and is related to network security can be accurately recognized, meanwhile, user behavior data are combined, a user behavior emotion association model is established, user behaviors are monitored in real time, the threat recognition accuracy is improved, and the threat recognition efficiency is improved. And powerful support is provided for taking safety protection measures in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network information security technology, and in particular to a network information security protection method based on big data. Background Art

[0002] In terms of technical background, with the rapid development of the Internet, cyberspace has become an important platform for information dissemination, social interaction and commercial transactions. However, the openness of cyberspace has also brought many security challenges, especially social media platforms. Due to their large user base and fast information dissemination speed, they have become the main channel for network attacks and online rumors. Therefore, how to effectively monitor and warn of network information security threats and protect user privacy and data security has become a technical problem that needs to be solved urgently.

[0003] Traditional technologies have many shortcomings in network information security protection. On the one hand, traditional methods mainly rely on technical means such as firewalls and intrusion detection systems. These means are often powerless when facing new network attacks and it is difficult to provide timely and effective protection. On the other hand, traditional methods ignore the in-depth analysis of user emotions and behaviors and cannot accurately identify potential network threats. For example, some online rumors or malicious information may be spread by disguising themselves as content posted by normal users, and it is difficult for traditional methods to effectively identify and intercept them.

[0004] Therefore, the development of network information security protection methods based on big data not only improves the accuracy and timeliness of network information security protection, but also provides users with a safer and more reliable network environment. Summary of the invention

[0005] The purpose of the present invention is to make up for the shortcomings of the existing technology and provide a network information security protection method based on big data. The method integrates social media sentiment mining, user behavior sentiment correlation analysis, and multi-source data fusion and threat warning technical means to achieve comprehensive and accurate protection against network information security threats. By real-time monitoring and analysis of user behavior and emotional state, the present invention can timely discover potential network security threats and take effective security protection measures, thereby greatly reducing network security risks.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a network information security protection method based on big data, the specific steps of the protection method are:

[0007] S100, data collection: Through web crawler technology and data interfaces, data on text, user behavior, and network traffic are collected from social media platforms, network service providers, and user terminals in real time, and the collected data is cleaned, denoised, and formatted;

[0008] S200, Social Media Sentiment Mining: Use natural language processing technology to segment, part-of-speech tag, and name entity recognize social media text data. Construct an exclusive dictionary according to the initial sentiment dictionary through the sentiment dictionary update formula. Calculate the sentiment score of the text data according to the exclusive dictionary through the sentiment analysis formula to determine whether the text contains negative emotions related to network security, mark potential threat remarks containing negative emotions and related to network security, and mine the characteristics of the theme, source, and spread range of threat information through clustering and association analysis;

[0009] S300, User Behavior Sentiment Association Analysis: Collect user behavior data such as logins, operations, and data accesses. Combine the results of social media sentiment mining to establish a user behavior sentiment association model, judge the degree of user behavior sentiment association, determine the sentiment baseline under different user behavior patterns, monitor user behavior in real time, calculate the deviation degree from the sentiment state baseline, mark abnormal users according to the deviation degree, and analyze the network security threats faced;

[0010] S400, Multi-source Data Fusion and Threat Warning: Integrate the results of social media sentiment mining and user behavior sentiment association analysis, construct a multi-source data fusion model to calculate threat indicators, set different levels of threat warning thresholds according to the threat indicators, and trigger a warning when the integrated threat indicator exceeds the threshold, sending a warning content containing information such as the threat type, source, and scope of influence to the administrator or user;

[0011] S500, Security Protection Measures: Take different measures according to the warning level: When there is a low-level warning, strengthen monitoring and the collection and analysis of user feedback; When there is a medium-level warning, conduct a comprehensive inspection of the server and optimize it, and at the same time conduct account security inspections and protection; When there is a high-level warning, immediately initiate an emergency response, perform system isolation and a comprehensive inspection for evidence collection.

[0012] Further, in the S200, social media sentiment mining, an exclusive dictionary is constructed according to the initial sentiment dictionary using the sentiment dictionary update formula. Let the initial sentiment dictionary be E0, and the sentiment words and corresponding sentiment scores included are represented as (w0, s0) ∈ E0, where w0 is the sentiment word and s0 is the sentiment score. According to the newly emerged text data D text Update the sentiment dictionary. For each newly emerged sentiment word w′ in D text Use the formula to update its sentiment score. The formula is:

[0013]

[0014] where, TF-IDF(w′, d j ) is the TF-IDF value of the word w′ in the text d j , s 0j (d j ) is the text d calculated according to the contextj and update the sentiment dictionary.

[0015] Furthermore, in S200, in social media sentiment mining, calculate the sentiment score of text data through the sentiment analysis formula, and determine whether the text contains negative emotions related to network security. For text data d, tokenize it into w1, w2, …, w k , and use the updated sentiment dictionary E to calculate its sentiment score S(d). The formula is:

[0016]

[0017] where s(w i ) is the sentiment score of sentiment word w i , e represents the natural constant, and g(w i , d) is an importance function of a word w i in text d. The formula is:

[0018]

[0019] POS(w i ) is the part-of-speech weight of word w i , α1 is the part-of-speech weight adjustment factor, dist(w i , d) is the relative position of word w i in text d, β1 is the position influence factor. Set the sentiment score threshold θ. If the calculated text sentiment score S(d) < θ, mark the text as potential threat remarks containing negative emotions related to network security.

[0020] Furthermore, in S200, in social media sentiment mining, perform clustering and association analysis on the text data marked as potential threats through the potential threat information clustering and association analysis formula, and mine the characteristics of the theme, source, and dissemination scope of threat information. Assume there is text data D threat marked as potential threats, and represent it as a vector set Use the spectral clustering algorithm to perform clustering on it. The calculation formula is:

[0021]

[0022] where is the Euclidean distance between vectors, sim(w ik , w jk ) is the semantic similarity of the k-th keyword in texts i and j, λ 1k is the weight of keyword similarity, γ is the distance influence factor, and q is the number of keywords considered.

[0023] Furthermore, in step S300, for the construction of the user behavior-emotion association model in user behavior-emotion association analysis, let the user behavior feature vector be The emotion score sequence is S = [s1, s2, …, s m . The calculation formula of the user behavior-emotion association model is:

[0024]

[0025] where α2 and β2 are weight coefficients, and α2 + β2 = 1, w i is the weight of the behavior feature b i , and r is the weight of the emotion score s j .

[0026] Furthermore, in step S300, for the determination of the emotion state baseline in user behavior-emotion association analysis, let the emotion score sequence S i of user U obtained from social media emotion mining within the time period t be i , t = {s i,t1 , s i,t2 , …, s i,tn}, where s i,tj is the emotion score of user U i at the time point tj. At the same time, integrate the login frequency F login , operation activity index A operate , and data access anomaly degree E data in the behavior data of the user within the same time period t into the behavior feature vector

[0027] For each user U i , use the data of m time periods in the past historical data to determine the emotion state baseline, calculate the mean and standard deviation of each behavior feature to establish the baseline model. For the login frequency F login , its mean and standard deviation The formulas are:

[0028] Similarly, calculate the mean and standard deviation of other behavior features, so as to obtain the emotion state baseline i of user U and

[0029] Furthermore, in step S300, for the detection and marking of abnormal users in user behavior-emotion association analysis, let the behavior feature vector i of user U at the current time t current be

[0030]

[0031] Calculate the deviation degree D from the emotional state baseline total , and the formula is:

[0032]

[0033] where is the k-th behavioral eigenvalue of user U i at the current time t current , μ ik and σ ik are the corresponding baseline mean and standard deviation, and A k is the deviation weight of the k-th behavioral feature. Obtain the latest emotional score of the user on the social media in real time When the comprehensive deviation degree D total exceeds the set threshold θ total , and the latest emotional score is negative, mark user U i as an abnormal user.

[0034] Furthermore, in the S400, the construction of the multi-source data fusion model in multi-source data fusion and threat warning, let the social media emotion mining result be S(d), and the user behavior-emotion correlation analysis result be D total , and the calculation formula of the threat index F is:

[0035] F = ω1×(1 + λ2×S(d))×D total + ω2×(1 + λ3×D total )×S(d), where

[0036] ω1 and ω2 are weight coefficients, and λ2 and λ3 are adjustment factors.

[0037] Furthermore, in the S400, the division of the warning level in multi-source data fusion and threat warning, let the low threat threshold be the medium threat threshold be and the high threat threshold be Th high ;

[0038] When no warning is triggered;

[0039] When a low-level warning is triggered;

[0040] When Th medium ≤ F < Th high , a medium-level warning is triggered;

[0041] When F ≥ Th high , a high-level warning is triggered.

[0042] Furthermore, for the S400, the low threat threshold in multi-source data fusion and threat warning Medium threat threshold and the high threat threshold TH high For the calculation, calculate the mean value μ of the threat index F in historical data Fj and the standard deviation μ Fb , the low threat threshold The calculation formula is:

[0043] The medium threat threshold The calculation formula is:

[0044] The high threat threshold Th high The calculation formula is:

[0045] Where ω3, ω4 and ω5 are coefficients.

[0046] Compared with the prior art, the network information security protection method based on big data has the following beneficial effects:

[0047] First, through social media sentiment mining and user behavior sentiment correlation analysis, the present invention realizes the accurate identification and warning of network information security threats. By using natural language processing technology and constructing an exclusive sentiment dictionary, it can accurately identify potential threat remarks containing negative emotions and related to network security. At the same time, combined with user behavior data, a user behavior sentiment correlation model is established to monitor user behavior in real time. Once the behavior deviates from the baseline and is accompanied by negative emotional fluctuations, it is marked as an abnormal user and the network security threats faced by it are analyzed, which not only improves the accuracy of threat identification, but also provides strong support for taking timely security protection measures.

[0048] Second, by constructing a multi-source data fusion model, the present invention realizes the comprehensive evaluation and warning of network information security threats. In the data collection stage, using web crawler technology and data interfaces, network data including text, user behavior, and network traffic are collected in real time from multiple sources. Subsequently, these data are fused with the results of social media sentiment mining and user behavior sentiment correlation analysis to construct a multi-source data fusion model to calculate threat indicators, and different levels of warning thresholds are set according to the threat indicators. When the fused threat indicator exceeds the threshold, an alarm is triggered, and warning content including threat type, source, and scope of influence information is sent to the administrator or user, which not only improves the comprehensiveness and accuracy of the warning, but also provides a scientific basis for formulating security protection strategies.

[0049] Other advantages, objects, and features of the present invention will be set forth in part in the following description, and in part will be obvious to those skilled in the art based on the examination of the following, or can be learned from the practice of the present invention. Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0051] Figure 1 It is a flowchart of a network information security protection method based on big data;

[0052] Figure 2 It is a process framework diagram of a network information security protection method based on big data. Detailed Embodiments

[0053] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, detail the specific embodiments, structures, features, and their effects of the present invention as follows.

[0054] Embodiment 1:

[0055] Enterprise Internal Network Security Protection

[0056] A large enterprise has many employees. Its internal network connects the office equipment of each department and has a certain interaction with the external Internet. The enterprise adopts the network information security protection method of the present invention to ensure network security.

[0057] Data collection (S100): Use web crawler technology to collect data from social media platforms used within the enterprise (such as enterprise internal forums, instant messaging groups). At the same time, cooperate with network service providers to obtain network traffic data, and install legal data collection plugins on employees' terminal devices to collect user behavior data. For example, within one month, a large amount of text information, employee login times, operation records (such as file upload / download, software usage), and data on the source and target addresses of network access are collected. Clean these data to remove duplicate, invalid, and interfering information, and convert data in different formats into a unified format for subsequent processing.

[0058] Social Media Sentiment Mining (S200): For the enterprise internal social media text data collected, natural language processing techniques are used for word segmentation, part-of-speech tagging, and named entity recognition. Taking the enterprise internal forum as an example, the initial sentiment dictionary contains sentiment words related to "fault", "vulnerability", and network security. When a new technology discussion post appears, such as "The system has been stuck recently, and it feels like there are potential problems", the sentiment dictionary is updated according to the sentiment dictionary update formula. Let the initial sentiment dictionary be E0, and the sentiment words and corresponding sentiment scores it contains are represented as (w0, s0) ∈ E0, where w0 is the sentiment word and s0 is the sentiment score. According to the newly emerged text data D text Update the sentiment dictionary. For each newly emerged sentiment word w′ in D text Use the formula to update its sentiment score. The formula is:

[0059] Calculate the sentiment score of the text data through the sentiment analysis formula to determine whether the text contains negative emotions related to network security. For the text data d, segment it into w1, w2, …, w k Use the updated sentiment dictionary E to calculate its sentiment score S(d). The formula is:

[0060] After analysis, the sentiment score S(d) of this text is lower than the set threshold θ. Mark it as a potential threat statement containing negative emotions related to network security. Then, through clustering and association analysis, mine the characteristics of the theme, source, and spread range of the threat information. Suppose there is text data D threat Marked as a potential threat, represent it as a vector set Use the spectral clustering algorithm to cluster it. The calculation formula is:

[0061] It is found that this type of statement mainly focuses on the usage problems of a specific software module, and the spread range is mainly within the R & D department.

[0062] User Behavior Sentiment Association Analysis (S300): Collect the behavior data of employees' logins, operations, and data accesses. For example, the login frequency of employee A suddenly increased from an average of 3 times a day to 8 times in the past week, and the operation activity index also increased significantly. At the same time, the data access involved some sensitive files outside its daily work scope. Combining the results of social media sentiment mining, establish a user behavior sentiment association model. Let the user behavior feature vector be The sentiment score sequence is S = [s1, s2, …, s m . The calculation formula of the user behavior sentiment association model is:

[0063] Judge the degree of emotional association of the user's behavior, determine the emotional baseline under different behavior patterns, and assume that user U is obtained from social media emotion mining i The emotional score sequence S within the time period t i , t = {s i,t1 , s i,t2 , …, s i,tn}, meanwhile, integrate the login frequency F login , operation activity index A operate , data access anomaly degree E data in the user's behavior data within the same time period t into a behavior feature vector

[0064] For each user U i Use the data of m time periods in the past historical data to determine the emotional state baseline, calculate the mean and standard deviation of each behavior feature to establish a baseline model. For the login frequency F login , its mean and standard deviation The formulas are as follows:

[0065]

[0066] Similarly, calculate the mean and standard deviation of other behavior features, so as to obtain the emotional state baseline i of user U

[0067] and Real-time monitor the user's behavior, calculate the deviation degree from the emotional state baseline. For user U i at the current time t current the behavior feature vector

[0068] Calculate its deviation degree D total from the emotional state baseline. The formula is as follows:

[0069] When the comprehensive deviation degree D total exceeds the set threshold θ total , and the latest emotional score is negative, mark the user as an abnormal user, and analyze the possible network security threats, such as whether there is an account being stolen or malicious operations

[0070] Multi-source data fusion and threat warning (S400): Integrate the results of social media emotion mining and the results of user behavior-emotion association analysis, construct a multi-source data fusion model, calculate threat indicators through the multi-source data fusion model. Assume that the result of social media emotion mining is S(d), and the result of user behavior-emotion association analysis is D total, the calculation formula of the threat indicator F is:

[0071] F = ω1 × (1 + λ2 × S(d)) × D total + ω2 × (1 + λ3 × D total ) × S(d). Let

[0072] The low threat threshold is The medium threat threshold is and the high threat threshold is Th high , and after calculation, the threat indicator exceeds the medium threat threshold triggering a medium-level warning and sending a warning message to the enterprise network security administrator, including that the threat type may be related to abnormal operations of internal personnel and potential system vulnerabilities, the source is the operation behavior and related speech dissemination of employees, and the affected range may involve the R & D department and related data.

[0073] Security protection measures (S500): For the medium-level warning, the enterprise network security team comprehensively checks the internal network services, optimizes the performance of relevant servers and scans for vulnerabilities, and at the same time conducts security checks on employee accounts, including resetting passwords and checking account permission settings for protection measures to prevent potential security risks from further expanding.

[0074] In summary, in the enterprise internal network security protection embodiment, by means of this protection method, the enterprise internal social media and employee behavior data are comprehensively collected and deeply analyzed, potential threats such as abnormal operations of employees are accurately identified, corresponding warnings are triggered according to the threat level and targeted protection is implemented, effectively maintaining the stability and security of the enterprise network environment and protecting the enterprise information assets from infringement.

[0075] Embodiment 2:

[0076] Financial institution network security protection

[0077] A financial institution processes a large amount of customer fund transaction information and has extremely high requirements for network security.

[0078] Data collection (S100): Collect text data related to financial transaction security and system stability from Internet financial forums and social media platforms through web crawlers, and at the same time obtain network traffic data and user operation behavior data in the transaction system from the internal network devices of the financial institution. For example, a large amount of text about the use experience of financial transaction software and network latency conditions, as well as behavior data such as user login time, transaction operation type, and amount are collected within a week, and these data are cleaned and format-converted to ensure the accuracy and availability of the data.

[0079] Social Media Sentiment Mining (S200): Process the collected text data. The initial sentiment dictionary contains words such as "risk" and "unsafe". When new text like "This trading platform has been experiencing frequent trading failures lately, and it feels very unsafe" appears, construct a dedicated dictionary according to the initial sentiment dictionary through the sentiment dictionary update formula. Let the initial sentiment dictionary be E0, and the sentiment words and corresponding sentiment scores it contains are represented as (w0, s0) ∈ E0, where w0 is the sentiment word and s0 is the sentiment score. According to the newly emerged text data D text Update the sentiment dictionary. For each newly emerged sentiment word w′ in D text Use the formula to update its sentiment score. The formula is:

[0080] Calculate the sentiment score of the text data according to the dedicated dictionary through the sentiment analysis formula to determine whether the text contains negative emotions related to network security. For the text data d, tokenize it into w1, w2, …, w k Use the updated sentiment dictionary E to calculate its sentiment score S(d). The formula is:

[0081] If the calculated text sentiment score S(d) is less than the sentiment score threshold θ, mark it as potentially threatening speech after analysis, and through clustering and association analysis, it is found that such speech mainly revolves around the stability issue of the trading system. Let the text data D threat that has been marked as potentially threatening be represented as a vector set Use the spectral clustering algorithm to cluster it. The calculation formula is:

[0082]

[0083] The sources are extensive, and the spread range involves many customers using the trading platform of this financial institution.

[0084] User Behavior Sentiment Association Analysis (S300): Collect the behavior data of users in the financial trading system. For example, users frequently conduct large - amount fund transfer operations in a short period of time, and the login time is abnormal (logging in during non - working hours), and at the same time express dissatisfaction with the security measures of the financial institution on social media. Combine this information to establish a user behavior sentiment association model. Let the user behavior feature vector be =[b1, b2, …, b n , and the sentiment score sequence be S = [s1, s2, …, s m . The calculation formula for the user behavior sentiment association model is:

[0085]

[0086] Judge the degree of user behavior sentiment association, and monitor user behavior in real - time. Let user Ui At the current time t current behavior feature vector

[0087] Calculate the deviation degree D from the emotional state baseline total , and the formula is:

[0088] When the comprehensive deviation degree D total exceeds the set threshold θ total , mark the user as an abnormal user and analyze possible cybersecurity threats, such as whether they have encountered online fraud or account hijacking.

[0089] Multi-source data fusion and threat warning (S400): Integrate the results of social media sentiment mining and user behavior sentiment correlation analysis, construct a multi-source data fusion model to calculate threat indicators, and the calculation formula of threat indicator F is:

[0090] F = ω1×(1 + λ2×S(d))×D total + ω2×(1 + λ3×D total ), given that the low threat threshold is the medium threat threshold is and the high threat threshold is TH high , and after calculation, if the threat indicator exceeds the high threat threshold Th high , trigger a high-level warning, send a warning message to the security management department of the financial institution and relevant regulatory agencies, indicating that the threat type is the possible risk of stolen funds, the source is the abnormal behavior of the user and the spread of relevant negative remarks, and the scope of influence may involve the entire financial trading system and the safety of the funds of many customers.

[0091] Security protection measures (S500): In response to the high-level warning, the financial institution immediately activates the emergency response mechanism, isolates the trading system to prevent the spread of risks, and at the same time comprehensively checks the system logs and transaction records for forensic analysis, traces the flow of abnormal funds, and takes timely measures to ensure the safety of customer funds, such as freezing relevant accounts, contacting affected customers, and reporting detailed situations to regulatory agencies.

[0092] In summary, in terms of the cybersecurity protection of financial institutions, by widely collecting financial-related network texts and user transaction behavior data, accurately identifying risk signs such as abnormal fund transfers, following a strict threat assessment and warning system, quickly responding to high-level risks, comprehensively isolating the system, and conducting in-depth forensic analysis, it effectively protects the safety of customer funds and the reputation of financial institutions, highlighting the important value of this protection method in key areas.

[0093] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments of equivalent changes within the scope of the technical solution of the present invention by using the above-disclosed technical content. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes and modifications made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A method for network information security protection based on big data, characterized in that, The specific steps of this protection method are as follows: S100, Data collection: Through web crawler technology and data interfaces, collect data on text, user behavior, and network traffic from social media platforms, network service providers, and user terminals in real time, and clean, denoise, and convert the format of the collected data. S200, Social media sentiment mining: Use natural language processing technology to segment words, perform part-of-speech tagging, and identify named entities in social media text data. Construct a dedicated dictionary according to the initial sentiment dictionary through the sentiment dictionary update formula. Calculate the sentiment score of the text data according to the dedicated dictionary through the sentiment analysis formula to determine whether the text contains negative emotions related to network security, mark potential threat remarks containing negative emotions and related to network security, and mine the characteristics of the theme, source, and spread range of threat information through clustering and association analysis. S300, User behavior sentiment correlation analysis: Collect behavior data of user logins, operations, and data access, establish a user behavior sentiment correlation model in combination with the results of social media sentiment mining, judge the degree of user behavior sentiment correlation, determine the sentiment baseline under different user behavior patterns, monitor user behavior in real time, calculate the deviation degree from the sentiment state baseline, mark abnormal users according to the deviation degree, and analyze the network security threats faced. S400, Multi-source data fusion and threat warning: Integrate the results of social media sentiment mining and user behavior sentiment correlation analysis, construct a multi-source data fusion model to calculate threat indicators, set different levels of threat warning thresholds according to the threat indicators, trigger a warning when the integrated threat indicator exceeds the threshold, and send a warning content containing information on the threat type, source, and scope of influence to the administrator or user. S500, Security protection measures: Take different measures according to the warning level: In the case of a low-level warning, strengthen monitoring and the collection and analysis of user feedback; in the case of a medium-level warning, conduct a comprehensive inspection of the server and optimize it, and at the same time conduct account security inspections and protection; in the case of a high-level warning, immediately initiate an emergency response, perform system isolation and a comprehensive inspection for evidence collection.

2. The network information security protection method based on big data according to claim 1, characterized in that The S200 constructs an exclusive dictionary according to the initial sentiment dictionary by using the sentiment dictionary update formula in social media sentiment mining. Let the initial sentiment dictionary be E0, and the sentiment words and corresponding sentiment scores included are expressed as (w0, s0) ∈ E0, where w0 is the sentiment word and s0 is the sentiment score. According to the newly emerged text data D text Update the sentiment dictionary. For each newly emerged sentiment word w′ in D text Among them, use the formula to update its sentiment score. The formula is: Among them, TF-IDF(w′, d j ) is the TF-IDF value of word w′ in text d j , and s 0j (d j ) is the sentiment score of text d calculated according to the context j , and the sentiment dictionary is updated.

3. The network information security protection method based on big data according to claim 1, characterized in that The S200 calculates the sentiment score of text data through a sentiment analysis formula in social media sentiment mining, and determines whether the text contains negative emotions related to network security. For the text data d, it is tokenized into w1, w2, …, w k , and uses the updated sentiment dictionary E to calculate its sentiment score S(d). The formula is: Among them, s(w i ) is the sentiment score of the sentiment word w i , e represents the natural constant, g(w i , d) is an importance function of a word w i in the text d, and the formula is: POS(w i ) is the part-of-speech weight of word w i , α1 is the part-of-speech weight adjustment factor, dist(w i , d) is the relative position of word w i in text d, β1 is the position influence factor. Set the sentiment score threshold θ. If the calculated text sentiment score S(d) < θ, mark the text as a potential threat statement related to negative emotions and network security.

4. The network information security protection method based on big data according to claim 1, wherein The S200 performs clustering and correlation analysis on the text data marked as potential threats through the potential threat information clustering and correlation analysis formula in social media sentiment mining, and mines the characteristics of the theme, source and spread scope of the threat information. There is text data D marked as potential threats threat , which is represented as a vector set Use the spectral clustering algorithm to perform clustering on it, and the calculation formula is: where is the Euclidean distance between vectors, sim(w ik , w jk ) is the semantic similarity of the k-th keyword in texts i and j, λ 1k is the weight of the keyword similarity, γ is the distance influence factor, and q is the number of keywords considered.

5. The network information security protection method based on big data according to claim 1, characterized in that In the S300, for the construction of the user behavior-emotion association model in user behavior-emotion association analysis, let the user behavior feature vector be The emotion score sequence is S = [s1, s2, …, s m , and the calculation formula of the user behavior-emotion association model is: where α2 and β2 are weight coefficients, and α2 + β2 = 1, w i is the weight of the behavioral feature b i and r is the weight of the sentiment score s j .

6. The network information security protection method based on big data according to claim 1, characterized in that For the above-mentioned S300, in the determination of the emotional state baseline in user behavior-emotion correlation analysis, it is assumed that user U is obtained from social media emotion mining. i The emotional score sequence S of user U within the time period t i , where t = {s i,t1 , s i,t2 , …, s i,tn}, and s i,tj is the emotional score of user U i at the time point tj. At the same time, the login frequency F login , operation activity index A operate , and data access abnormality E data in the behavior data of the user within the same time period t are integrated into a behavior feature vector [F login , A operate , E data , …]. For each user U i , the data of m time periods in the past historical data are used to determine the emotional state baseline, and the mean and standard deviation of each behavior feature are calculated to establish a baseline model. For the login frequency F login , its mean and standard deviation are calculated by the following formulas: Similarly, the means and standard deviations of other behavior features are calculated, so as to obtain the emotional state baseline i of user U and 7. The network information security protection method based on big data according to claim 1, wherein The S300 is used for detecting and marking abnormal users in the analysis of the association between user behavior and emotion. Let user U i at the current time t current have a behavior feature vector Calculate the deviation degree D from the emotional state baseline total , and the formula is: Among them is the user U i at the current time t current of the k-th behavioral eigenvalue, μ ik and σ ik are the corresponding baseline mean and standard deviation, A k is the deviation weight of the k-th behavioral feature, and the latest sentiment score of the user on the social media is obtained in real time When the comprehensive deviation degree D total exceeds the set threshold θ total , and the latest sentiment score is negative, the user U i is marked as an abnormal user.

8. The network information security protection method based on big data according to claim 1, characterized in that, Regarding the S400, in the construction of the multi-source data fusion model in multi-source data fusion and threat warning, let the social media sentiment mining result be S(d), and the user behavior sentiment correlation analysis result be D total , the calculation formula for the threat index F is as follows: F = ω1×(1 + λ2×S(d))×D total + ω2×(1 + λ3×D total )×S(d), Where ω1 and ω2 are weight coefficients, and λ2 and λ3 are adjustment factors.

9. The network information security protection method based on big data according to claim 1, characterized in that In the S400, for the division of the warning levels in multi-source data fusion and threat warning, let the low-threat threshold be the medium-threat threshold be and the high-threat threshold be Th high ; When No warning is triggered; When a low-level warning is triggered; When Th medium ≤F < Th high , trigger a medium-level warning; When F ≥ Th high , a high-level warning is triggered.

10. The network information security protection method based on big data according to claim 9, wherein The S400, low threat threshold in multi-source data fusion and threat warning Medium threat threshold Th medium and high threat threshold Th high are calculated by calculating the mean value μ Fj and standard deviation μ Fb of the threat index F in historical data. The formula for the low threat threshold is as follows: μ Fb The formula for the medium threat threshold Th medium is as follows: The formula for the high threat threshold Th high is as follows: where ω3, ω4, and ω5 are coefficients.

Citation Information

Cited By

  • Cybersecurity risk calculation and reporting model

    US20260025399A1