Method and system for evaluating danger level of dark web text, electronic equipment and storage medium

By combining models such as BERT, CNN and BiLSTM, the content, users and dissemination mode of dark web text are evaluated, and the problem of the existing technology failing to fully identify potentially dangerous content in dark web text is achieved, and a more comprehensive risk level assessment and higher evaluation results are achieved.

CN119938926APending Publication Date: 2025-05-06BEIJING UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510023680.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When evaluating the hazard level of dark web text, the prior art mainly relies on static content analysis, failing to fully consider potential information such as user behavior and text propagation patterns, resulting in the inability to fully identify potential dangerous content.

Method used

By obtaining dark web text and using the pre-trained language model BERT to generate high-dimensional embedded representations, combining the classification model CNN and the neural network BiLSTM, the text content hazard score, user hazard score and text propagation hazard score are calculated, and the hazard level is comprehensively evaluated and divided.

Benefits of technology

Analyzing the potential dangers of dark web text from multiple perspectives, providing more comprehensive information, helping to discover risk points that deep learning models may ignore, enhance the interpretability of the model, and improve evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938926A_ABST
    Figure CN119938926A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for evaluating the danger level of a dark web text, electronic equipment and a storage medium. The method comprises the following steps: acquiring the dark web text; inputting the dark web text into a language model BERT, and outputting a high-dimensional embedded representation; inputting the high-dimensional embedded representation into a classification model CNN, and outputting features; inputting the features into a neural network BiLSTM, and outputting a text content danger score and a user danger score; calculating a text propagation danger score based on the text forwarding amount, the like amount and the comment amount; and dividing danger levels of the dark web text based on the text content danger score, the user danger score and the text propagation danger score. According to the method, the potential risks of the dark web text can be analyzed from multiple angles, including content features, user behaviors and propagation modes. The dimensions provide more comprehensive information, and can help to find risk points which may be neglected by a deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of danger level of dark web texts, and in particular to a method, system, electronic device and storage medium for evaluating the danger level of dark web texts. Background Art

[0002] With the continuous development of information technology, network information has become more complex and diverse. The dark web, as the hidden layer of the Internet, has become a hotbed for various illegal activities and crimes due to its anonymity and decentralized characteristics. Unlike the surface web, the content in the dark web is usually difficult to monitor and track. Therefore, monitoring and analyzing dark web texts has become one of the key tasks to improve network security, combat criminal activities and protect public safety. However, since dark web texts usually use obscure language or code words to bypass supervision, direct analysis of these texts often faces great challenges. In order to effectively identify potential dangerous content and provide early warnings, it is necessary to systematically classify the danger levels of texts in the dark web.

[0003] At present, the research on risk detection of dark web texts mainly focuses on dark web text topic modeling, machine learning, deep learning and other aspects:

[0004] In terms of dark web text topic modeling, related research mainly focuses on topic modeling of social media content or forum posts on the dark web, using the LDA algorithm to automatically mine hidden topics, identify potential violent speech, extremist ideas or criminal behavior, and conduct risk assessment and grading.

[0005] As for the use of machine learning classification algorithms, some studies usually combine keywords, emotions and context in dark web texts to assess the risks of dark web texts and manually annotate dark web texts. Finally, machine learning classification algorithms are used to automatically analyze dark web texts and classify them according to the degree of danger marked.

[0006] As for deep learning classification algorithms, researchers need to manually label the danger level of dark web texts first, then use pre-trained language models such as BERT and GPT to learn text representation, and combine appropriate classification models such as CNN, LSTM, Transformer, etc. to predict the danger level of dark web texts. Common model combinations include BERT+CNN, BERT+BiLSTM, BERT+CNN+BiLSTM, BERT+BiLSTM+CRF, etc. These combined model combinations can be used to make a good classification of the danger level of dark web texts.

[0007] The existing methods for identifying the danger level of dark web texts are mainly based on static content analysis, such as keyword matching and sentiment analysis, which emphasizes the dangerous content of the text itself. Although these methods can detect directly threatening language such as violence, terror, and drugs, they do not consider other characteristics that can affect the danger of the text, such as the user behavior and historical background of the posting of the text, the degree of spread of the text in social networks, and other potential information. Summary of the invention

[0008] In view of the shortcomings existing in the above-mentioned problems, the present invention provides a method, system, electronic device and storage medium for evaluating the danger level of dark web text.

[0009] To achieve the above object, the present invention provides a method for evaluating the danger level of dark web text, comprising:

[0010] Get dark web text;

[0011] Input the dark web text into the language model BERT and output a high-dimensional embedding representation;

[0012] Input the high-dimensional embedding representation into a classification model CNN and output features;

[0013] Input the features into the neural network BiLSTM, and output the text content risk score and the user risk score;

[0014] At the same time, the text communication risk score is calculated based on the number of text forwarding, likes and comments;

[0015] The dark web text is classified into a danger level based on the text content danger score, the user danger score and the text propagation danger score.

[0016] Preferably, dark web text content, user behavior of posting texts, historical background, text forwarding volume, text comment volume, and text like volume information are obtained by crawling dark web forums and Twitter platforms.

[0017] Preferably, the calculation formula of the risk score is:

[0018] S sum =ω1·S content +ω2·S user +ω3·S spread ;

[0019] Where: S content S is the risk score of the text content; user is the user's risk score; S spread is the text transmission risk score; ω1+ω2+ω3=1, 0<ω i <1.

[0020] Preferably, the text content risk score ranges from 0 to S content ≤10.

[0021] Preferably, the calculation formula of the user risk score is:

[0022]

[0023] Where: M post is the number of all posts of the user, M post ′ is the number of risk control posts of users; N fans The number of followers of the user.

[0024] Preferably, the calculation formula of the text transmission risk score is:

[0025] S spread =G spread ·(α1·N1+α2·N2+α3·N3);

[0026] Where: α1+α2+α3=1,0<α i <1,0 <G spread ≤10; N i It is the ratio between the number of reposts, likes and comments.

[0027] Preferably, N i The calculation formula is:

[0028]

[0029] Where: f1 is the forwarding / like ratio; f2 is the forwarding / comment ratio; f3 is the like / comment ratio.

[0030] This application also provides a system for evaluating the danger level of dark web texts, including:

[0031] Acquisition module, used to obtain dark web text;

[0032] A BERT module, used to input the dark web text into the language model BERT and output a high-dimensional embedding representation;

[0033] A CNN module, used for inputting the high-dimensional embedding representation into a classification model CNN and outputting features;

[0034] A BiLSTM module, used for inputting the features into the neural network BiLSTM, and outputting a text content risk score and a user risk score;

[0035] A calculation module, used to calculate the text transmission risk score based on the text forwarding volume, the number of likes and the number of comments;

[0036] A classification module is used to classify the danger level of the dark web text based on the text content danger score, the user danger score and the text dissemination danger score.

[0037] The present invention also provides an electronic device, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.

[0038] The present invention also provides a storage medium storing a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the above method.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The present invention can analyze the potential dangers of dark web texts from multiple perspectives, including content features, user behavior, and propagation patterns. These dimensions provide more comprehensive information and can help discover risk points that deep learning models may overlook. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flow chart of the method for evaluating the danger level of dark web texts of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] Reference Figure 1 The present invention provides a method for evaluating the danger level of dark web text, comprising:

[0044] Get dark web text;

[0045] Input the dark web text into the language model BERT and output a high-dimensional embedding representation;

[0046] Input the high-dimensional embedding representation into the classification model CNN and output the features;

[0047] Input the features into the neural network BiLSTM and output the text content risk score and user risk score;

[0048] At the same time, the text communication risk score is calculated based on the number of text forwarding, likes and comments;

[0049] Dark web texts are classified into risk levels based on the text content risk score, user risk score and text dissemination risk score.

[0050] In this embodiment, the application crawls dark web forums, Twitter and other platforms to obtain dark web text content, user behavior and historical background of the text, the number of forwarding of the text, the number of comments on the text, the number of likes on the text, and other information. First, the pre-trained model BERT+CNN+BiLSTM is used to judge the risk of the text to determine whether the text is a risk-controlled text. If it is a risk-controlled text, the text content risk score S is obtained. content . At the same time, calculate the risk score S of the user who posted the text user , the spread score S of the text spread By assigning different weights ω1, ω2, ω3 to each indicator, using the weighted formula S sum =ω1S content +ω2S user +ω3S spread , the calculated risk score. sum ≤4, the corresponding dark web text danger level is low. sum ≤7, the corresponding dark web text danger level is medium. sum When ≤10, the corresponding dark web text danger level is high.

[0051] Specifically, the dark web texts are manually labeled for risk. The dark web texts are labeled into four categories: no risk, low risk, medium risk, and high risk, and the corresponding risk scores are S content , 0≤S content ≤10. BERT+CNN+BiLSTM is used to identify content risks of dark web texts. First, the input dark web text is encoded at the word level through BERT. BERT converts each input word into a high-dimensional embedding representation, which contains rich semantic information. The embedding representation of each token generated by BERT is then used as the input of CNN, and multiple convolution kernels of different sizes are applied to the CNN layer to capture features of different granularities. Information is then aggregated through the maximum pooling operation to obtain a more concise text representation. The features extracted by CNN are then further input into the BiLSTM layer. BiLSTM can model contextual information from both ends of the text, thereby better capturing long-term dependencies in the text. Finally, the output of BiLSTM is further converted into a feature representation through a fully connected layer, and then input into the softmax layer to obtain the danger score S of the text content. content .

[0052] ​​​Furthermore, the risk score of the user who posted the text is calculated. First, all posts of the user need to be crawled to obtain the total number M post , use BERT+CNN+BiLSTM to identify each post separately and obtain the total number of all risk control posts M post ′, calculate the risk control post proportion weight Then crawl the number of fans of the user N fans , normalize the maximum value of the number of user fans Calculate the user's risk score

[0053] Furthermore, the forwarding amount N of the text is obtained. forwarding 、N number of likes likes and the number of comments N comments . In the dark web, the amount of forwarding, likes, and comments on a text reflects the diffusion and influence of the text, especially when it involves illegal activities or other high-risk content. However, relying solely on the absolute value of each indicator may not be enough to fully reflect the dissemination effect of the text, so we introduced ratios to quantify the text dissemination characteristics: forwarding / like ratio f1, forwarding / comment ratio f2, and like / comment ratio f3. The higher f1 is, the more likely the text is to be spread. This type of text usually contains inflammatory, extremist, or malicious topics. The higher f2 is, the text is not only widely spread, but also can stimulate discussion among users. This type of text may be controversial. The higher f3 is, the text is easy to be supported or recognized, but it may be simple content, such as extreme views or emotional speech, which is dangerous.

[0054] Specifically, each ratio is normalized to obtain The final calculated dark web text transmission risk score is S spread =G spread ·(α1·N1+α2·N2+α3·N3),α1+α2+α3=1,0<α i <1,0 <G spread ≤10.

[0055] In this embodiment, the text content risk scores S and S are obtained respectively. content , user risk score S user , the spread score S of the text spread . Calculate the total risk score S of the text sum =ω1·S content +ω2·S user +ω3·S spread ,ω1+ω2+ω3=1,0<ω i <1. This is the total risk score of the dark web text. By determining the range of the total risk score, the risk level of the dark web text can be determined.

[0056] This application can analyze the potential dangers of dark web texts from multiple angles, including content features, user behavior, and propagation patterns. These dimensions provide more comprehensive information and can help discover risk points that deep learning models may overlook. Multi-dimensional solutions enhance the interpretability of the model. Deep learning models are usually "black box" models. Especially in complex text classification tasks, it is often very difficult to understand the decision-making process of the model. The multi-dimensional evaluation scheme can provide a more interpretable risk assessment by combining different features and analysis methods. A single deep learning model may be affected by problems such as data distribution, noise, and label inconsistency, resulting in unstable or large deviations in the evaluation results. The multi-dimensional solution can effectively reduce the vulnerability of a single model in specific scenarios by integrating multiple features and technologies.

[0057] This application also provides a system for evaluating the danger level of dark web texts, including:

[0058] Acquisition module, used to obtain dark web text;

[0059] BERT module, which is used to input dark web text into the language model BERT and output high-dimensional embedding representation;

[0060] CNN module, used to input high-dimensional embedding representation into the classification model CNN and output features;

[0061] The BiLSTM module is used to input features into the neural network BiLSTM and output the text content risk score and the user risk score;

[0062] A calculation module, used to calculate the text transmission risk score based on the text forwarding volume, the number of likes and the number of comments;

[0063] The classification module is used to classify the danger level of dark web texts based on the text content danger score, user danger score and text propagation danger score.

[0064] The present invention also provides an electronic device, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.

[0065] The present invention also provides a storage medium storing a computer program executable by an electronic device. When the program runs on the electronic device, the electronic device executes the above method.

[0066] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for evaluating the danger level of dark web text, characterized in that: include: Get dark web texts; Input the dark web text into the language model BERT and output a high-dimensional embedding representation; Input the high-dimensional embedding representation into a classification model CNN and output features; Input the features into the neural network BiLSTM, and output the text content risk score and the user risk score; At the same time, the text communication risk score is calculated based on the number of text forwarding, likes and comments; The dark web text is classified into a danger level based on the text content danger score, the user danger score and the text propagation danger score.

2. The method for evaluating the danger level of dark web text according to claim 1 is characterized in that: By crawling dark web forums and Twitter platforms, we obtain information about dark web text content, user behavior of posting texts, historical background, text forwarding volume, text comment volume, and text like volume.

3. The method for evaluating the danger level of dark web text according to claim 2 is characterized in that: The calculation formula of the risk score is: S sum =ω1·S content +ω2·S user +ω3·S spread ; Where: S content S is the risk score of the text content; user is the user's risk score; S spread is the text transmission risk score; ω1+ω2+ω3=1, 0<ω i <1.

4. The method for evaluating the danger level of dark web text according to claim 1 is characterized in that: The range of the text content risk score is 0≤S content ≤10.

5. The method for evaluating the danger level of dark web text according to claim 4 is characterized in that: The calculation formula of the user risk score is: Where: M post is the number of all posts of the user, M post ′ is the number of risk control posts of users; N fans The number of followers of the user.

6. The method for evaluating the danger level of dark web text according to claim 5 is characterized in that: The calculation formula of the text transmission risk score is: S spread =G spread ·(α1·N1+α2·N2+α3·N3); Where: α1+α2+α3=1,0<α i <1,0 <G spread ≤10; N i It is the ratio between the number of reposts, likes and comments.

7. The method for evaluating the danger level of dark web text according to claim 6 is characterized in that: N i The calculation formula is: Where: f1 is the forwarding / like ratio; f2 is the forwarding / comment ratio; f3 is the like / comment ratio.

8. A system for evaluating the danger level of dark web texts, characterized in that: include: Acquisition module, used to obtain dark web text; A BERT module, used to input the dark web text into the language model BERT and output a high-dimensional embedding representation; A CNN module, used for inputting the high-dimensional embedding representation into a classification model CNN and outputting features; A BiLSTM module, used for inputting the features into the neural network BiLSTM, and outputting a text content risk score and a user risk score; A calculation module, used to calculate the text transmission risk score based on the text forwarding volume, the number of likes and the number of comments; A classification module is used to classify the danger level of the dark web text based on the text content danger score, the user danger score and the text dissemination danger score.

9. An electronic device, characterized in that: The method comprises at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: It stores a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Risk level determination method and device, electronic equipment and storage medium

    CN117670516A