Semi-private domain intelligence information acquisition method and device, electronic equipment and storage medium

By generating keyword databases and building a semi-private domain intelligence knowledge map, the problem that existing technology is difficult to identify and track black industry behavior in a semi-private domain environment is solved, and efficient and comprehensive supervision and protection of black industry activities are achieved.

CN120336316APending Publication Date: 2025-07-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510316194.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When dealing with complex and changing black industry activities, it is difficult for the existing technology to provide comprehensive and real-time protection measures in a semi-private domain environment. The existing methods have limited effects when facing encrypted traffic, and rely on a large amount of training data and high-quality feature extraction technology adaptability and real-time performance, and threat intelligence sharing platforms have problems with information lag and data accuracy.

Method used

By generating keyword databases based on intelligence topics, using pre-trained keyword models to extract original public intelligence from multi-source data, compute threat values to screen high-threat intelligence, build a semi-private domain intelligence knowledge graph, and realize the identification and tracking of black industry behavior.

Benefits of technology

Effectively identifying and tracking black industry behavior in a semi-private domain environment improves the comprehensiveness and real-time nature of intelligence analysis, enhances the ability to supervise black industry activities, and provides more comprehensive protective measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336316A_ABST
    Figure CN120336316A_ABST
Patent Text Reader

Abstract

The invention provides a semi-private domain intelligence information acquisition method and device, electronic equipment and a storage medium, relates to the technical field of intelligence analysis, in particular to the fields of threat intelligence mining, intelligence association analysis and the like, and can be used for application scenes of security risk assessment, risk early warning and the like. According to the specific implementation scheme, a keyword library is generated by utilizing a pre-trained keyword model based on an intelligence theme; extracting original public intelligence from the data source based on the keyword library; calculating a threat value of the original public intelligence to obtain a high threat public intelligence set; determining semi-private domain content based on the high threat public intelligence set; and according to the semi-private domain content, creating a semi-private domain intelligence knowledge graph. According to the scheme, black production behaviors can be effectively identified and tracked in a semi-private domain environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligence analysis technology, especially to fields such as threat intelligence mining and intelligence correlation analysis, and can be used in application scenarios such as security risk assessment and risk warning. Specifically, it relates to a method, device, electronic device, and storage medium for obtaining semi-private domain intelligence information. Background Art

[0002] In the current network environment, many black production behaviors are carried out in a private or semi-private environment, which makes the spread of black production technology more concealed and increases the difficulty of supervision and tracking. Existing technologies are unable to cope with complex and changeable black production activities and are difficult to provide comprehensive and real-time protection measures. Summary of the Invention

[0003] The present disclosure provides a method, device, electronic device, and storage medium for obtaining semi-private domain intelligence information.

[0004] According to a first aspect of the present disclosure, there is provided a method for obtaining semi-private domain intelligence information, including: generating a keyword library based on an intelligence theme by using a pre-trained keyword model; extracting original public intelligence from a data source based on the keyword library; calculating a threat value of the original public intelligence to obtain a high-threat public intelligence set; determining semi-private domain content based on the high-threat public intelligence set; and creating a semi-private domain intelligence knowledge graph according to the semi-private domain content.

[0005] According to a second aspect of the present disclosure, there is provided a device for obtaining semi-private domain intelligence information, including: a thesaurus generation module for generating a keyword library based on an intelligence theme by using a pre-trained keyword model; an intelligence collection module for extracting original public intelligence from a data source based on the keyword library; a threat determination module for calculating a threat value of the original public intelligence to obtain a high-threat public intelligence set; a content acquisition module for determining semi-private domain content based on the high-threat public intelligence set; and a graph creation module for creating a semi-private domain intelligence knowledge graph according to the semi-private domain content.

[0006] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.

[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.

[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method according to any one of the embodiments of the present disclosure.

[0009] By adopting the solution of the present disclosure, it is possible to effectively identify and track black production behaviors in a semi-private domain environment.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is a schematic flowchart of a method for obtaining semi-private domain intelligence information according to an embodiment of the present disclosure;

[0013] Figure 2 is a schematic flowchart of generating semi-private domain content according to an embodiment of the present disclosure;

[0014] Figure 3 is a schematic structural diagram of a device for obtaining semi-private domain intelligence information according to an embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram of a scenario of a method for obtaining semi-private domain intelligence information according to an embodiment of the present disclosure;

[0016] Figure 5 is a structural diagram of an electronic device for implementing the method for obtaining semi-private domain intelligence information according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0018] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. As used herein, the term "at least one" means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may mean including any one or more elements selected from the set consisting of A, B, and C. As used herein, the terms "first" and "second" are used to refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature may be one or more, and the second feature may also be one or more.

[0019] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.

[0020] Before introducing the technical solutions of the embodiments of the present disclosure, further explanations are made on the technical terms that may be used in the present disclosure:

[0021] Semi-private domain: refers to data or content that is not completely public but also not completely private. Such data or content may require specific permissions, authorizations, or other conditions to be accessed and is not completely open to all the public. Moreover, such data or content is not strictly confidential classified data and may have been disseminated or shared within a certain range, but the range is limited. In the embodiments of the present disclosure, the semi-private domain often refers to those intelligence data sources or closed communities that require authorized access.

[0022] In related technologies, the techniques for combating black production activities mainly include: methods based on network traffic analysis, that is, by detecting abnormal traffic patterns or specific communication protocol features to identify potential black production activities, which are commonly used to identify malware activities and abnormal access behaviors in the network; methods based on data mining and machine learning technologies, that is, extracting features from a large amount of log and behavior data and establishing models to automatically identify and classify black production behaviors. This method usually requires a large amount of training data and computing resources, and has limited recognition ability for new attacks; methods based on threat intelligence sharing platforms, that is, exchanging black production-related information between different organizations and security vendors to improve the overall threat perception ability. This method relies on the cooperation of participants and the timeliness of information. Although the existing technologies provide effective solutions in some aspects, there are still many deficiencies. First, the network traffic analysis method has limited effectiveness in the face of encrypted traffic and is easily circumvented by black production gangs through obfuscation techniques. Second, the technologies based on data mining and machine learning rely on a large amount of labeled data and high-quality feature extraction, and their adaptability and real-time performance are challenged in the face of rapidly changing attack methods. In addition, there are problems of information lag and data accuracy in the threat intelligence sharing platform, and the information asymmetry between different organizations also affects its effectiveness. These deficiencies make the existing technologies appear powerless in dealing with complex and ever-changing black production activities and are difficult to provide comprehensive and real-time protection measures. Therefore, new methods are urgently needed to make up for these defects, especially in the semi-private domain environment, to more effectively identify and track black production behaviors.

[0023] In order to at least partially solve one or more of the above problems and other potential problems, the present disclosure proposes a method for obtaining semi-private domain intelligence information, which can effectively identify and track black production behaviors in the semi-private domain environment.

[0024] An embodiment of the present disclosure provides a method for obtaining semi-private domain intelligence information. Figure 1 FIG. is a schematic flowchart of the method for obtaining semi-private domain intelligence information according to an embodiment of the present disclosure. The method for obtaining semi-private domain intelligence information can be applied to a semi-private domain intelligence information obtaining device. The semi-private domain intelligence information obtaining device is located in an electronic device. The electronic device includes but is not limited to a fixed device and / or a mobile device. For example, the fixed device includes but is not limited to a server, and the server can be a cloud server or a general server. For example, the mobile device includes but is not limited to a risk assessment device and a risk warning device, and the risk assessment device and the risk warning device can be a mobile phone, a tablet computer, etc. In some possible implementation manners, the method for obtaining semi-private domain intelligence information can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, the method for obtaining semi-private domain intelligence information includes:

[0025] S101. Generate a keyword library based on the intelligence theme using a pre-trained keyword model.

[0026] S102. Extract the original public intelligence from the data source based on the keyword library.

[0027] S103. Calculate the threat value of the original public intelligence to obtain a high-threat public intelligence set.

[0028] S104. Determine the semi-private domain content based on the high-threat public intelligence set.

[0029] S105. Create a semi-private domain intelligence knowledge graph according to the semi-private domain content.

[0030] Here, the intelligence theme is the core target or direction of intelligence analysis or mining, that is, the area or problem that needs to be focused on. In the embodiments of the present disclosure, the intelligence theme can represent the specific target of the intelligence mining task. Exemplarily, the intelligence theme can be set as "cyber attack activities", "data leakage incidents", or "potential threats of a specific organization", etc.

[0031] Here, the keyword is the core vocabulary or phrase related to the intelligence theme, used to extract relevant intelligence information from the data source. In the embodiments of the present disclosure, the keyword can be used to filter or match the information in the data source, thereby helping to extract the content directly related to the intelligence theme.

[0032] Here, the keyword model is a model based on machine learning, natural language processing (Natural Language Processing, NLP) or expert experience, specifically used to generate or optimize the keyword library. In the embodiments of the present disclosure, the keyword model can automatically generate keywords related to the intelligence theme by analyzing existing data or corpora, avoiding the deficiencies of manual maintenance of the keyword library.

[0033] Here, the keyword library is a set of keywords related to the intelligence theme, usually stored in a structured form. In the embodiments of the present disclosure, the keyword library is the core tool for mining intelligence, providing a series of keywords for screening and analyzing information to ensure the accuracy and comprehensiveness of intelligence extraction.

[0034] In the embodiments of the present disclosure, the goal and scope of the intelligence task can be clarified first, the direction to be mined can be determined, and then the intelligence theme can be defined. Subsequently, the determined intelligence theme is input into the keyword model to generate a keyword set. Finally, the keyword set is further cleaned and optimized to form a high-quality keyword library.

[0035] Here, the data source refers to the original data source on which intelligence mining depends. In the embodiments of the present disclosure, the data source may include public data from various sources, which can provide content for intelligence analysis and provide a basis for obtaining threat intelligence.

[0036] Here, the original public intelligence refers to the unprocessed and unanalyzed intelligence content extracted from the data source. In the embodiments of the present disclosure, the original public intelligence is the initial result of intelligence mining and needs to be further processed, screened, and analyzed before it can be transformed into useful threat intelligence.

[0037] In the embodiments of the present disclosure, web crawler technology can be used to automatically crawl the content of the data source through a program to obtain the original data. Subsequently, the keywords in the keyword library can be matched with the original data content, and the data entries containing the keywords can be screened out, and then the original public intelligence can be generated. Finally, the original public intelligence can be stored in a structured manner for subsequent analysis and processing.

[0038] Here, the threat value is an indicator for quantitatively evaluating the threat level of a piece of intelligence content. In the embodiments of the present disclosure, the threat value can help screen out high-threat intelligence information, reduce information noise, and focus on key threat events. The threat value can be quantitatively evaluated from multiple factors such as the confidentiality, integrity, availability, scope of influence, and attacker capabilities of the intelligence. Among them, confidentiality represents the degree of exposure of sensitive information that may be caused by the leakage of intelligence. Integrity represents the possibility of the intelligence being tampered with or damaged and its impact. Availability represents the potential impact of the interruption of the intelligence or related system services. The scope of influence represents the scope and severity that may be affected by the leakage or attack of the intelligence. Attacker capabilities represent the technical level, resources, and motivation of potential attackers. After comprehensively evaluating these factors, the threat value of the intelligence can be quantified. Exemplarily, to evaluate the threat value of a certain intelligence leakage, assume that the intelligence contains highly sensitive customer data and the confidentiality score is 9 (10 is the highest threat); the intelligence may be tampered with, but the existing protection measures are relatively strong, and the integrity score is 5; the leakage of the intelligence may cause a short-term interruption of the system, but the impact is limited, and the availability score is 4; the leakage of the intelligence may affect 1000 customers and involve financial information, and the scope of influence score is 7. The attacker has a medium technical level and limited resources, and the attacker capability score is 6. Assuming that each item has the same weight, the average score is (9 + 5 + 4 + 7 + 6) / 10 = 3.1 points, indicating that this intelligence leakage has a relatively low threat. It should be noted that the weight ratio of each item can be adjusted according to requirements. The above is only an exemplary illustration and does not limit all possible situations of multiple factors. Here, exhaustive enumeration is not done. In the embodiments of the present disclosure, the evaluation criteria can be defined first according to the actual intelligence theme and application scenario. Exemplarily, the evaluation criteria can include relevance, credibility, severity, immediacy, scope of dissemination, etc. Subsequently, according to the evaluation criteria, the threat value of the original publicly available intelligence is calculated. Exemplarily, the threat value can be calculated based on set rules, or a machine learning model can be used to predict the threat value of new intelligence in combination with historical data. Finally, according to the calculated threat value, a set of high-threat publicly available intelligence can be screened out. Exemplarily, a fixed threat value threshold can be set or a dynamic threshold can be set according to the actual situation. By comparing the calculated threat value with the threshold, a set of high-threat publicly available intelligence can be screened out; similarly, the intelligence can also be sorted from high to low according to the threat value, and a certain proportion of the highest-threat intelligence can be selected as high-threat intelligence, and then the high-threat intelligence can be grouped into a set. In particular, the set of high-threat publicly available intelligence screened out can be stored in a structured manner for convenient subsequent analysis and processing.

[0039] In the embodiments of the present disclosure, the high-threat intelligence set can provide clues pointing to semi-private content. Therefore, associated clues can be extracted from high-threat public intelligence, and then the semi-private content can be gradually explored. Specifically, the high-threat public intelligence may contain technical indicators, and these indicators can be used to further track relevant content in the semi-private domain. Exemplarily, the technical indicators may include Internet Protocol (IP) addresses, domain names, file hash values, vulnerability numbers, etc., which can be obtained through self-built crawlers or data scraping attacks. Further, entities that may be related to the semi-private domain can also be extracted through text mining of the public intelligence content. Exemplarily, the entities may include organization names, tool or software names, implementation methods, attack targets, etc. Further still, the content in the semi-private domain usually revolves around hot events, and through timeline analysis, high-threat intelligence alerts can be associated with the discussions or activities in the semi-private domain. Exemplarily, based on the release time of the public intelligence, the activities in the semi-private domain around that time point can be focused on. Finally, through the extracted clues, the potentially hidden semi-private domain can be gradually explored, and then the semi-private content can be determined. Exemplarily, exploration can be carried out through in-depth search engines or web crawler technologies, or through closed social media groups and forums. The above is only an exemplary illustration and does not serve as a limitation on all possible situations for determining semi-private content, and exhaustive listing is not done here.

[0040] Here, the semi-private domain intelligence knowledge graph is an intelligence data visualization and association analysis tool constructed using knowledge graph technology, which can be used to reveal the entities and their relationships related to semi-private domain intelligence. In particular, the semi-private domain intelligence knowledge graph can structure, associate, and visualize complex semi-private domain intelligence data to support threat tracking, intelligence association, and security decision-making.

[0041] In the embodiments of the present disclosure, since the sources of semi-private content are complex, the semi-private content can be first cleaned and standardized to ensure data quality, and then the semi-private content can be analyzed to extract key entities and the relationships between the entities, and then a semi-private domain intelligence knowledge graph can be created. Specifically, after cleaning and standardizing the semi-private content, specific types of entities can be first extracted from the semi-private content. Exemplarily, the entities may include attackers, attack targets, attack indicators, attack methods, etc., and regular expressions or named entity recognition tools can be used for extraction. Further, the relationships between the entities can be extracted from the semi-private content. Exemplarily, the relationships may include the relationship between an attacker and a tool, an attacker and a target, a tool and a target, etc., and methods such as syntactic dependency analysis, text matching, or deep learning models can be used for extraction. Further still, according to the predefined entity types, relationship types, and attributes, the extracted entities and relationships can be used as the nodes and edges of the knowledge graph respectively, thus creating a semi-private domain intelligence knowledge graph.

[0042] In the technical solution of the embodiments of the present disclosure, through the keyword model, a keyword library related to a specific intelligence theme can be quickly generated, ensuring the comprehensiveness and pertinence of the keyword library, and can dynamically adapt to the changes in the intelligence field. By using the keyword library and threat value calculation, important information related to the intelligence theme can be quickly identified from a large amount of data, ensuring the extensiveness and diversity of the intelligence coverage. At the same time, the threat value calculation assigns a quantitative standard to the intelligence, which helps to quickly screen out high-threat intelligence and provides a clear priority basis for subsequent analysis and processing. Further mining semi-private content from the high-threat public intelligence set can obtain deeper and more valuable intelligence information. Using the semi-private intelligence knowledge graph to structure and relate complex intelligence data can systematically and structurally present the intelligence relationship network, significantly improving the comprehensiveness and accuracy of the knowledge graph.

[0043] In some embodiments, based on the intelligence theme, a keyword library is generated using a pre-trained keyword model, including: analyzing historical data to determine the intelligence theme; using the keyword model to extract the topic words in the intelligence theme and establish a topic word library; using the keyword model to diverge the semantic meaning of any topic word in the topic word library to generate a keyword library.

[0044] In the embodiments of the present disclosure, the determination process of the intelligence theme depends on the analysis of historical data. Specifically, the historical data sources related to the intelligence can be obtained first. After appropriate data cleaning and format unification, word segmentation and language processing can be performed on the historical data to extract basic words or phrases. Further, the word frequency of the segmented text can be counted to find the words with higher frequencies, and combined with the context, the words representative of the intelligence theme can be screened out. Then, multiple themes can be extracted from the text using the topic modeling method to determine the intelligence theme.

[0045] Here, the topic word refers to the core keyword of the intelligence theme, which is a key term that can highly summarize or represent the intelligence theme and can be used to describe the core content of the intelligence. In the embodiments of the present disclosure, the topic word is directly highly related to the intelligence theme, can represent the core information of the theme, and at the same time, the topic word is often a common term in the field and has a certain universality. In particular, the topic word sets of different intelligence themes are different.

[0046] In the embodiments of the present disclosure, the topic words of the intelligence theme can be extracted from historical data based on the keyword model. Exemplarily, for each intelligence theme, the high-frequency words or high-weight words related to the theme can be calculated first. After filtering out irrelevant words, combined with the prior knowledge pre-integrated, the keyword set can be refined.

[0047] In the embodiments of the present disclosure, a larger keyword library can be generated by expanding the topic words, covering more terms related to the intelligence topic. Specifically, synonyms or near-synonyms can be generated for each topic word using a keyword model first, then extended keywords can be generated using context semantics, and finally the topic words can be extended in combination with the domain to generate a keyword library. Exemplarily, the process of expanding topic words can be combined with capitalization and lowercase, homophony, misspelled words, slang, abbreviations, long-tail keywords, etc., or Chinese words can be converted into corresponding English terms and slang. In particular, relevant priorities can also be set between keywords according to their derivation methods for subsequent viewing.

[0048] In this way, by deeply analyzing historical data, the current intelligence topic can be accurately identified and determined, ensuring the directionality and pertinence of subsequent keyword extraction and intelligence mining. By extracting the core topic words in the intelligence topic through a keyword model, a high-quality topic word library can be quickly constructed, significantly improving the accuracy and comprehensiveness of topic word extraction, and providing reliable data support for the subsequent expansion and application of the topic word library. By expanding topic words through semantic divergence, a keyword library with a wider coverage and stronger relevance can be generated, not only enriching the content of the keyword library, but also improving the diversity and relevance of keywords, providing a more comprehensive data basis for subsequent intelligence mining and analysis, and significantly enhancing the coverage and relevance of the knowledge graph.

[0049] In some embodiments, the data sources at least include: surface web information, dark web information, and instant messaging application data.

[0050] Here, surface web information refers to data that can be publicly accessed through the Internet and can be obtained without special permissions or special tools. In the embodiments of the present disclosure, surface web information may include public websites, social media, public databases, and search engine index data, etc.

[0051] Here, dark web information refers to data existing in the dark web, content that needs to be accessed through special tools or permissions, and these data are usually not indexed by conventional search engines. In the embodiments of the present disclosure, dark web information may include illegal transactions, hacker forums, data leaks, hidden services, etc.

[0052] Here, instant messaging (IM) application data refers to the information content transmitted through IM applications. In the embodiments of the present disclosure, IM application data may include chat group data, private chat messages, and file transfer data, etc.

[0053] In the embodiments of the present disclosure, the data sources may also include: network traffic data, log data, open vulnerability databases, government intelligence and industry reports, social media data, etc., which can be selected according to actual situations.

[0054] In this way, by integrating clear web information, dark web information, and instant messaging tool application data, it is possible to comprehensively cover intelligence content in the public, semi-public, and concealed domains, make up for the possible blind spots of a single data source, significantly improve the comprehensiveness, depth, and timeliness of the knowledge graph, and provide a more efficient and accurate tool for intelligence analysis and decision-making.

[0055] In some embodiments, based on a keyword library, raw public intelligence is extracted from a data source, including: extracting public content from the data source based on the keyword library; and marking the public content according to a preset tag type to obtain raw public intelligence.

[0056] In the embodiments of the present disclosure, the keyword library can be used as a content screening condition to traverse the content of the data source, match the vocabulary in the keyword library, extract the content that meets the conditions, and use the extracted content as the public content.

[0057] Here, the tag type refers to a preset category for classifying and identifying the data of the data source, and can be used to mark the theme, nature, or field of the public content. The role of the tag is to classify a large amount of crawled content according to its semantics, theme, or use for subsequent screening, analysis, and retrieval. In the embodiments of the present disclosure, the tag type can be a relevant black production vocabulary library, that is, terms or keywords related to network black production, which can reflect the specific category or theme of the public content.

[0058] In the embodiments of the present disclosure, the process of marking the public content can be achieved through keyword matching and semantic analysis, and then the public content is associated with the preset tag type. Specifically, matching rules can be defined for each tag based on the black production vocabulary library. Exemplarily, keyword matching rules or semantic matching rules can be adopted. Further, each piece of content can be traversed and sequentially matched with the tag keywords. If certain keywords or highly relevant semantics are included in the content, the corresponding tag is assigned to it. Exemplarily, natural language analysis tools such as keyword extraction tools or pre-trained models can be used to achieve this. Finally, the marked public content can be saved as structured data for storage.

[0059] In this way, by extracting public content from multi-source data through the keyword library, it is possible to quickly and accurately screen out high-value information related to the intelligence theme. At the same time, the comprehensiveness and relevance of the keyword library ensure the coverage range and depth of the extracted content, enhancing the comprehensiveness and pertinence of intelligence analysis. By marking the public content with a preset tag type, it is possible to systematically and structurally classify and organize intelligence data, significantly improving the manageability and analyzability of intelligence data, providing an efficient and accurate tool for the construction and application of the knowledge graph, and ultimately improving the comprehensiveness and practicality of intelligence analysis.

[0060] In some embodiments, public content is extracted from a data source based on a keyword library, including: performing keyword matching on the data source information using natural language processing techniques to obtain first public content; performing semantic analysis on the data source information using a deep learning model to obtain second public content; and performing data processing on the first public content and the second public content to obtain target public content.

[0061] Here, the first public content is the public information extracted by performing keyword matching on the data source information through NLP techniques. The keyword matching is based on a preset keyword library, which matches the text in the data source with the words in the keyword library and filters out the content containing these keywords.

[0062] In the embodiments of the present disclosure, after preprocessing the data source information, a string search algorithm or a regular expression can be used to identify keywords, and each piece of data is scanned item by item to extract the fragments containing the keywords, thereby obtaining the first public content.

[0063] Here, the second public content is the public information extracted by performing semantic analysis on the data source information through a deep learning model. Semantic analysis not only relies on keyword matching but can also use the deep learning model to understand the context and implicit semantics of the text, thereby extracting content that is not explicitly covered by the keyword library but is semantically relevant.

[0064] In the embodiments of the present disclosure, after preprocessing the data source information, a deep learning model can be used to convert the text into a semantic vector representation, and at the same time, each keyword in the keyword library is vectorized, and then the similarity with the text semantic vector is calculated, and the second public content is obtained by setting a similarity threshold for screening. Similarly, a deep learning model can also be used to comprehensively analyze the text context, so as to extract the content related to the keyword library after semantic analysis as the second public content.

[0065] In the embodiments of the present disclosure, there may be duplicate or redundant parts between the first public content and the second public content, and they may also need to be further classified, formatted, or cleaned. Through the data processing step, the two can be integrated into the final public content.

[0066] In this way, the first public content quickly extracts explicitly relevant content through keyword matching, covering the clear information in the intelligence data; the second public content mines implicit information through semantic analysis, supplementing the intelligence content not covered by the keywords. By comprehensively processing the first public content and the second public content, the results of keyword matching and semantic analysis can be integrated, eliminating duplicate information and supplementing missing content, and finally obtaining more comprehensive and accurate public content. This method significantly improves the integrity and consistency of the intelligence data and provides high-quality and highly reliable data support for subsequent intelligence analysis and knowledge graph construction.

[0067] In some embodiments, data processing is performed on the first disclosure and the second disclosure to obtain a disclosure, including: combining and deduplicating the first disclosure and the second disclosure to obtain the disclosure to be processed; and performing data cleaning on the disclosure to be processed to obtain the target disclosure.

[0068] In the embodiments of the present disclosure, the text similarity algorithm can be first used to compare the first disclosure and the second disclosure to remove exactly the same or highly similar entries. Subsequently, the first disclosure and the second disclosure can be supplemented and integrated, and then the integrated content can be classified and labeled for subsequent analysis.

[0069] In this way, through data combination and deduplication, redundant information is eliminated, and the quality and accuracy of intelligence are improved. Through data cleaning, the degree of structuring and usability of the data are improved, so that the output structured data can be directly used for threat intelligence analysis, report generation, support for intelligence sharing, and subsequent business requirements.

[0070] In some embodiments, the target disclosure is labeled according to a preset tag type to obtain the original public intelligence, including: obtaining the basic information of the target disclosure according to the preset tag content; the basic information at least includes: the platform where the target disclosure is located, the link of the target disclosure, the time of the target disclosure, and the user account related to the target disclosure; generating metadata of the disclosure according to the basic information; and combining the target disclosure and the metadata to generate the original public intelligence.

[0071] Here, the basic information refers to the core description fields extracted from the disclosure, which can be used to identify key information such as the source, time, publisher, etc. of the disclosure. In the embodiments of the present disclosure, the basic information is a summary information extraction of the target disclosure, which can help quickly understand the background and attributes of the content. Exemplarily, the basic information may include the platform where the disclosure is located, the link of the disclosure, the time of the disclosure, the user account related to the disclosure, as well as the content type, IP address, geographical location, associated content, etc.

[0072] In the embodiments of the present disclosure, the basic information of the disclosure can be obtained according to the preset tag content. Specifically, since the process of labeling the target disclosure can be realized through keyword matching and semantic analysis, that is, the target disclosure is associated with the preset tag type. Therefore, the basic information of the target disclosure can be directly extracted from the tag content.

[0073] In the embodiments of the present disclosure, the extracted basic information can be mapped to metadata fields, and the standardized fields can be organized into a unified structure to generate the metadata of the target disclosure content. In particular, if some fields are missing, they can be filled by inference or default values.

[0074] In the embodiments of the present disclosure, the target disclosure content includes the original text information, and the metadata includes the standardized descriptive information. In the process of merging the target disclosure content and the metadata, the target disclosure content can be used as the core content of the intelligence, and the metadata can be nested in the intelligence as descriptive information. Finally, according to the intelligence storage requirements, the merged intelligence can be saved in a structured data format. Exemplarily, the crawled posts can be saved in a structured form of "tag: platform: link: title + time: user: content: comment user + comment + time".

[0075] In this way, by obtaining the basic information according to the preset tags, generating the metadata, and integrating the target disclosure content and the metadata to generate the original disclosure intelligence, the method realizes the structured, standardized, and automated processing of intelligence data, significantly improves the accuracy and efficiency of intelligence mining, and provides high-quality data support for intelligence analysis and threat detection.

[0076] In some embodiments, calculate the threat value of the original disclosure intelligence to obtain a high-threat disclosure intelligence set, including: calculating the threat value of the original disclosure intelligence, sorting the threat values of the original disclosure intelligence to obtain an original disclosure intelligence sequence; extracting the original disclosure intelligence with a threat value greater than a preset threat threshold in the original disclosure intelligence sequence to generate a high-threat disclosure intelligence set.

[0077] In the embodiments of the present disclosure, the threat value of the original disclosure intelligence needs to be quantified by comprehensively considering various factors according to the intelligence content and its associated information. Specifically, the threat value can be calculated based on aspects such as tag weight, content sensitivity, source credibility, and time factor. Among them, the tag weight can refer to the weight assigned according to the tags of the intelligence content; the content sensitivity can refer to whether high-sensitive information is mentioned in the intelligence; the source credibility can refer to the evaluation of its credibility based on the historical threat behavior of the intelligence source platform or user account; the time factor can refer to that the intelligence closer to the current time may pose a greater threat. Further, in the process of calculating the threat value of the original disclosure intelligence, first, for the original disclosure intelligence, extract its corresponding tags, sensitive content, source information, and timestamp. Subsequently, the threat value can be calculated according to the predefined weight table and coefficient table.

[0078] In the embodiments of the present disclosure, in the process of sorting the original disclosure intelligence in descending order according to its threat value, sorting algorithms such as quick sort or merge sort can be used to sort the threat value calculation results.

[0079] In the embodiments of the present disclosure, a threat value threshold may be preset according to business requirements first. When the threat value is greater than the threat value threshold, it may indicate that the intelligence poses a relatively high threat to the business. Subsequently, the intelligence with a threat value greater than the preset threshold may be screened from the sorted original publicly available intelligence sequence. Finally, the set of extracted high-threat intelligence may be stored as structured data, and all the finally generated structured data may be used as the set of high-threat publicly available intelligence.

[0080] In this way, by calculating the threat value, sorting the intelligence, and screening the high-threat intelligence, the efficient extraction and processing of high-threat intelligence from the original intelligence can be achieved, improving the efficiency and accuracy of intelligence processing, enhancing the usability and interpretability of the intelligence, and providing effective support for security decision-making and threat response.

[0081] In some embodiments, calculating the threat value of the original publicly available intelligence, sorting the threat values of the original publicly available intelligence to obtain the original publicly available intelligence sequence, includes: calculating the platform threat value according to the original publicly available intelligence and the preset platform threat coefficient; calculating the relevance threat value according to the original publicly available intelligence and the preset relevance threat coefficient; calculating the timeliness threat value according to the original publicly available intelligence and the preset timeliness threat coefficient; calculating the negative keyword threat value according to the original publicly available intelligence and the preset negative keyword threat coefficient; calculating the threat value of the original publicly available intelligence by using the platform threat value, the relevance threat value, the timeliness threat value, and the negative keyword threat value according to the preset threat ratio.

[0082] In the embodiments of the present disclosure, the platform threat coefficients corresponding to different platforms may be predefined. Exemplarily, relatively high platform threat coefficients may be assigned to high-risk platforms such as dark web forums and hacker communities, and relatively low platform threat coefficients may be assigned to ordinary social media platforms or public forums. In particular, the platform threat coefficients may be preset according to historical records or business experience. In the actual scenario, the source platform of the original publicly available intelligence may be extracted first, the corresponding platform threat coefficient of the platform may be found, and the platform threat coefficient may be used as the platform threat value. Alternatively, according to other metadata information of the original publicly available intelligence, combined with the platform threat coefficient corresponding to the original publicly available intelligence, the calculation result may be used as the platform threat value.

[0083] In the embodiments of the present disclosure, different correlation threat coefficients corresponding to different correlations between intelligence content and historical records, target systems, or tags can be predefined. Exemplarily, if the content mentioned in the original public intelligence is highly relevant to a certain known attack tool, attack means, or target system, a relatively high correlation threat coefficient can be assigned. For other irrelevant content, a relatively low correlation threat coefficient can be assigned. In particular, the correlation threat coefficient can be preset according to historical records or business experience. In an actual scenario, keyword matching or semantic analysis techniques can be first used to calculate the correlation of the intelligence content. According to the analysis results, the correlation threat coefficient is assigned, and the correlation threat coefficient is used as the correlation threat value. Alternatively, based on other metadata information of the original public intelligence, the calculation can be performed in combination with the correlation threat coefficient corresponding to the original public intelligence, and the calculation result is used as the correlation threat value.

[0084] In the embodiments of the present disclosure, different timeliness threat coefficients corresponding to the difference between the release time and the current time can be predefined. Exemplarily, the closer the release time of the original public intelligence is to the current time, the greater the threat possibility. In particular, the timeliness threat coefficient can be preset according to historical records or business experience. In an actual scenario, the release time of the intelligence can be first extracted, the difference from the current time is calculated, and the predefined timeliness threat coefficient is found according to the time difference, and the timeliness threat coefficient is used as the timeliness threat value. Alternatively, based on other metadata information of the original public intelligence, the calculation can be performed in combination with the timeliness threat coefficient corresponding to the original public intelligence, and the calculation result is used as the timeliness threat value.

[0085] Here, negative keywords refer to specific words or phrases that need to be excluded or marked as irrelevant, unimportant, or even potentially reducing the intelligence value. These keywords are usually predefined, have nothing to do with the target intelligence analysis requirements, or may interfere with the accuracy of threat intelligence screening. Negative keywords can help eliminate low-value intelligence or reduce its threat assessment score, so that the finally screened intelligence is more accurate.

[0086] In the embodiments of the present disclosure, a negative keyword set can be predefined, and a corresponding negative keyword threat coefficient is set for each negative keyword. In particular, the negative keyword threat coefficient is usually negative, indicating the weight for reducing the intelligence threat value. In an actual scenario, the negative keyword list of the original public intelligence can be first matched to obtain the negative keyword threat coefficient corresponding to the original public intelligence. In particular, if one or more negative keywords are matched in the original public intelligence, the corresponding threat coefficients can be accumulated. Subsequently, the negative keyword threat coefficient is used as the negative keyword threat value. Alternatively, based on other metadata information of the original public intelligence, the calculation can be performed in combination with the negative keyword threat coefficient corresponding to the original public intelligence, and the calculation result is used as the negative keyword threat value.

[0087] In the embodiments of the present disclosure, weights can be set in advance for the platform threat value, the relevance threat value, the timeliness threat value, and the negative keyword threat value, and can be dynamically adjusted according to actual business needs. In an actual scenario, a weighted formula can be used to calculate the threat value of the original public intelligence.

[0088] In this way, the threat value is divided into multiple dimensions to comprehensively evaluate the threat level of the intelligence, avoiding biases caused by a single indicator, and the weight of each dimension can be adjusted, enabling flexible adaptation to different business requirements and scenarios and improving the accuracy of the threat value.

[0089] In some embodiments, based on the high-threat public intelligence set, semi-private domain content is determined, including: screening a target user set based on the high-threat public intelligence set; determining a target group set using the target user set; and extracting the semi-private domain content of the target group set using a large model.

[0090] Here, a target user refers to a user associated with high-threat public intelligence. These users may have directly or indirectly participated in high-threat activities, or have an important association with high-threat intelligence content. Through these users, more potential threat information and related groups can be further explored.

[0091] In the embodiments of the present disclosure, the process of screening a target user set based on the high-threat public intelligence set can be achieved by extracting all involved user account information from the high-threat intelligence set, or by using graph algorithms or social network analysis methods to identify users directly or indirectly associated with the high-threat intelligence.

[0092] Here, a target group refers to communities, chat groups, forum sections, etc. where target users gather, and can be semi-private private groups that require joining permissions.

[0093] In the embodiments of the present disclosure, in the process of determining the target group set, for each user in the target user set, the group information they participated in can be extracted first. Exemplarily, the group information a user participated in can be queried from the social media application programming interface (API), or a crawler tool can be used to locate the active groups of the user on a specific platform. Subsequently, the groups can be screened according to indicators such as group activity and the number of associated users. Exemplarily, if the number of users related to high-threat intelligence in a group is larger, the priority is higher; at the same time, if the discussion content in the group is highly relevant to the tags of high-threat intelligence, the priority is higher.

[0094] In the embodiments of the present disclosure, in the process of using a large model to extract semi-private content of a target group set, data such as chat records and post content can be first obtained from the target group set, and then the obtained data is input into the large model. The large model analyzes the captured content to extract semi-private content related to high-threat intelligence. Exemplarily, analysis can be performed through keyword extraction, topic classification, user behavior analysis, etc.

[0095] In this way, by screening target users and groups and locating users and communities highly related to high-threat intelligence, potential attack behaviors and organizational activities can be deeply explored. Using a large model to extract semi-private content expands the analysis of high-threat intelligence from the public domain to the semi-private domain, improving the depth and coverage of intelligence. Automatically screening target users and groups significantly reduces the manual intervention cost of intelligence analysis, especially when dealing with large-scale intelligence data. The introduction of the large model realizes the automatic parsing and classification of semi-private content, avoiding the cumbersome process of manual sorting and screening and greatly improving the efficiency. Through the screening of target users and target groups, the source and dissemination path of threat intelligence can be clearly traced.

[0096] In some embodiments, based on a high-threat public intelligence set, a target user set is screened, including: obtaining the user account in the corresponding metadata according to any high-threat intelligence in the high-threat public intelligence set; traversing each high-threat intelligence in the high-threat public intelligence set to generate a high-threat user set; sorting the high-threat user set according to the hit times, and extracting the user accounts with hit times greater than the preset user hit times threshold to generate a target user set.

[0097] In the embodiments of the present disclosure, the acquisition methods of user accounts include but are not limited to the following legal ways: 1) Account information actively disclosed by users; 2) Account data legally and compliantly obtained based on explicit authorization by users or in accordance with relevant privacy policies and laws and regulations.

[0098] In the embodiments of the present disclosure, in the process of obtaining the user account in the corresponding metadata, explicit user identification information can be extracted from the metadata of high-threat intelligence, or potential account information can be extracted from the content of high-threat intelligence through natural language processing technology.

[0099] In the embodiments of the present disclosure, in the process of generating a high-threat user set, for each piece of intelligence in the high-threat public intelligence set, the user account in the metadata can be extracted respectively, and then the extracted user accounts are added to a unified set to form a high-threat user set.

[0100] In the embodiments of the present disclosure, in the process of generating a target user set, the user accounts can be sorted first in descending order of the hit times. Specifically, the user accounts and their hit times can be used as key-value pairs first, and then the key-value pairs of the high-threat user set can be sorted in descending order according to the corresponding values. Subsequently, a threshold for the user hit times can be preset to screen high-correlation user accounts. Finally, the sorted high-threat user set can be traversed to extract the user accounts with hit times greater than or equal to the threshold, thereby generating the target user set.

[0101] In this way, by extracting and analyzing the user account information in high-threat intelligence, users highly relevant to threat events can be accurately located, improving the accuracy of intelligence analysis. By sorting the hit times of user accounts, the most relevant high-threat users can be intuitively processed first, optimizing resource allocation. By adjusting the threshold for user hit times, different business requirements can be flexibly adapted.

[0102] In some embodiments, using the target user set to determine the target group set includes: for each target user in the target user set, extracting the information of the associated groups in their social network and identifying the associated groups based on social network analysis methods; traversing the information of the associated groups of all target users, counting the hit times of each group in the target user set, and generating a set of suspicious groups; forming the target group set according to the groups in the set of suspicious groups whose hit times exceed the preset threshold.

[0103] Here, social network analysis technology is a set of theories and methods for studying and analyzing the nodes in a network and the relationships between them, and can be used to reveal the structure, roles, influence of nodes in the network, as well as the interaction and dissemination patterns between groups.

[0104] In the embodiments of the present disclosure, social network analysis technology is used to identify the relationships between target users and other users or groups. In particular, these relationships can be direct relationships such as the groups where the users are located or indirect relationships such as being associated through common friends or groups where they have spoken together. Exemplarily, the group information that the target user participates in can be directly extracted from the records of the target user, or the behavior and social activities of the target user can be located, and the groups to which their interaction objects belong can be analyzed. Also, a relationship network graph between users and groups can be constructed, and the groups associated with the target user can be analyzed through the topological structure.

[0105] In the embodiments of the present disclosure, for each user in the target user set, social network analysis technology is used to identify the relationships between the target user and other users or groups, and the information of the groups associated with all traversed users is uniformly summarized to obtain a set of suspicious groups.

[0106] In the embodiments of the present disclosure, the occurrence times of each group in the target user set can be counted, and then all the groups in the suspicious group set can be sorted in descending order according to the hit times, so that the key groups associated with more target users can be processed preferentially. Then, a hit time threshold for suspicious groups can be preset to screen out the important groups closely associated with multiple target users. By traversing the sorted suspicious group set, the groups with hit times greater than or equal to the threshold are extracted, and then the target group set is generated.

[0107] In this way, through the analysis of the association between target users and groups, the high-risk groups closely related to threat activities can be accurately screened out, and the coverage of intelligence analysis can be improved. Through sorting by hit times and threshold screening, the priority groups highly associated with multiple target users can be focused on, the interference of non-key groups can be reduced, so that the intelligence analysis resources can be concentrated on the most important targets, and the analysis process can be optimized. Through automatic extraction, traversal, and statistical group relationships, the rapid screening of a large number of intelligence groups is realized, which is suitable for processing large-scale intelligence data.

[0108] In some embodiments, according to the semi-private content, a semi-private intelligence knowledge graph is created, including: performing data processing on the semi-private content to obtain semi-private intelligence; identifying the semi-private intelligence to generate semi-private entities; using the semi-private entities to construct a semi-private intelligence database; and constructing a semi-private intelligence knowledge graph based on the semi-private intelligence database.

[0109] Here, the semi-private entity refers to an information unit extracted from the semi-private intelligence that has a clear semantics or represents an actual object. These entities are the basic building blocks of the intelligence knowledge graph, usually used to represent certain key objects or concepts and are connected by relationships. Exemplarily, the semi-private entities can include users, groups, IP addresses, domain names, files, events, etc., depending on the theme and source of the intelligence.

[0110] In the embodiments of the present disclosure, the process of performing data processing on the semi-private content can achieve denoising processing by cleaning the information irrelevant to the intelligence, filtering out the content with incorrect formats or that cannot be parsed, and then converting the unstructured data into structured data, thereby obtaining the semi-private intelligence.

[0111] In the embodiments of the present disclosure, the process of generating semi-private entities can first extract named entities from the semi-private intelligence using natural language processing techniques, then classify the extracted entities into different types, and finally standardize the entities, thereby generating semi-private entities.

[0112] In the embodiments of the present disclosure, in the process of constructing a semi-private domain intelligence database, the extracted semi-private domain entities can be stored in the intelligence database, and a reasonable structure can be designed for subsequent querying and analysis. Exemplarily, a relational database can be used for tabular storage. Subsequently, the entities, their attributes, and the relationships between entities are stored in the semi-private domain intelligence database.

[0113] In the embodiments of the present disclosure, in the process of constructing a semi-private domain intelligence knowledge graph, the semi-private domain entities can first be used as nodes in the graph for node construction. Subsequently, according to the relationship data in the database, edges in the graph are added to achieve relationship construction. Finally, attribute information is added to each node and each edge respectively.

[0114] In this way, by extracting entities and constructing a database, they can be transformed into structured and queryable data for easy and quick analysis. Through the construction of the knowledge graph, intelligence is presented in a visual manner, which helps to discover potential threats and trace the source. The semi-private domain intelligence knowledge graph can track the source of threats, analyze the threat propagation path, and predict potential threat nodes or groups through graph analysis techniques.

[0115] Figure 2 The flowchart showing the generation of semi-private domain content is as Figure 2 shown, and this process includes:

[0116] S201: Construct a thesaurus. Specifically, a thesaurus for searching black production-related vocabulary can be constructed first. Exemplarily, the topics can include the following: customization, accounts, network disk cracking, map location tampering, etc. Subsequently, a keyword library can be constructed. The formulation of the keyword library needs to be carried out according to different topics, and several keyword libraries may need to be generated for a certain topic. Exemplarily, combinations can be made from capitalization and lowercase combinations, homophones, misspelled words, slang, abbreviations, long-tail keywords, etc. At the same time, since it is oriented to overseas websites, corresponding English terms and slang need to be added. The keywords are also set with relevant priorities according to their derivation methods for convenient subsequent viewing.

[0117] S202: Extract the original public intelligence. Specifically, relevant posts or users can be searched on key platforms based on the vocabulary in the thesaurus first. Then, different crawler rules are formulated for different platforms. In particular, for some platforms, different delay and time strategies may need to be created for the crawler strategy. Finally, the crawled posts can be saved in several units of "tag: platform: link: title + time: user: content: comment user + comment + time".

[0118] S203: Filter according to different categories based on business needs. The amount of data from multiple platforms may be relatively large, and can be further sorted and viewed from the perspectives of platform priority, keyword relevance, timeliness, and negative keywords. Subsequently, the filtered users can be saved, and the saved format can be "label: user nickname: user account: IP address: physical location". In particular, for information that cannot be obtained, it can be stored in the form of default values or empty values. Then, different users can also be prioritized according to various dimensions such as the number of fans, number of personal posts, and level of activity. In particular, since the same user may comment on different posts in a short period of time, appropriate deduplication can be performed based on the user account.

[0119] S204: Use the selected users to determine the target group, and then crawl the information in the group with the help of crawler technology. If you can enter the relevant group, try to interact and trade with the black market users to further mine more valuable information.

[0120] S205: Use the big model to extract, organize and understand the crawled natural language information, sort out the information, and generate semi-private content.

[0121] In this way, with the help of the constructed multi-level vocabulary, highly relevant information can be extracted from massive data. Through information crawling, user interaction and deep mining, it is not limited to the analysis of surface data, but also can deeply explore the hidden relationships and behavior patterns of users, so that a more detailed black industry user portrait can be constructed, effectively assisting in the formulation of accurate risk control strategies. The use of large model language understanding technology has greatly improved the efficiency of information sorting and reduced the time and cost of manual participation.

[0122] The present disclosure provides a semi-private domain intelligence information acquisition device, such as Figure 3 As shown, the device may include: a vocabulary generation module 301, which is used to generate a keyword library based on the intelligence theme using a pre-trained keyword model; an intelligence collection module 302, which is used to extract original public intelligence from the data source based on the keyword library; a threat determination module 303, which is used to calculate the threat value of the original public intelligence and obtain a high-threat public intelligence set; a content acquisition module 304, which is used to determine the semi-private domain content based on the high-threat public intelligence set; a graph creation module 305, which is used to create a semi-private domain intelligence knowledge graph based on the semi-private domain content.

[0123] In some embodiments, the thesaurus generation module 301 includes: a theme determination sub-module for analyzing historical data to determine the intelligence theme; a thematic thesaurus sub-module for using a keyword model to extract the theme words in the intelligence theme and establish a thematic thesaurus; and a keyword library sub-module for using the keyword model to diverge the semantic meaning of any theme word in the thematic thesaurus to generate a keyword library.

[0124] In some embodiments, the data source at least includes: surface web information, dark web information, and instant messaging tool application data.

[0125] In some embodiments, the intelligence collection module 302 includes: an intelligence extraction sub-module for extracting target public content from the data source based on the keyword library; and an intelligence marking sub-module for marking the target public content according to a preset tag type to obtain the original public intelligence.

[0126] In some embodiments, the intelligence extraction sub-module is used for: performing keyword matching on the data source information by using natural language processing technology to obtain the first public content; performing semantic analysis on the data source information by using a deep learning model to obtain the second public content; and performing data processing on the first public content and the second public content to obtain the target public content.

[0127] In some embodiments, the intelligence extraction sub-module is further used for: merging and deduplicating the first public content and the second public content to obtain the public content to be processed; and performing data cleaning on the public content to be processed to obtain the target public content.

[0128] In some embodiments, the intelligence marking sub-module is used for: obtaining the basic information of the target public content according to the preset tag content; the basic information at least includes: the platform where the target public content is located, the link of the target public content, the time of the target public content, and the user account related to the target public content; generating metadata of the target public content according to the basic information; and merging the target public content and the metadata to generate the original public intelligence.

[0129] In some embodiments, the threat determination module 303 includes: a threat calculation sub-module for calculating the threat value of the original public intelligence, sorting the threat values of the original public intelligence to obtain an original public intelligence sequence; and a threat set sub-module for extracting the original public intelligence with a threat value greater than a preset threat threshold in the original public intelligence sequence to generate a high-threat public intelligence set.

[0130] In some embodiments, the threat calculation sub-module is configured to: calculate a platform threat value according to the original open intelligence and a preset platform threat coefficient; calculate a relevance threat value according to the original open intelligence and a preset relevance threat coefficient; calculate a timeliness threat value according to the original open intelligence and a preset timeliness threat coefficient; calculate a negative keyword threat value according to the original open intelligence and a preset negative keyword threat coefficient; and calculate the threat value of the original open intelligence by using the platform threat value, the relevance threat value, the timeliness threat value, and the negative keyword threat value according to a preset threat ratio.

[0131] In some embodiments, the content acquisition module 304 includes: a user screening sub-module, configured to screen a target user set based on a high-threat open intelligence set; a group determination sub-module, configured to determine a target group set by using the target user set; and a content extraction sub-module, configured to extract semi-private domain content of the target group set by using a large model.

[0132] In some embodiments, the user screening sub-module is configured to: obtain a user account in the corresponding metadata according to any high-threat intelligence in the high-threat open intelligence set; traverse each high-threat intelligence in the high-threat open intelligence set to generate a high-threat user set; sort the high-threat user set according to the hit times, and extract user accounts with hit times greater than a preset user hit times threshold to generate a target user set.

[0133] In some embodiments, the group determination sub-module is configured to: for each target user in the target user set, extract the social network associated group information thereof, and identify the associated groups based on a social network analysis method; traverse the associated group information of all target users, count the hit times of each group in the target user set to generate a suspicious group set; sort the suspicious group set according to the hit times, and extract groups with hit times exceeding a preset threshold to form a target group set.

[0134] In some embodiments, the graph creation module 305 includes: an intelligence generation sub-module, configured to perform data processing on the semi-private domain content to obtain semi-private domain intelligence; an entity generation sub-module, configured to identify the semi-private domain intelligence to generate semi-private domain entities; an entity construction sub-module, configured to construct a semi-private domain intelligence database by using the semi-private domain entities; and a graph construction sub-module, configured to construct a semi-private domain intelligence knowledge graph based on the semi-private domain intelligence database. For the specific functions and example descriptions of the modules and sub-modules of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.

[0135] The semi-private domain intelligence information acquisition device in the embodiments of the present disclosure can significantly improve the comprehensiveness and accuracy of the semi-private domain intelligence knowledge graph.

[0136] An embodiment of the present disclosure provides a schematic diagram of a scenario for a method of obtaining semi-private domain intelligence information, as Figure 4 shown.

[0137] As mentioned above, the method for obtaining semi-private domain intelligence information provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.

[0138] Specifically, the electronic device can specifically perform the following operations:

[0139] Generate a keyword library based on the intelligence theme using a pre-trained keyword model;

[0140] Extract the original publicly available intelligence from the data source based on the keyword library;

[0141] Calculate the threat value of the original publicly available intelligence to obtain a set of publicly available intelligence with high threats;

[0142] Determine the semi-private domain content based on the set of publicly available intelligence with high threats;

[0143] Create a semi-private domain intelligence knowledge graph according to the semi-private domain content.

[0144] It should be understood that Figure 4 the shown scenario diagram is merely illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 4 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0145] In the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0146] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0147] Figure 5FIG. 0 shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0148] As Figure 5 shown, the device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0149] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0150] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various dedicated Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the semi-private domain intelligence information acquisition method. For example, in some embodiments, the semi-private domain intelligence information acquisition method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the semi-private domain intelligence information acquisition method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the semi-private domain intelligence information acquisition method by any other suitable means (e.g., by means of firmware).

[0151] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application-Specific Standard Products (ASSPs), System on Chip (SOC) systems, Complex Programmable Logic Devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0152] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0153] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0154] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0155] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.

[0156] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0157] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0158] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for obtaining semi-private domain intelligence information, comprising: Generating a keyword library based on the intelligence theme using a pre-trained keyword model; Extracting original public intelligence from a data source based on the keyword library; Calculating the threat value of the original public intelligence to obtain a high-threat public intelligence set; Determining semi-private domain content based on the high-threat public intelligence set; Creating a semi-private domain intelligence knowledge graph according to the semi-private domain content.

2. The method according to claim 1, wherein, The generating a keyword library based on the intelligence theme using a pre-trained keyword model includes: Analyzing historical data to determine the intelligence theme; Using the keyword model to extract theme words in the intelligence theme and establishing a theme word library; Using the keyword model to perform semantic divergence on any theme word in the theme word library to generate the keyword library.

3. The method according to claim 1, wherein, The data source at least includes: clear web information, dark web information, and instant messaging tool application data.

4. The method according to claim 1, wherein The extracting original public intelligence from a data source based on the keyword library includes: Extracting target public content from the data source based on the keyword library; Marking the target public content according to a preset tag type to obtain the original public intelligence.

5. The method according to claim 4, wherein The extracting target public content from the data source based on the keyword library includes: Performing keyword matching on the data source information using natural language processing technology to obtain the first public content; Performing semantic analysis on the data source information using a deep learning model to obtain the second public content; Performing data processing on the first public content and the second public content to obtain the target public content.

6. The method according to claim 5, wherein The performing data processing on the first public content and the second public content to obtain the target public content includes: Merging and removing duplicates from the first public content and the second public content to obtain the public content to be processed; Performing data cleaning on the public content to be processed to obtain the target public content.

7. The method according to claim 4, wherein The marking the target public content according to a preset tag type to obtain the original public intelligence includes: Obtaining the basic information of the target public content according to the preset tag content; the basic information at least includes: the platform where the target public content is located, the link of the target public content, the time of the target public content, and the user account related to the target public content; Generating metadata of the target public content according to the basic information; Merging the target public content and the metadata to generate the original public intelligence.

8. The method according to claim 1, wherein The calculating the threat value of the original public intelligence to obtain a high-threat public intelligence set includes: Calculating the threat value of the original public intelligence, sorting the threat values of the original public intelligence to obtain an original public intelligence sequence; Extracting the original public intelligence with a threat value greater than a preset threat threshold in the original public intelligence sequence to generate the high-threat public intelligence set.

9. The method according to claim 8, wherein The calculating the threat value of the original public intelligence, sorting the threat values of the original public intelligence to obtain an original public intelligence sequence includes: Calculate the platform threat value according to the original public intelligence and the preset platform threat coefficient; Calculate the relevance threat value according to the original public intelligence and the preset relevance threat coefficient; Calculate the timeliness threat value according to the original public intelligence and the preset timeliness threat coefficient; Calculate the negative keyword threat value according to the original public intelligence and the preset negative keyword threat coefficient; Calculate the threat value of the original public intelligence according to the preset threat ratio, using the platform threat value, the relevance threat value, the timeliness threat value, and the negative keyword threat value.

10. The method according to claim 1, wherein Determine semi-private content based on the high-threat public intelligence set, including: Screen the high-threat public intelligence set to obtain a target user set; Use the target user set to determine a target group set; Use a large model to extract the semi-private content of the target group set.

11. The method according to claim 10, wherein, The screening to obtain a target user set based on the high-threat public intelligence set includes: According to any high-threat intelligence in the high-threat public intelligence set, obtain the user account in the corresponding metadata; Traverse each high-threat intelligence in the high-threat public intelligence set to generate a high-threat user set; Sort the high-threat user set according to the hit times, and extract the user accounts with hit times greater than the preset user hit times threshold to generate the target user set.

12. The method according to claim 10, wherein, The using the target user set to determine a target group set includes: For each target user in the target user set, extract their social network associated group information and identify the associated groups based on social network analysis methods; Traverse the associated group information of all target users, count the hit times of each group in the target user set, and generate a suspicious group set; Form a target group set according to the groups in the suspicious group set whose hit times exceed the preset threshold.

13. The method according to claim 1, wherein, Create a semi-private intelligence knowledge graph according to the semi-private content, including: Perform data processing on the semi-private content to obtain semi-private intelligence; Identify the semi-private intelligence to generate semi-private entities; Use the semi-private entities to construct a semi-private intelligence database; Based on the semi-private intelligence database, construct the semi-private intelligence knowledge graph.

14. A semi-private intelligence information acquisition device, including: A thesaurus generation module for generating a keyword library based on the intelligence theme using a pre-trained keyword model; An intelligence collection module for extracting original public intelligence from a data source based on the keyword library; A threat determination module for calculating the threat value of the original public intelligence to obtain a high-threat public intelligence set; A content acquisition module for determining semi-private content based on the high-threat public intelligence set; A graph creation module for creating a semi-private intelligence knowledge graph according to the semi-private content.

15. An electronic device, including: At least one processor; And A memory communicatively connected to at least one processor; wherein, The memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable the at least one processor to execute the method of any one of claims 1-13.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are for causing a computer to execute the method of any one of claims 1-13.

17. A computer program product, comprising a computer program stored on a storage medium, the computer program implementing the method of any one of claims 1-13 when executed by a processor.