A multi-source heterogeneous threat intelligence data automatic processing method and system
Patent Information
- Application Number
- CN202610726480.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-09-11
AI Technical Summary
基于文本内容指纹对原始文本对象进行自适应解析输出中间文本对象;
本发明提出的多源异构威胁情报数据自动处理方法,与现有技术相比,形成了从原始异构情报到可直接服务于下游模块多模态智能分析的高质量结构化数据的端到端能力,本发明的创新具体体现在:配置表元数据驱动的"零代码"多源接入扩展机制,无需修改核心代码即可接入新数据源;基于内容指纹探测的自适应格式识别与解析器选择,解决格式标记错误问题;多层融合的安全实体抽取管线,兼顾覆盖率与精确率;跨文档实体归一化的安全实体别名映射机制;"一次处理、多模态复用"的统一情报结构模型UTIDM。通过统一接入、统一解析、统一抽取和统一表达,提高了威胁情报数据的采集效率、结构化质量和下游适配能力,利用自动化多维探查提升情报获取时效性,为后续智能问答、攻击链分析和知识图谱构建提供高质量基础数据,实现一站式自动化流水线式执行。
Smart Images

Figure CN122734962A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network security technology, specifically relating to an automatic processing method and system for multi-source heterogeneous threat intelligence data. Background Technology
[0002] With the continued growth of security threats such as APT attacks, ransomware attacks, supply chain delivery attacks, and mass exploitation of vulnerabilities, cyber threat intelligence has become a crucial supporting information for cybersecurity operations, threat hunting, incident response, and risk warning. Threat intelligence includes not only IOC-type metrics but also higher-level security knowledge such as the background of the attacking organization, characteristics of malicious samples, attack stages, delivery methods, vulnerability exploitation paths, target industries, and the scope of victims. Currently, actual threat intelligence data mainly comes from the following sources: open-source threat intelligence platforms and vulnerability platforms, such as OTX, MITRE ATT&CK, and NVD; analysis reports, blogs, and security bulletins published on the official websites of security vendors; vulnerability disclosure and patch notes pages; security-related content in technical communities, forums, and code repositories; and historical reports and local documents accumulated within enterprises or laboratories.
[0003] These data exhibit distinct characteristics of being "multi-source, heterogeneous, dynamic, and unstructured." On the one hand, the intelligence updates rapidly, requiring the system to possess continuous collection and incremental update capabilities; on the other hand, the intelligence is expressed in complex ways, with the same APT organization, malware, or vulnerability potentially being expressed using different names or formats in different documents, which places higher demands on subsequent automated processing.
[0004] Existing technical solution one: Patent document CN118174971A provides "a method and system for governing multi-source heterogeneous data of network threats," which obtains structured intelligence through standard API interfaces and stores the returned results in a database for querying. Based on the data exploration results, a data standard for multi-source heterogeneous network threat data is edited; according to the data standard, a custom task is configured to standardize the processing of multi-source heterogeneous network threat data in the data warehouse, completing data cleaning, data association, and data backfilling, and finally storing the processed data in the corresponding original intelligence database. This method enhances the governance capability of multi-source heterogeneous data. Existing technical solution two: Patent document CN121418178A provides "an intelligent analysis method for threat intelligence based on large models and knowledge graphs," involving automated cleaning, structured extraction and consistent storage, and multimodal intelligence retrieval and association methods for multi-source heterogeneous data based on a hybrid retrieval framework and evidence backtracking design. It proposes a feasible, lightweight threat scoring and grading mechanism, as well as a structured prompt and interpretable output scheme with evidence alignment, enabling the conclusion to be traced back to the original text and the chain of events, and generating disposal recommendations for assets and business operations.
[0005] However, the drawbacks of the existing technical solution one are: over-reliance on standard interfaces and existing structured intelligence sources, resulting in insufficient processing capabilities for unstructured security reports, blogs, forum posts, and local historical documents, and limited intelligence coverage. Meanwhile, network threat data, such as alert logs, malware samples, and C2 domains, are highly time-sensitive. Using methods typically involving batch writing and temporary storage in analysis data warehouses, even with subsequent cleaning and correlation, can introduce delays of several minutes or even hours, making it unsuitable for scenarios requiring real-time or near-real-time responses. The drawbacks of the existing technical solution two are: significant differences in semantics, format, credibility, and timeliness between different threat intelligence sources (such as open-source, commercial, and internal logs). Automated cleaning and structured extraction may lead to misalignment or information loss, especially for unstructured or adversarial samples. The quality of the knowledge graph highly depends on the initial ontology design, the accuracy of entity links, and the continuous update mechanism. For novel threats (zero-day vulnerabilities, unknown attack patterns), the graph lacks corresponding nodes and relationships; in this case, the contribution of the "graph retrieval" channel is limited, and the system may degenerate into pure vector retrieval. Summary of the Invention
[0006] To address the aforementioned problems in the existing technology, this invention provides an automatic processing method and system for multi-source heterogeneous threat intelligence data. In a first aspect, embodiments of the present invention provide an automatic processing method for multi-source heterogeneous threat intelligence data, the method comprising: Establish a threat intelligence source configuration table corresponding to different types of threat intelligence sources, and collect the corresponding raw text objects using the collection methods corresponding to different types of threat intelligence sources; The original text object is adaptively parsed and an intermediate text object is output based on the text content fingerprint. Perform intelligence content cleaning and standardization preprocessing on the intermediate text object; Security elements are identified and extracted from cleaned and preprocessed intermediate text objects through a multi-layered hybrid extraction pipeline. The extracted security elements are mapped to a pre-defined unified threat intelligence data model.
[0007] In one embodiment of the present invention, the threat intelligence source configuration table includes the access method, update cycle, collection rules, authentication information, file type and priority of different types of threat intelligence sources.
[0008] In one embodiment of the present invention, different types of threat intelligence sources include at least open intelligence platform interfaces, vulnerability announcements and security bulletin pages, security vendor blogs and report pages, local historical intelligence file directories, user-uploaded data, and other semi-structured security content sources.
[0009] In one embodiment of the present invention, the text content fingerprint includes a text type field and text content features in the original text object; correspondingly, adaptive parsing of the original text object to output an intermediate text object includes: Based on the text type field, the original text object is parsed for the first time; Based on the text content features, the original text object after the first parsing is parsed a second time to output an intermediate text object.
[0010] In one embodiment of the present invention, the intermediate text object undergoes intelligence content cleaning and standardization preprocessing, including: Remove webpage navigation, copyright notices, advertising content, irrelevant footnotes, abnormal line breaks, and duplicate paragraphs; Standardize time format, URL format, vulnerability number format, and text encoding; Correct common OCR errors or formatting issues; Normalize and deduplicate duplicate IOCs in the same type of threat intelligence sources; The different writing formats of APT organization aliases, malware aliases, and vulnerability numbers are normalized, and cross-document normalization mapping is performed with reference to the built-in security entity alias dictionary.
[0011] In one embodiment of the present invention, security elements are identified and extracted from cleaned and preprocessed intermediate text objects through a multi-layer hybrid extraction pipeline, including: A three-layer hybrid extraction pipeline consisting of a rule layer, a knowledge layer, and a model layer is constructed. Based on the extraction order of the rule layer, knowledge layer, and model layer in the three-layer hybrid extraction pipeline, the intermediate text objects after cleaning and preprocessing are identified and extracted. The extraction results are then normalized and mapped across documents with reference to the built-in secure entity alias dictionary to obtain the security elements.
[0012] In one embodiment of the present invention, the unified threat intelligence data model includes an identity group, a time information group, a content description group, a security element group, and a quality control group, so as to achieve field mapping and unified structured expression through five types of field groups.
[0013] In one embodiment of the present invention, the identity identification group includes at least an intelligence identifier, an intelligence source, and a source category; The time information group should include at least the release time and the collection time; The content description group should include at least the intelligence type, summary information, and text excerpt; A security element group should include at least an entity list, an IOC list, a TTP list, and a target object list. The quality control group should include at least a credibility or confidence score and the location of the original evidence.
[0014] In one embodiment of the present invention, the automatic processing method for multi-source heterogeneous threat intelligence data further includes: Based on a pre-defined unified threat intelligence data model, unified storage management is implemented for different downstream modules.
[0015] Secondly, embodiments of the present invention provide an automatic processing system for multi-source heterogeneous threat intelligence data, the automatic processing system for multi-source heterogeneous threat intelligence data comprising: The data collection module is used to create a threat intelligence source configuration table for different types of threat intelligence sources, and to collect the corresponding raw text objects using the data collection methods corresponding to different types of threat intelligence sources. The parsing module is used to adaptively parse the original text object based on the text content fingerprint and output an intermediate text object; The preprocessing module is used to perform intelligence content cleaning and standardization preprocessing on intermediate text objects; The identification and extraction module is used to identify and extract security elements from cleaned and preprocessed intermediate text objects through a multi-layer hybrid extraction pipeline. The mapping module is used to map the extracted security elements to a pre-defined unified threat intelligence data model.
[0016] The beneficial effects of this invention are: The proposed automatic processing method for multi-source heterogeneous threat intelligence data in this invention, compared with existing technologies, forms an end-to-end capability from raw heterogeneous intelligence to high-quality structured data that can directly serve downstream modules for multimodal intelligent analysis. The innovations of this invention are specifically reflected in: a configuration table metadata-driven "zero-code" multi-source access extension mechanism, allowing access to new data sources without modifying core code; adaptive format recognition and parser selection based on content fingerprint detection, solving the problem of format marking errors; a multi-layered fusion security entity extraction pipeline, balancing coverage and accuracy; a cross-document entity normalization security entity alias mapping mechanism; and a unified intelligence structure model (UTIDM) for "one-time processing and multimodal reuse." Through unified access, unified parsing, unified extraction, and unified expression, the efficiency of threat intelligence data collection, the quality of structured data, and downstream adaptability are improved. Automated multi-dimensional exploration enhances the timeliness of intelligence acquisition, providing high-quality foundational data for subsequent intelligent question answering, attack chain analysis, and knowledge graph construction, achieving one-stop automated pipeline execution.
[0017] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating an automatic processing method for multi-source heterogeneous threat intelligence data provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the adaptive parsing framework provided in an embodiment of the present invention; Figure 3 This is a flowchart illustrating another method for automatically processing multi-source heterogeneous threat intelligence data provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an automatic processing system for multi-source heterogeneous threat intelligence data provided in an embodiment of the present invention. Detailed Implementation
[0019] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0020] Firstly, please see Figure 1 This invention provides an automatic processing method for multi-source heterogeneous threat intelligence data, which includes: S10. Establish a threat intelligence source configuration table corresponding to different types of threat intelligence sources, and collect the corresponding raw text objects using the collection methods corresponding to different types of threat intelligence sources.
[0021] In this embodiment of the invention, the threat intelligence source configuration table includes the access method, update cycle, collection rules, authentication information, file type, and priority of different types of threat intelligence sources. Different types of threat intelligence sources include at least open intelligence platform interfaces, vulnerability announcements and security bulletin pages, security vendor blogs and report pages, local historical intelligence file directories, user-uploaded materials, and other semi-structured security content sources. More specifically: Each threat intelligence source record in the threat intelligence source configuration table of this invention includes at least the following metadata fields: intelligence source ID, access method type (API / crawl / directory / upload), update cycle (real-time / hour / day), authentication credential field name, file format constraints, priority weight, etc., to support subsequent unified scheduling and management.
[0022] This invention utilizes open intelligence platform interfaces, such as OTX and VirusTotal API, which employ API key authentication for periodic fetching, supporting both full initialization and incremental update modes. Vulnerability announcements and security bulletin pages, such as NVD and CNVD, are crawled and subscribed to via RSS feeds. Security vendor blogs and report pages, such as those of QiAnXin, 360, and Microstep Online, are crawled periodically. Local historical intelligence file directories are loaded in batches via directory scanning. User-uploaded data, such as single or batch files uploaded through the system's front-end interface, is automatically placed into a processing queue upon receipt, supporting common formats like PDF, DOCX, TXT, and XLSX, triggering an immediate parsing process without waiting for scheduled cycles. Other semi-structured security content sources include: PoC vulnerability exploit code repositories on GitHub, whose README and issues contain semi-structured information such as vulnerability descriptions, affected versions, and CVE numbers, obtained through the GitHub API or web crawling; and malicious sample analysis summaries appearing on code-sharing platforms like Pastebin, whose text contains key information such as IOC, sample hashes, and C2 addresses, collected through keyword monitoring and page snapshots.
[0023] Different collection methods are employed for different types of intelligence sources: for API-based intelligence sources, a scheduled method is used to call APIs and retrieve incremental data; for web-based intelligence sources, the page content, title, publication time, and tags are extracted according to preset templates or text positioning rules; for local file-based intelligence sources, documents to be processed are loaded through directory polling or batch import; for user-uploaded intelligence sources, the upload queue is monitored and the processing flow is triggered immediately; and so on. After collection, for intelligence sources that have already been collected, the system automatically extracts the summary hash of the current content during each exploration and compares it with the locally stored records. Download and subsequent processing are only triggered when the content changes, avoiding duplicate collection. The system also automatically extracts external reference links appearing in the collected documents and determines whether they belong to trusted intelligence sources according to preset domain whitelist rules. Those that meet the criteria are automatically added to the collection queue without manual intervention, realizing the autonomous expansion of intelligence sources. After collection, a unique original record identifier is generated for each piece of original intelligence, along with metadata such as collection time, source URL, original file name, and source category, ultimately forming the original text object.
[0024] Existing technologies for accessing threat intelligence data sources generally adopt a "hard-coded + individual adaptation" approach. This means that a separate collection module is developed for each platform, and adding a new data source requires code modification and redeployment, resulting in high expansion costs. There is no readily available solution in the industry that integrates heterogeneous access methods such as API-based, crawling-based, file-based, and user-uploaded methods into a unified configurable framework. This invention designs a threat intelligence source configuration table-driven approach that uniformly describes the access parameters of all intelligence sources. The scheduling engine dynamically selects the collection strategy based on the configuration table. Adding a new intelligence source only requires filling in the configuration items to access it, without modifying the core code. This achieves "zero-code expansion" capability for the intelligence collection infrastructure, which is unprecedented in existing threat intelligence platform technical solutions.
[0025] S20. Adaptively parse the original text object based on the text content fingerprint and output an intermediate text object.
[0026] In this embodiment of the invention, the text content fingerprint includes a text type field and text content features in the original text object; correspondingly, adaptive parsing of the original text object to output an intermediate text object includes: performing a first parsing of the original text object based on the text type field; and performing a second parsing of the original text object after the first parsing based on the text content features to output an intermediate text object. More specifically: This invention automatically selects and configures the appropriate parser based on text type fields and content characteristics: For PDF files, it extracts the main text, page numbers, titles, paragraphs, and tables, prioritizing title levels and main text blocks using rule-based layout analysis, and additionally calling the OCR module for scanned PDFs; for HTML web pages, it extracts the main text area, title, publication time, and tags, using predefined text positioning rule sets to identify the main content area and filter out noisy areas such as navigation bars and sidebars; for structured files such as JSON, CSV, and STIX, it extracts field content, parsing according to field path mapping rules; for plain text files such as TXT and Markdown, it directly loads and performs paragraph segmentation; for complex layouts such as two-column PDFs and nested tables, it can further perform format normalization and paragraph reorganization. After parsing, a standard intermediate text object is uniformly generated, including the main text content, title information, source information, and original metadata. All fields will be used as input for subsequent cleaning and preprocessing. Figure 2 This illustrates the process of automatically selecting the appropriate parser for formats such as PDF, HTML, JSON, TXT, and STIX.
[0027] Existing general-purpose text parsing tools such as Apache Tika and TeXtract support multiple formats, but their format identification relies on file extensions or MIME types. They cannot correctly process security documents with missing extensions or incorrect formatting markers, such as security reports uploaded as TXT files disguised as PDFs. This invention introduces a content fingerprinting mechanism: first, it detects the document's header to determine its true format; then, it combines this with document content features such as text density, table proportions, and font diversity to fine-tune the parsing strategy. Furthermore, it designs a special processing branch to address common formatting issues in security reports, such as two-column PDFs, nested tables, and OCR scans. This two-layer parsing strategy, combining format recognition and content awareness, is a specialized design for the characteristics of security intelligence documents, transcending the scope of general document parsing approaches.
[0028] S30. Perform intelligence content cleaning and standardization preprocessing on the intermediate text object.
[0029] In embodiments of this invention, the intermediate text object undergoes intelligence content cleaning and standardization preprocessing, including: removing webpage navigation, copyright notices, advertising content, irrelevant footnotes, abnormal line breaks, and duplicate paragraphs; standardizing time format, URL format, vulnerability number format, and text encoding; correcting common OCR errors or formatting anomalies; standardizing and deduplicating duplicate IOCs in the same type of threat intelligence source; and normalizing different writing formats of APT organization aliases, malware aliases, and vulnerability numbers, using a cross-document normalization mapping with reference to the built-in security entity alias dictionary. More specifically: After completing the intelligence content cleaning and standardization preprocessing, to ensure the quality of subsequent extraction, it is necessary to perform special security-oriented cleaning processing on the intermediate text objects (i.e., the structured intermediates containing the main text, title information, source information, and original metadata generated after parsing various heterogeneous documents). This includes: removing webpage navigation, copyright notices, advertising content, irrelevant footnotes, abnormal line breaks, and duplicate paragraphs, such as through batch matching and filtering using a preset noise pattern rule set; and unifying the time format (converting various expressions such as "Jan 2025", "2025 / 01 / 01", and "20250101" into ISO format). The following aspects were cleaned: The standard format "2025-01-01" was used; URL format (removing tracking parameters and standardizing protocol prefixes); vulnerability number format (unifying variants such as "cve2024-1234" and "CVE_2024_1234" to the standard "CVE-2024-1234" format); and text encoding (standardizing UTF-8). Common OCR errors and formatting anomalies were corrected (e.g., misidentifying "0day" as "Oday", confusing the letter "I" with the number "1"). Duplicate IOCs in the same type of threat intelligence were normalized and deduplicated (using the normalized IOC value as the key for set deduplication). Different writing styles of APT organization aliases (e.g., "APT28" and "Fancy Bear"), malware aliases (e.g., "WannaCry" and "WannaCrypt"), and vulnerability numbers were initially normalized and mapped uniformly using the built-in security entity alias dictionary. These cleansing results served as input for subsequent security element extraction.
[0030] General text cleaning methods, such as stop word removal, HTML tag filtering, and whitespace compression, cannot solve the unique noise problems of security intelligence texts. Security reports contain numerous proprietary noise patterns that general cleaning methods cannot handle, such as: multiple spelling variations of CVE numbers ("CVE_2024_1234", "cve2024-1234", "CVE:2024-1234"), mixed use of APT group aliases ("Fancy Bear" / "APT28" / "Sofacy" referencing the same group), confusion in identifying "0day" / "0day" / "0-day" when OCR scans security documents, and errors in truncating and splicing malicious hash values at the end of lines in security report tables. This invention designs a specialized cleaning rule set to address these proprietary noise problems in security intelligence. These rules are derived from in-depth analysis of real security intelligence data, rather than simple applications of general cleaning methods, and have clear domain-specific relevance.
[0031] S40. Through a multi-layered hybrid extraction pipeline, identify and extract security elements from the cleaned and preprocessed intermediate text objects.
[0032] In this embodiment of the invention, a multi-layered hybrid extraction pipeline is used to identify and extract security elements from cleaned and preprocessed intermediate text objects. This includes: constructing a three-layered hybrid extraction pipeline comprising a rule layer, a knowledge layer, and a model layer; identifying and extracting security elements from the cleaned and preprocessed intermediate text objects according to the extraction order of the rule layer, knowledge layer, and model layer in the three-layered hybrid extraction pipeline; and performing cross-document normalization mapping on the extraction results with reference to a built-in security entity alias dictionary to obtain the security elements. More specifically: The security elements in this invention include, but are not limited to: APT organization and attack group name; vulnerability number and vulnerability alias; malware family, Trojan name, and sample name; IOC indicators such as IP address, domain name, URL, email address, and file hash; attack techniques and tactics, exploitation methods, delivery methods, and persistence methods; attack targets, affected industries, geographical regions, attack time, and scope of impact. This invention adopts a three-layer fusion extraction architecture of "rule layer + knowledge layer + model layer": (1) Rule layer: Design a high-precision regular expression rule set for IOC elements (IP, domain name, hash, CVE number, etc.), and achieve millisecond-level batch matching through finite state machine to ensure high recall rate; (2) Knowledge layer: Based on a special dictionary in the field of network security (including APT organization name library, malware family name library, ATT&CK technical number reference table, industry classification dictionary, etc.), perform sliding window scanning on the text to achieve fast matching of known named entities, and use alias mapping table to complete normalization; (3) Model layer: For new threat entities that the rule layer and knowledge layer fail to identify (such as newly emerging APT organization names, malware families not included in the database), call the large language model fine-tuned by the security corpus to perform context-aware entity recognition and relation extraction, and output formatted JSON extraction results. The three-layer results are merged and deduplicated according to priority: the rule layer results have the highest priority (accuracy is guaranteed), followed by the knowledge layer, and the model layer results are attached with confidence scores. Results below the threshold are marked as "awaiting manual review". The aforementioned fusion mechanism effectively controls the false sampling rate while maintaining high coverage, thus solving the inherent defects of pure rule-based methods in identifying novel threats and pure model-based methods in having a high false sampling rate.
[0033] Entities in the cybersecurity field are characterized by a high degree of homonymy and synonymy (the same APT organization may have more than ten aliases, and the same vulnerability may have different numbering formats in different reports), which general NLP extraction methods cannot handle. This invention constructs a three-layer fusion extraction pipeline of "rule layer + knowledge layer + model layer" and performs cross-document normalization mapping on the extraction results through a built-in security entity alias dictionary, so that the same entity from different intelligence sources is ultimately standardized into the same standard expression, laying the foundation for entity alignment in the subsequent knowledge graph. At the same time, a unique ID derived from the combination of content digest hash and source URL is generated for each intelligence, along with metadata such as collection time, source URL, and source category, ensuring traceability throughout the entire process; such cross-document and cross-source unified entity normalization capability has not yet been fully implemented in existing threat intelligence processing solutions.
[0034] S50: Map the extracted security elements to a pre-defined unified threat intelligence data model.
[0035] The pre-defined unified threat intelligence data model in this embodiment of the invention includes an identity identification group, a time information group, a content description group, a security element group, and a quality control group, to achieve field mapping and unified structured expression through five types of field groups. Specifically: the identity identification group includes at least an intelligence identifier, intelligence source, and source category; the time information group includes at least the release time and collection time; the content description group includes at least the intelligence type, summary information, and text fragments; the security element group includes at least an entity list, an IOC list, a TTP list, and a target object list; and the quality control group includes at least the credibility or confidence score and the location of the original evidence. More specifically: UTIDM (Unified Threat Intelligence Data Model) is a standardized data model for cybersecurity scenarios. It aims to integrate diverse intelligence sources and formats using a unified field system, enabling structured results to be directly integrated with downstream retrieval, question answering, graphing, and analysis modules. The model is divided into five field groups: (1) Identity group, which includes at least the intelligence identifier (SHA-256 hash generated from source URL + publication time + content digest, ensuring global uniqueness), intelligence source, and source category; (2) Time information group, which includes at least the publication time and collection time (both in ISO format). (8601 standard format storage); (3) Content description group, including at least the intelligence type (vulnerability / malware / APT report / IOC intelligence, etc.), summary information (Chinese summary of no more than 200 characters automatically generated by LLM), and text fragments (a list of original text fragments divided by semantic paragraphs for vectorized retrieval); (4) Security element group, including at least the entity list (named entities such as APT organizations, malware, and vulnerabilities and their standardized names), IOC list (various IOC indicators and their type labels), TTP list (ATT&CK tactical and technical numbers), and target object list (victim industry, geographical region, and target asset type); (5) Quality control group, including at least the credibility or confidence score (calculated by a comprehensive weighted average of source authority, content completeness, and timeliness) and the location of original evidence (pointing to the record identifier in the original text citation library, supporting conclusion backtracking). UTIDM is designed to be compatible with the core fields of the STIX 2.1 specification, while extending the language fields and Chinese entity standardization fields required for Chinese intelligence processing. Table 1 shows the comparison between the core fields of UTIDM and STIX 2.1.
[0036] Table 1 Comparison of core fields between UTIDM and STIX 2.1
[0037] Existing intelligence processing solutions typically design data structures for a single downstream task (such as keyword annotation structures designed specifically for inverted indexes, or triple structures designed specifically for knowledge graphs). This results in the same intelligence needing to be processed and stored separately for different downstream modules, leading to data redundancy and difficulty in ensuring consistency. The UTIDM proposed in this invention can simultaneously serve keyword retrieval (through inverted index fields), semantic retrieval (through text fragment vectorization fields), knowledge graph construction (through entity list and TTP list fields), and intelligent question answering (through summary fields and original evidence location fields) after a single processing step, achieving the goal of "one-time processing, multi-modal reuse." While compatible with the core fields of the STIX 2.1 specification, this structural model extends language fields and Chinese entity normalization fields specifically for Chinese security intelligence scenarios, providing a clear technological increment beyond existing standards.
[0038] Further, please see Figure 3 The automatic processing method for multi-source heterogeneous threat intelligence data provided in this embodiment of the invention further includes: S60: Based on a pre-defined unified threat intelligence data model, unified storage management is performed on different downstream modules.
[0039] The unified threat intelligence data model mapped by this invention can be directly adapted to subsequent search, question answering, and graph modules, resulting in stronger data reuse capabilities. It can achieve multi-purpose storage and intermediate result output, with structured results simultaneously written to: an inverted index for keyword retrieval; text blocks and vectorized input sets for semantic retrieval; an entity relationship intermediate table for knowledge graph construction; and a source text citation library for original source tracing. For repeatedly collected or subsequently updated intelligence, it is merged, updated, or incrementally supplemented based on source address, title similarity, publication time, content summary hash, and key entity consistency to maintain the real-time performance and consistency of the intelligence database.
[0040] In summary, the automatic processing method for multi-source heterogeneous threat intelligence data proposed in this invention, compared with existing technologies, forms an end-to-end capability from raw heterogeneous intelligence to high-quality structured data that can directly serve downstream modules for multimodal intelligent analysis. The innovations of this invention are specifically reflected in: a "zero-code" multi-source access extension mechanism driven by configuration table metadata, allowing access to new data sources without modifying core code; adaptive format recognition and parser selection based on content fingerprint detection, solving the problem of format marking errors; a multi-layered fusion security entity extraction pipeline, balancing coverage and accuracy; a cross-document entity normalization security entity alias mapping mechanism; and a unified intelligence structure model (UTIDM) for "one-time processing and multimodal reuse." Through unified access, unified parsing, unified extraction, and unified expression, the efficiency of threat intelligence data collection, the quality of structured data, and downstream adaptability are improved. Automated multi-dimensional exploration enhances the timeliness of intelligence acquisition, providing high-quality foundational data for subsequent intelligent question answering, attack chain analysis, and knowledge graph construction, achieving one-stop automated pipeline execution.
[0041] Secondly, please see Figure 4 This invention provides an automatic processing system for multi-source heterogeneous threat intelligence data, which includes: The data collection module is used to create a threat intelligence source configuration table for different types of threat intelligence sources, and to collect the corresponding raw text objects using the data collection methods corresponding to different types of threat intelligence sources. The parsing module is used to adaptively parse the original text object based on the text content fingerprint and output an intermediate text object; The preprocessing module is used to perform intelligence content cleaning and standardization preprocessing on intermediate text objects; The identification and extraction module is used to identify and extract security elements from cleaned and preprocessed intermediate text objects through a multi-layer hybrid extraction pipeline. The mapping module is used to map the extracted security elements to a pre-defined unified threat intelligence data model.
[0042] Furthermore, such as Figure 4 As shown, the multi-source heterogeneous threat intelligence data automatic processing system provided in this embodiment of the invention also includes a storage module; the storage module is used to perform unified storage management of different downstream modules based on a preset unified threat intelligence data model.
[0043] As the system embodiment of the second aspect is basically similar to the method embodiment of the first aspect, the description is relatively simple, and relevant details can be found in the description of the method embodiment of the first aspect.
[0044] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0045] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the specification and accompanying drawings, will understand and implement other variations of the disclosed embodiments in carrying out the claimed invention. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. While certain measures are described in different embodiments, this does not mean that these measures cannot be combined to produce good results.
[0046] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A multi-source heterogeneous threat intelligence data automatic processing method, characterized in that, The automatic processing method for multi-source heterogeneous threat intelligence data includes: Establish a threat intelligence source configuration table corresponding to different types of threat intelligence sources, and collect the corresponding raw text objects using the collection methods corresponding to different types of threat intelligence sources; The original text object is adaptively parsed and an intermediate text object is output based on the text content fingerprint. Perform intelligence content cleaning and standardization preprocessing on the intermediate text object; Security elements are identified and extracted from cleaned and preprocessed intermediate text objects through a multi-layered hybrid extraction pipeline. The extracted security elements are mapped to a pre-defined unified threat intelligence data model.
2. The method of claim 1, wherein, The threat intelligence source configuration table includes the access method, update cycle, collection rules, authentication information, file type, and priority of different types of threat intelligence sources.
3. The method of claim 1, wherein, Different types of threat intelligence sources include at least open intelligence platform interfaces, vulnerability announcements and security bulletin pages, security vendor blogs and report pages, local historical intelligence file directories, user-uploaded data, and other semi-structured security content sources.
4. The method of claim 1, wherein, The text content fingerprint includes the text type field and text content characteristics in the original text object; The corresponding adaptive parsing of the original text object to output an intermediate text object includes: Based on the text type field, the original text object is parsed for the first time; Based on the text content features, the original text object after the first parsing is parsed a second time to output an intermediate text object.
5. The method of claim 1, wherein, The intermediate text object undergoes intelligence content cleaning and standardization preprocessing, including: Remove webpage navigation, copyright notices, advertising content, irrelevant footnotes, abnormal line breaks, and duplicate paragraphs; Standardize time format, URL format, vulnerability number format, and text encoding; Correct common OCR errors or formatting issues; Normalize and deduplicate duplicate IOCs in the same type of threat intelligence sources; The different writing formats of APT organization aliases, malware aliases, and vulnerability numbers are normalized, and cross-document normalization mapping is performed with reference to the built-in security entity alias dictionary.
6. The method of claim 1, wherein, Security elements are identified and extracted from cleaned and preprocessed intermediate text objects using a multi-layered hybrid extraction pipeline, including: A three-layer hybrid extraction pipeline consisting of a rule layer, a knowledge layer, and a model layer is constructed. Based on the extraction order of the rule layer, knowledge layer, and model layer in the three-layer hybrid extraction pipeline, the intermediate text objects after cleaning and preprocessing are identified and extracted. The extraction results are then normalized and mapped across documents with reference to the built-in secure entity alias dictionary to obtain the security elements.
7. The method of claim 1, wherein, The unified threat intelligence data model includes an identity group, a time information group, a content description group, a security element group, and a quality control group, which achieve field mapping and unified structured expression through five types of field groups.
8. The method of claim 7, wherein, An identification group includes at least an intelligence identifier, intelligence source, and source category; The time information group should include at least the release time and the collection time; The content description group should include at least the intelligence type, summary information, and text excerpt; A security element group should include at least an entity list, an IOC list, a TTP list, and a target object list. The quality control group should include at least a credibility or confidence score and the location of the original evidence.
9. The automatic processing method for multi-source heterogeneous threat intelligence data according to claim 1, characterized in that, The automatic processing method for multi-source heterogeneous threat intelligence data also includes: Based on a pre-defined unified threat intelligence data model, unified storage management is implemented for different downstream modules.
10. An automatic processing system for multi-source heterogeneous threat intelligence data, characterized in that, The automated processing system for multi-source heterogeneous threat intelligence data includes: The data collection module is used to create a threat intelligence source configuration table for different types of threat intelligence sources, and to collect the corresponding raw text objects using the data collection methods corresponding to different types of threat intelligence sources. The parsing module is used to adaptively parse the original text object based on the text content fingerprint and output an intermediate text object; The preprocessing module is used to perform intelligence content cleaning and standardization preprocessing on intermediate text objects; The identification and extraction module is used to identify and extract security elements from cleaned and preprocessed intermediate text objects through a multi-layer hybrid extraction pipeline. The mapping module is used to map the extracted security elements to a pre-defined unified threat intelligence data model.
Citation Information
Patent Citations
Multi-source heterogeneous data management method and system for network threats
CN118174971A
Threat intelligence intelligent analysis method based on large model and knowledge graph
CN121418178A