A threat data resource directory updating method and management system

By using large-scale pre-trained language models for intelligent change inference and conflict resolution, the problem of low efficiency in updating and management of threat data resource catalogs is solved, automated updates and efficient data retrieval are realized, and the efficiency of network security threat analysis is improved.

CN118939651BActive Publication Date: 2025-05-13INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410951885.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-05-13
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively update and manage threatened data resource catalogs, resulting in low flexibility in resource catalogs, low efficiency in data retrieval and update, and lack of semantic considerations in data changes.

Method used

Large-scale pre-trained language model is adopted to conduct intelligent change inference and conflict resolution from the perspective of data content, and realize dynamic updates of threat data resource catalogs.

Benefits of technology

By deeply mining data semantic information, improving data change and digestion technology, automatic update of data resource catalogs is achieved, data retrieval and update efficiency is improved, manual intervention is reduced, costs are reduced, and network security threat analysis is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118939651B_ABST
    Figure CN118939651B_ABST
Patent Text Reader

Abstract

The present invention discloses an update method and management system for a threat data resource directory, and belongs to the field of network security technology. The present invention generates a semantic index matrix of deterministic data by using a large-scale pre-trained language model, compares it with the semantic index matrix of the threat data resource directory, and performs data change inference and conflict resolution based on similarity; for uncertain data, the mapping comparison between field names and metadata is performed, the data structure or data type is updated according to the mapping ratio, and the uncertain data is converted into deterministic data before data change inference. The present invention can deeply mine data semantic information, improve data change and resolution technology from the perspective of content understanding, realize automatic update of data resource directory types, dynamically adapt to different types of data, improve the efficiency of data retrieval and update, reduce manual intervention, reduce labor costs, and improve the efficiency of network security threat analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention specifically relates to a threat data resource directory update method and management system, belonging to the technical field of network security. Background Art

[0002] With the continuous development of Internet technology and the increasing speed of informatization construction, new types of network attacks emerge in an endless stream, and network security confrontation is becoming more and more intense. In the past, network security threat analysis and monitoring based solely on network traffic data can no longer cope with the current new types of network attacks and source tracing analysis. As a result, existing researchers and enterprises have continuously introduced more data in the process of network security threat analysis, including network threat intelligence such as vulnerability data, sample data, reputation data, network asset data such as units, cloud platforms, equipment, information systems, and network security business data such as security incidents, security inspections, response disposal, etc. In order to effectively organize and associate large amounts of scattered network data and business data, and improve the efficiency and effectiveness of threat analysis and security monitoring, network security laboratories and enterprises have tried to build a network security big data processing platform and construct a relatively comprehensive threat data resource directory to support the management and operational analysis of network security data. However, with the upgrade of various network traffic probe devices, network security monitoring equipment, network threat analysis technology, and the strengthening of security management requirements, the data types, data structures, and data levels of various threat data resources are constantly adjusted and changed. In the context of the continuous iteration and update of data resources in the current network security field and the ever-changing needs of data application and analysis, building a definite and complete threat data resource directory and solidifying data sources and data storage and update methods can no longer meet actual needs.

[0003] The dynamic update technology of data resource directory is one of the key links in data management. It involves how to update and manage the data resource directory in a timely manner to adapt to the continuous growth and change of data. This technology mainly includes two aspects: one is the identification and update of data, and the other is the update and maintenance of the resource directory. The data update and identification technology uses polling comparison to identify the newly added and modified data items and update them to the data resource directory. Polling comparison includes full-text polling comparison, unique identifier polling comparison or combination keyword comparison. The effects of these three types of comparisons grow inversely proportional to the comparison performance, and cannot achieve semantic level recognition and update. Resource directory update and maintenance technology is usually formulated and adjusted by domain experts when expanding new data sources, new data structures and new data management requirements. Although this method is stable and accurate, it cannot cope with changes in data sources and adjustments to data structures. It can be seen that the existing technology has the disadvantages of low flexibility of resource directories, slow efficiency in data retrieval and update, and lack of semantic considerations for data changes. Therefore, facing the network security field where various types of data analysis and application technologies are constantly iterating and updating, designing and implementing a technology that can automatically and intelligently identify changed data and automatically update and maintain data resource directories can largely solve the problems of poor threat data management performance and low threat data analysis utilization, and indirectly support the analysis effect in the network security application field. Summary of the invention

[0004] The purpose of the present invention is to propose a threat data resource directory update method and management system, which uses a large-scale pre-trained language model to intelligently change inference and resolve conflicts from the perspective of data content understanding, so as to realize dynamic update of threat data resource directory data.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for updating a threat data resource directory comprises the following steps:

[0007] Identify whether the accessed data is deterministic data. Deterministic data refers to data types that can be accepted by the threat data resource directory.

[0008] If it is deterministic data, data change inference is performed, and the steps include: using a large-scale pre-trained language model to generate a semantic index matrix of deterministic data; comparing the generated semantic index matrix with the semantic index matrix of the threat data resource directory; if the semantic index matrices are the same, no operation is performed; if the cosine similarity of the semantic index matrix is ​​equal to or greater than the similarity threshold, data change and conflict resolution operations are performed; if the cosine similarity of the semantic index matrix is ​​less than the similarity threshold, an insertion operation is performed;

[0009] If it is uncertain data, a data type comparison analysis of the threat data resource directory is performed, and the steps include: extracting the field name of the uncertain data and mapping it with the metadata of the threat data resource directory; if the mapping ratio reaches or exceeds the threshold, updating the data structure of the mapped data type; if the mapping ratio does not exceed the threshold, updating the data type of the threat data resource directory; after the update, the uncertain data is regarded as deterministic data and data change inference is performed.

[0010] Furthermore, the step of identifying whether the access data is deterministic data includes:

[0011] Determine whether the data source of the accessed data is a known source in the threat data resource directory. If so, proceed to the next step;

[0012] Determine whether the data type and data field of the accessed data match the known data type and corresponding data field of the threat data resource directory. If so, proceed to the next step;

[0013] Determine whether the access data meets the data field requirements of the corresponding data type of the threat data resource directory. The data field includes attributes, mandatory fields and data formats. If all are met, it is determined to be deterministic data, otherwise it is determined to be uncertain data.

[0014] Furthermore, the steps of generating a semantic index matrix of deterministic data using a large-scale pre-trained language model include:

[0015] Assemble the field names and data values ​​of deterministic data into strings to form text;

[0016] Calculate the sum of the vectors of symbol embedding, fragment embedding, and position embedding of each character in the above text to obtain the vector representation of the text, and use a large-scale pre-trained language model to encode the vector representation of the text to generate a high-dimensional vector space representation;

[0017] The hash value represented by the above vector space is calculated using a locality-sensitive hashing algorithm, and the binary hash code of the hash value is used as a semantic index matrix of deterministic data.

[0018] Furthermore, if the cosine similarity of the semantic index matrix is ​​equal to or greater than the threshold, the specific steps of performing data change and conflict resolution operations include:

[0019] If the cosine similarity of the index matrix is ​​equal to or greater than the threshold, then select a piece of original data whose cosine similarity value reaches or exceeds the similarity threshold and has the highest cosine similarity;

[0020] Extracting a unique identifier of the same data type from the deterministic data and the above original data, where the unique identifier is a metadata and its attribute value or a group of metadata and its attribute value;

[0021] Compare the two extracted unique identifiers to see if they are consistent. If they are consistent, perform data changes and resolve conflicts.

[0022] Furthermore, the steps of performing data change and conflict resolution include:

[0023] For the default items of the original data, directly insert the corresponding data fields of the deterministic data to make data changes;

[0024] For the conflicting items between the original data and the deterministic data, the historical data of the corresponding data type in the threat data resource directory is extracted to train the Transformer model, and the attribute values ​​of the conflicting items are predicted with the maximum accuracy by adjusting the model parameters;

[0025] Remove the conflicting attribute values ​​of the original data and the deterministic data, input the remaining text string into the Transformer model to predict the conflicting attribute values, and obtain the predicted values;

[0026] Calculate the cosine similarity between the conflicting attribute value of the deterministic data and the predicted value, as well as the cosine similarity between the conflicting attribute value of the original data and the predicted value, and select the attribute value with the highest cosine similarity as the data change result.

[0027] Furthermore, the step of extracting the field name of the uncertainty data and mapping and comparing it with the metadata of the threat data resource directory includes:

[0028] Extract the field names of uncertain data and use a large-scale pre-trained language model to represent the field names in vector space;

[0029] The cosine similarity between the vector space representation of the field name and the metadata of the threat data resource directory is calculated, and the metadata with the largest cosine similarity and exceeding the similarity threshold is selected as the attribute representation of the uncertainty data to obtain the converted data structure;

[0030] Compare the converted data structure with the data type definitions in the threat data resource directory to determine the consistency of the attribute fields.

[0031] Furthermore, the large-scale pre-trained language model is a BERT model or an XLNet model.

[0032] Furthermore, if the mapping ratio reaches or exceeds the threshold, the data structure of the mapping type is updated, which means that if the required attribute fields are consistent and more than half of the other attribute fields are the same, the newly added attributes in the uncertainty data are added to the metadata set of the threat data resource directory, and supplemented to the data structure under the corresponding data type, and the data type definition is updated.

[0033] Furthermore, if the mapping ratio does not exceed the threshold, the type of the threat data resource directory is updated, which means that if the required attribute fields are inconsistent, or more than half of the other attribute fields are the same, the data type with the largest number of overlapping data attributes is found through comparison, and a new data type is created at the same level as the data type. The type is named the uncertain data name, and the type structure is the converted data structure.

[0034] A management system for a threat data resource directory, comprising:

[0035] The threat data resource directory management module is used to view the types of threat data resource directories, view and interactively analyze data type changes, and view and interactively analyze field type changes;

[0036] A threat data retrieval module is used to perform semantic retrieval on threat data in the threat data resource directory based on a semantic index matrix;

[0037] The threat data intelligent update module is used to access various types of threat data and dynamically update the data resource directory according to the above update method.

[0038] The present invention has the following advantages:

[0039] The present invention utilizes large-scale pre-trained models to deeply mine data semantic information, improves data change and resolution technology from the perspective of content understanding, and proposes an automatic update mechanism for data resource directory types, which can dynamically adapt to different types of data, making data retrieval and update more efficient, and making up for the shortcomings of insufficient semantic considerations in existing methods. It can automatically respond to the operational management of various new types of network security threat-related data, reduce manual intervention, and reduce labor costs. Through intelligent data updating and management, the efficiency of network security threat analysis is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The present invention is a flowchart of a method for updating a threat data resource directory according to an embodiment of the present invention.

[0041] Figure 2 This is a diagram of threat data resource catalog - data type definition example.

[0042] Figure 3 This is a schematic diagram of access data-text conversion. DETAILED DESCRIPTION

[0043] In order to make the various technical features and advantages or technical effects in the above technical solutions of the present invention more obvious and easy to understand, they are described in detail below with reference to the accompanying drawings.

[0044] The embodiment of the present invention specifically discloses a method for updating a threat data resource directory, and its brief processing flow is as follows: Figure 1 The threat data resource directory targeted by this method for updating is composed of multi-level network security related data types, data type definitions, metadata and threat data.

[0045] 1) Multi-level network security related data types effectively organize different types of data according to the way the data is generated, the content, and the update method, including: ① original data generated in network communications, such as communication logs and protocol analysis logs; ② threat alarm logs and sample analysis reports that have been initially screened by network security monitoring equipment rule matching; ③ network security events and threat intelligence that have been analyzed and judged by the network threat monitoring system; and ④ network assets and network event handling data generated in network security management.

[0046] 2) Data type definition, which is used to describe and define the source of the current type of data and the attribute fields it should have. The attribute fields are composed of a set of metadata and can be further divided into mandatory attributes and non-mandatory attributes. One or more mandatory attribute combinations can be used as the unique identifier of the current type of data. For example, a data type can be defined as Figure 2 The examples shown are each a data type.

[0047] 3) Metadata is a set of network security threat data elements, including descriptions of network space subject attributes such as IP, IP segment, domain name, link, fingerprint, sample hash, virtual account, region, etc., as well as elements describing network behavior such as time, threat type, exploitation tools, technical means, impact description, etc.

[0048] The specific processing steps of this method are described as follows:

[0049] Step 1: Identify whether the incoming data is deterministic data

[0050] By judging whether the access data is a data type that can be received by the current threat data resource directory, it is further determined whether the access data is deterministic data. Specifically, it includes the following two steps:

[0051] 1) Determine whether the source of the accessed data exists in the threat data resource directory and is mapped to one or more corresponding data types. If so, it is determined to be deterministic data, otherwise it is uncertain data;

[0052] 2) Determine whether the fields and field values ​​of the accessed data meet the attribute requirements, mandatory requirements, and data format requirements of the corresponding data type. If all are met, then make deterministic data changes and resolve conflicts. If not, then the data is deemed to be of non-quality and discarded directly.

[0053] Step 2: Intelligently change and resolve conflicts on deterministic data

[0054] Generate a retrieval index (as a semantic index matrix) for deterministic data using a large-scale pre-trained language model, and compare and analyze it with the data index (also a semantic index matrix) in the threat data resource directory. If the indexes are the same, no operation is performed. If the indexes are similar, data change and conflict resolution operations are carried out. If the indexes are different, an insertion operation is executed. If the indexes are similar, the data change and conflict resolution operations specifically include the following steps:

[0055] 1) Assemble the field names and data values of the deterministic data InputData into a string to form a text S describing the data content, in the specific form as Figure 3 shown.

[0056] 2) To obtain richer semantic features, calculate the sum of the vectors of each character in the text S in three dimensions: symbol embedding, segment embedding, and position embedding, to obtain the vector representation of the text, and then input it into the large-scale pre-trained language model BERT or XLNet for vector embedding, to obtain the vector space representation X=(x1,x2,…x n ), where n is the maximum length of the text, and x i (1 < i < n) is a 1024-dimensional vector space.

[0057] 3) Since directly using the space vector as the storage space of the data index has a large overhead and complex calculations, this method uses Locality-Sensitive Hashing (LSH) that can capture the spatial relationship between vectors to reduce the dimension of the vector space representation X of InputData, and the calculation steps are as follows:

[0058] First, define n 1024-dimensional random vectors a1,a2,…a n , such that each vector element follows a standard normal distribution;

[0059] Secondly, for the vector space representation X of InputData, calculate n hash values;

[0060] h1(x) = sign(a1·x1)

[0061] h2(x) = sign(a2·x2)

[0062] …

[0063] hn(x) = sign(a n ·x n )

[0064] Then, form an n-dimensional binary hash code h(x)=(h1(x),h2(x),…,hn(x)) from the n hash values as the semantic index matrix of InputData.

[0065] 4) Compare the semantic index matrix of InputData with the semantic index matrix of all data under the corresponding data type in the threat data resource directory. If the two semantic index matrices are the same, no operation is performed; if the cosine similarity of the two semantic index matrices is equal to or greater than the threshold, select a piece of data SourceData with a cosine similarity value that reaches or exceeds a similarity threshold and has the highest cosine similarity, where the similarity threshold is empirically adjusted by the usage scenario. Further extract the unique identifiers of InputData and SourceData and compare whether they are consistent. If they are consistent, the two are actually the same data, and data changes and conflict resolution are carried out. Otherwise, it is regarded as new data insertion.

[0066] 5) For the default items of SourceData, directly insert the corresponding data fields of InputData to implement data changes. For the conflicting items of InputData and SourceData with consistent unique identifiers, use the generative model Transformer to predict the data value, and then select the data value that best conforms to the semantic logic for update. The specific steps are as follows:

[0067] First, the Transformer model is trained based on the existing historical data under the data type of the threat data resource directory. This is mainly done by masking a non-mandatory attribute value in the data text and continuously learning and adjusting the Transformer parameters so that it can predict a default attribute value with the greatest possible accuracy.

[0068] Secondly, remove the conflicting attribute values ​​in InputData and SourceData, and input the remaining text string into Transformer to predict the conflicting attribute values ​​to obtain the predicted value vector t preditct ;

[0069] Then, calculate the conflicting attribute values ​​t of InputData respectively input 、SourceData conflict item attribute value t source With t preditct The cosine similarity of , takes the attribute value with the largest similarity as the data update result.

[0070] Step 3: Uncertainty data comparison and data resource catalog type update and maintenance

[0071] Compare the consistency of the uncertainty data structure and the threat data resource directory type from the semantic level, update and maintain the type or type structure of the threat data resource directory, convert the uncertainty data into deterministic data, and then update the data. The specific steps include:

[0072] 1) Extract the field name of the uncertainty data and use the BERT or XLNet model to represent it in vector space, perform similarity calculation with the metadata of the threat data resource directory, select the metadata with the largest cosine similarity and exceeding a similarity threshold as the attribute representation of the uncertainty data, and obtain the converted data structure, where the similarity threshold is empirically adjusted by the usage scenario.

[0073] 2) Compare the converted data structure with the data type definitions of the threat data resource directory. If the required attribute fields are inconsistent, go to step 3) to create a new data type; if all the required attribute fields are consistent, further compare whether the other attribute fields are the same. If more than half of the other attribute fields are the same, introduce the new attributes in the uncertainty data into the metadata set of the threat data resource directory and add them to the data structure under the corresponding data type to update the data type definition; if less than half of the other attribute fields are the same, it is considered a new data type, and step 3) to create a new data type is executed.

[0074] 3) Compare and find the data type with the largest number of overlapping data attributes, and create a new data type at the same level as this data type. The type is named the name of the uncertain data, and the type structure is the data structure of the uncertain data after metadata comparison and conversion.

[0075] 4) After the threat data resource directory is updated and maintained, the uncertain data is treated as deterministic data and inserted into the corresponding data type to complete the data update.

[0076] Although the present invention has been disclosed as above by way of embodiments, it is not intended to limit the present invention. Appropriate modifications or equivalent substitutions of the technical solutions of the present invention made by ordinary technicians in the field should all be included in the protection scope of the present invention. The protection scope of the present invention shall be based on what is defined in the claims.

Claims

1. A method for updating a threat data resource directory, characterized in that: The following steps are involved: Identify whether the accessed data is deterministic data. Deterministic data refers to data types that can be accepted by the threat data resource directory. If it is deterministic data, data change inference is performed, and the steps include: using a large-scale pre-trained language model to generate a semantic index matrix for deterministic data; comparing the generated semantic index matrix with the semantic index matrix of the threat data resource directory; if the semantic index matrices are the same, no operation is performed; if the cosine similarity of the semantic index matrix is ​​equal to or greater than the similarity threshold, data change and conflict resolution operations are performed on the data in the threat data resource directory; if the cosine similarity of the semantic index matrix is ​​less than the similarity threshold, an insertion operation is performed; if the cosine similarity of the semantic index matrix is ​​equal to or greater than the threshold, data change and conflict resolution operations are performed on the data in the threat data resource directory. The specific steps include: if the cosine similarity of the index matrix is ​​equal to or greater than the threshold, selecting a piece of original data whose cosine similarity value reaches or exceeds the similarity threshold and has the highest cosine similarity; extracting a unique identifier of the same data type in the deterministic data and the above original data, and the unique identifier is An identifier is a metadata and its attribute value or a group of metadata and its attribute value; compare whether the two extracted unique identifiers are consistent, and if they are consistent, perform data change and conflict resolution; the steps of performing data change and conflict resolution include: for the default items of the original data, directly insert the corresponding data fields of the deterministic data to perform data change; for the conflicting items of the original data and the deterministic data, extract the historical data under the corresponding data type of the threat data resource directory to train the Transformer model, and predict the attribute values ​​of the conflicting items with the maximum accuracy by adjusting the model parameters; remove the attribute values ​​of the conflicting items of the original data and the deterministic data, input the remaining text string into the Transformer model to predict the attribute values ​​of the conflicting items, and obtain the predicted values; calculate the cosine similarity between the attribute values ​​of the conflicting items of the deterministic data and the predicted values, and the cosine similarity between the attribute values ​​of the conflicting items of the original data and the predicted values, and select the attribute value with the highest cosine similarity as the data change result; If it is uncertain data, a data type comparison analysis of the threat data resource directory is performed, and the steps include: extracting the field name of the uncertain data and mapping it with the metadata of the threat data resource directory; if the mapping ratio reaches or exceeds the threshold, updating the data structure of the mapped data type; if the mapping ratio does not exceed the threshold, updating the data type of the threat data resource directory; after the update, the uncertain data is regarded as deterministic data and data change inference is performed.

2. The updating method according to claim 1, characterized in that: The step of identifying whether the access data is deterministic data includes: determining whether the data source of the access data is a known source of the threat data resource directory, and if so, proceeding to the next step; Determine whether the data type and data field of the accessed data match the known data type and corresponding data field of the threat data resource directory. If so, proceed to the next step; Determine whether the access data meets the data field requirements of the corresponding data type of the threat data resource directory. The data field includes attributes, mandatory fields and data formats. If all are met, it is determined to be deterministic data, otherwise it is determined to be uncertain data.

3. The updating method according to claim 1, characterized in that: The steps of generating a semantic index matrix of deterministic data using a large-scale pre-trained language model include: Assemble the field names and data values ​​of deterministic data into strings to form text; Calculate the sum of the vectors of symbol embedding, fragment embedding, and position embedding of each character in the above text to obtain the vector representation of the text, and use a large-scale pre-trained language model to encode the vector representation of the text to generate a high-dimensional vector space representation; The hash value represented by the above vector space is calculated using a locality-sensitive hashing algorithm, and the binary hash code of the hash value is used as a semantic index matrix of the deterministic data.

4. The updating method according to claim 1, characterized in that: The steps of extracting the field names of uncertainty data and mapping and comparing them with the metadata of the threat data resource directory include: Extract the field names of uncertain data and use a large-scale pre-trained language model to represent the field names in vector space; The cosine similarity between the vector space representation of the field name and the metadata of the threat data resource directory is calculated, and the metadata with the largest cosine similarity and exceeding the similarity threshold is selected as the attribute representation of the uncertainty data to obtain the converted data structure; Compare the converted data structure with the data type definitions in the threat data resource directory to determine the consistency of the attribute fields.

5. The updating method according to claim 1, 3 or 4, characterized in that: The large-scale pre-trained language model is the BERT model or the XLNet model.

6. The updating method according to claim 4, characterized in that: If the mapping ratio reaches or exceeds the threshold, the data structure of the mapping type is updated. This means that if the required attribute fields are consistent and more than half of the other attribute fields are the same, the newly added attributes in the uncertainty data are added to the metadata set of the threat data resource directory and supplemented to the data structure under the corresponding data type to update the data type definition.

7. The updating method according to claim 4, characterized in that: If the mapping ratio does not exceed the threshold, the type of the threat data resource directory is updated, which means that if the required attribute fields are inconsistent, or more than half of the other attribute fields are the same, the data type with the largest number of overlapping data attributes is found through comparison, and a new data type is created at the same level as the data type. The type is named the uncertain data name, and the type structure is the converted data structure.

8. A management system for a threat data resource directory, characterized in that: include: The threat data resource directory management module is used to view the types of threat data resource directories, view and interactively analyze data type changes, and view and interactively analyze field type changes; A threat data retrieval module is used to perform semantic retrieval on threat data in the threat data resource directory based on a semantic index matrix; The threat data intelligent update module is used to access various types of threat data and dynamically update the data resource directory according to the update method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Dynamically organizing cloud computing resources to facilitate discovery

    CN103733194A

  • Multi-source threat intelligence fusion method and device, equipment and storage medium

    CN114925757A