Method and apparatus for constructing a service data set
Patent Information
- Application Number
- CN202211321107.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-10-26
AI Technical Summary
[0003]本申请实施例的目的在于提供一种业务数据集的构造方法及装置,能够解决数据要素之间的无关联问题,从而能够让测试数据集聚焦于被纳管的终端资产风险要素上,进而能够保障数据集构造的合理性、有效性、多样性,方便业务效果快速呈现
[0034]在上述实现过程中,该业务数据集的构造装置可以通过获取单元获取业务数据;通过提取单元提取业务数据中的数据要素;通过梳理单元对数据要素进行梳理,得到要素集合;通过划分单元对要素集合进行关系层级划分,得到要素层级划分结果;通过转换单元将要素层级划分结果转换为数据集配置文件;通过构造单元来基于数据集配置文件和业务数据,构造业务数据集。可见,该装置能够解决数据要素之间的无关联问题,从而能够让测试数据集聚焦于被纳管的终端资产风险要素上,进而能够保障数据集构造的合理性、有效性、多样性,方便业务效果快速呈现。
Smart Images

Figure CN115495754B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a method and apparatus for constructing a business dataset. Background Technology
[0002] When analyzing and demonstrating endpoint security, endpoint security involves not only the threats faced by the endpoint, but also various other factors such as the endpoint type, its security domain, vulnerabilities, and the stage and level of the attack. Therefore, given so many factors, how to correlate and freely combine these business datasets to create a dataset that can quickly demonstrate business results is one of the technical problems that business personnel are eager to solve. Summary of the Invention
[0003] The purpose of this application is to provide a method and apparatus for constructing a business dataset, which can solve the problem of the lack of correlation between data elements, thereby enabling the test dataset to focus on the risk elements of the managed terminal assets, and thus ensuring the rationality, effectiveness and diversity of the dataset construction, and facilitating the rapid presentation of business results.
[0004] The first aspect of this application provides a method for constructing a business dataset, including:
[0005] Obtain business data;
[0006] Extract data elements from the business data;
[0007] The data elements are sorted to obtain an element set;
[0008] The element set is hierarchically divided to obtain the element hierarchy division result;
[0009] Convert the feature hierarchy partitioning results into a dataset configuration file;
[0010] Based on the dataset configuration file and the business data, construct the business dataset.
[0011] In the above implementation process, this method prioritizes acquiring business data and extracting data elements from it. This extracts multiple elements relevant to the current business from each data set, facilitating subsequent categorization based on these elements. Next, the method organizes these data elements into element sets, completing the first step of data element organization. Then, it performs a hierarchical relationship partitioning of the element sets, resulting in a hierarchical partitioning of elements. This step confirms the relationships between all element sets and constructs a relationship graph, which in turn establishes a hierarchical partitioning of the element sets, resulting in a strongly correlated data element graph. Subsequently, the method converts the hierarchical partitioning results into a dataset configuration file, creating a data configuration document. This document enables subsequent steps to construct the business dataset based on the relationships between the data elements. Finally, the method constructs a reasonable and effective business dataset based on the dataset configuration file and the business data. In summary, this method can solve the problem of lack of correlation between data elements, thereby enabling the test dataset to focus on the risk elements of the managed terminal assets, and thus ensuring the rationality, effectiveness, and diversity of the dataset construction, facilitating the rapid presentation of business results.
[0012] Furthermore, the step of sorting the data elements to obtain an element set includes:
[0013] Identify special elements among the data elements that require special processing;
[0014] Remove the special elements from the data elements to obtain the elements to be processed;
[0015] The elements to be processed are sorted out to obtain an element set.
[0016] In the above implementation process, this method can prioritize identifying special elements requiring special processing during the process of sorting data elements to obtain an element set; then, it removes these special elements from the data elements to obtain the elements to be processed; and finally, it sorts these elements to obtain the element set. It is evident that this method can eliminate all elements requiring special processing, thereby preventing such data elements from participating in subsequent hierarchical divisions, and thus improving the effectiveness of business data construction.
[0017] Furthermore, the step of sorting the elements to be processed to obtain an element set includes:
[0018] Based on the principle of high cohesion and low coupling, the elements to be processed are organized into an element set.
[0019] In the above implementation process, this method can organize the elements to be processed into element sets based on the principle of grouping elements with high cohesion into one element set and distributing elements with loose coupling into different element sets, thereby packaging and processing all scattered data elements, which is beneficial to the construction of the dataset.
[0020] Furthermore, the step of performing a hierarchical division of the element set to obtain the element hierarchical division result includes:
[0021] Based on the association relationships and association order of the element set, the element set is divided into multiple element levels;
[0022] Each of the aforementioned element levels is aggregated and merged to obtain the element level division result.
[0023] In the above implementation process, when dividing the element set into hierarchical relationships to obtain the element hierarchy division result, this method can first divide the element set into multiple element levels based on the association relationships and association order of the element set; then, it merges the sets of each element level to obtain the element hierarchy division result. It can be seen that this method can sort and organize the element set based on the association relationships and association order between the element sets, thereby arranging the element set into multiple levels sequentially from beginning to end, and thus constructing an element hierarchy architecture to obtain the element hierarchy division result.
[0024] Furthermore, the step of converting the feature hierarchy partitioning results into a dataset configuration file includes:
[0025] The feature hierarchy division results are converted into a dataset configuration file using the YAML language.
[0026] In the above implementation process, the method can use the YAML language to express the relationship between elements and element sets, thereby ensuring that each leaf node is an element, which facilitates the subsequent steps of building the business dataset.
[0027] A second aspect of this application provides an apparatus for constructing a business dataset, the apparatus comprising:
[0028] The acquisition unit is used to acquire business data;
[0029] Extraction unit, used to extract data elements from the business data;
[0030] The sorting unit is used to sort the data elements to obtain an element set;
[0031] A partitioning unit is used to perform hierarchical partitioning of the element set to obtain element hierarchical partitioning results;
[0032] A conversion unit is used to convert the feature hierarchy partitioning results into a dataset configuration file.
[0033] The construction unit is used to construct a business dataset based on the dataset configuration file and the business data.
[0034] In the above implementation process, the device for constructing the business dataset can acquire business data through an acquisition unit; extract data elements from the business data through an extraction unit; organize the data elements through a sorting unit to obtain an element set; perform hierarchical division of the element set through a partitioning unit to obtain an element hierarchical partitioning result; convert the element hierarchical partitioning result into a dataset configuration file through a transformation unit; and construct the business dataset based on the dataset configuration file and the business data through a construction unit. It is evident that this device can solve the problem of lack of correlation between data elements, thereby enabling the test dataset to focus on the risk elements of the managed terminal assets, thus ensuring the rationality, effectiveness, and diversity of the dataset construction and facilitating the rapid presentation of business results.
[0035] Furthermore, the sorting unit includes:
[0036] The identification subunit is used to identify special elements among the data elements that require special processing.
[0037] The elimination sub-unit is used to remove the special element from the data elements to obtain the element to be processed.
[0038] The sorting sub-unit is used to sort the elements to be processed to obtain an element set.
[0039] In the above implementation process, the sorting unit can identify special elements in the data elements that require special processing by identifying sub-units; remove special elements from the data elements by removing sub-units to obtain the elements to be processed; and sort the elements to be processed by the sorting sub-units to obtain the element set. It is evident that this device can remove all elements requiring special processing, thereby preventing such data elements from participating in subsequent hierarchical divisions, thus improving the effectiveness of business data construction.
[0040] Furthermore, the sorting subunit is specifically used to sort the elements to be processed into an element set based on the principle of high cohesion and low coupling.
[0041] In the above implementation process, the device can organize the elements to be processed into element sets based on the principle of grouping elements with high cohesion into one element set and distributing elements with loose coupling into different element sets, thereby packaging and processing all scattered data elements, which is beneficial to the construction of datasets.
[0042] Furthermore, the partitioning unit includes:
[0043] Sub-units are used to divide the element set into multiple element levels based on the association relationship and association order of the element set;
[0044] The merge sub-unit is used to merge each of the aforementioned element levels to obtain the element level division result.
[0045] In the above implementation process, the partitioning unit can divide the element set into multiple element levels based on the association relationship and association order of the element set by dividing it into sub-units; by merging sub-units, each element level is merged to obtain the element level partitioning result. It can be seen that this method can sort and organize the element set based on the association relationship and association order between the element sets, thereby arranging the element set into multiple levels sequentially from beginning to end, thus constructing an element level architecture and obtaining the element level partitioning result.
[0046] Furthermore, the conversion unit is specifically used to convert the feature hierarchy partitioning results into a dataset configuration file using the YAML language.
[0047] In the above implementation process, the device can use the YAML language to express the relationship between elements and sets of elements, thereby ensuring that each leaf node is an element, which facilitates the device in building business datasets.
[0048] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor runs the computer program to cause the electronic device to perform the method for constructing a business dataset as described in any one of the first aspects of this application.
[0049] A fourth aspect of this application provides a computer-readable storage medium storing computer program instructions, which, when read and executed by a processor, perform the method for constructing a business dataset as described in any one of the first aspects of this application. Attached Figure Description
[0050] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 A flowchart illustrating a method for constructing a business dataset as provided in an embodiment of this application;
[0052] Figure 2 A schematic diagram of a device for constructing a business dataset provided in an embodiment of this application;
[0053] Figure 3 A flowchart illustrating a method for constructing a business dataset as provided in an embodiment of this application;
[0054] Figure 4 A schematic diagram of an element set provided for an embodiment of this application;
[0055] Figure 5 This is a schematic diagram of the decomposition of an element set provided in an embodiment of this application;
[0056] Figure 6 This application provides a schematic diagram of a merging set of elements.
[0057] Figure 7 This application provides a schematic diagram of an element hierarchy division process.
[0058] Figure 8 This is a schematic diagram of the element hierarchy division result provided in an embodiment of this application;
[0059] Figure 9 An example diagram of a dataset configuration file provided in an embodiment of this application;
[0060] Figure 10 An example diagram of a business dataset provided in an embodiment of this application;
[0061] Figure 11 This is a schematic diagram illustrating a method for randomly selecting and generating multiple test datasets using code, as provided in an embodiment of this application. Detailed Implementation
[0062] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0063] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0064] Example 1
[0065] Please refer to Figure 1 , Figure 1 This embodiment provides a flowchart illustrating a method for constructing a business dataset. The method for constructing the business dataset includes:
[0066] S101. Obtain business data.
[0067] S102. Extract data elements from business data.
[0068] In this embodiment, the method can prioritize organizing the necessary elements of the data.
[0069] In this embodiment, each type of data in the method is composed of various elements. The elements need to be sorted out according to the specified business and represented by a unique identifier.
[0070] S103. Identify special elements in the data elements that require special processing.
[0071] S104. Remove special elements from the data elements to obtain the elements to be processed.
[0072] S105. Based on the principle of high cohesion and low coupling, the elements to be processed are sorted into an element set.
[0073] In this embodiment, the method can organize the relationships between elements based on the principle of high cohesion and low coupling, and prioritize the elimination of elements that require special processing.
[0074] In this embodiment, the principle of high cohesion and low coupling is:
[0075] a) Grouping elements with high cohesion into a single set of elements;
[0076] b) Divide loosely coupled elements into different element sets;
[0077] c) Use a unique identifier to represent the set of elements.
[0078] S106. Based on the association relationship and association order of the element set, divide the element set into multiple element levels.
[0079] S107. Merge each element level to obtain the element level division result.
[0080] In this embodiment, the method can divide the element relationship hierarchy and adjust the element set according to the relationship hierarchy.
[0081] In this embodiment, the method completes the relationship sorting of the element set. This step is relatively complex, but the general principle is:
[0082] a) If there are sequential constraints between elements in a set, such as a one-to-many relationship, then the element set needs to be separated.
[0083] b) If there is a strong correlation between elements, such as a one-to-one relationship, it can be regarded as one element;
[0084] c) Perform hierarchical and sorting of the element set to determine the order between elements, which will facilitate the writing of element relationships later;
[0085] d) Merging the element sets of the first layer;
[0086] e) Re-adjust the unique identifier to represent the set of elements.
[0087] S108. Using the YAML language, convert the feature hierarchy partitioning results into a dataset configuration file.
[0088] In this embodiment, the method can generate a dataset configuration file by defining the hierarchy and relationships between elements and using machine-readable YAML language.
[0089] In this embodiment, the method can use the YAML language to express the relationship between elements and element sets, and ensure that each leaf node is an element.
[0090] S109. Construct a business dataset based on the dataset configuration file and business data.
[0091] In this embodiment, the method can express the element values and the relationships between elements at the beginning using directory and file name, and set the value of each leaf node, that is, the element and put it into the corresponding txt file.
[0092] In this embodiment, after the above-mentioned element sorting work is completed, the method can use a program to extract and integrate element values.
[0093] In this embodiment, the method can achieve a reasonable combination and generation of business test datasets by encoding according to the dataset configuration information and the number of dataset entries.
[0094] Please refer to Figure 3 , Figure 3 A flowchart illustrating a method for constructing a business dataset is shown.
[0095] For example, this method can be applied to the accumulation of any business data. Specifically, this method can be used to ensure the rapid presentation of business data through an enterprise situational awareness system. When conducting risk perception, the enterprise situational awareness system needs to perform comprehensive risk calculations based on information such as the terminal's own attributes, the threats faced by the terminal, and the vulnerabilities existing on the terminal. There are three main types of data here: terminal, threat, and vulnerability. This example will use the method mentioned in the invention to generate threat data that can be associated with the terminal. The specific example process is as follows:
[0096] (1) Threat identification is a type of log data. The log contains the following core elements and unique identifiers:
[0097] Threat IP (att_ip), threat region (att_addr), victim asset IP (aset_ip), victim asset name (asset_name), victim security domain (sec_dom), attack type (att_type), attack details (att_msg), attack stage (att_kc), severity level (severity), attack result (att_result), etc.
[0098] (2) Based on the principles of high cohesion and loose coupling, we classify the elements into three sets, which are represented as follows: Figure 4 .
[0099] (3) The relationships within and between element sets need to be sorted out.
[0100] ① First, determine the constraints within a set of elements and decompose the set of elements with constraints. For example, in the "External Threat Information" set, IP addresses within each region should not be duplicated. Therefore, each threat source IP has a sequential relationship with its region, so we need to split this set of elements into two sets. Similarly, in attack information, attack types within each attack stage should not be duplicated, and attack stages and attack types have a sequential relationship. Therefore, we need to decompose the attack stages within this set of elements into a separate set of elements. And so on. This method can decompose the set of elements according to constraints, resulting in the following set of elements: Figure 5 As shown.
[0101] ② Determine if there are strong relationships between elements in each set. For example, if there is a one-to-one relationship between a victim's asset IP and a victim's asset name, then we represent these two elements as a single element. Specifically... Figure 6 As shown.
[0102] ③ Next, we begin to analyze the order and hierarchy of the elements. For example, the threat's geographical location must be determined before the threat's source IP can be identified; therefore, the threat's geographical location is located above the threat's source IP. The final hierarchical relationship is as follows: Figure 7 As shown.
[0103] ④ Next, we will merge the first-level feature sets and assign unique identifiers to each set. The result is as follows: Figure 8 As shown.
[0104] (4) Use the YAML language to express the relationships between these collections, such as... Figure 9 As shown.
[0105] (5) After obtaining the feature set and feature relationships, we express the feature values and relationships using directory and file names, and set the value of each leaf node (i.e., feature) and put it into the corresponding txt file. The results can be found in [reference needed]. Figure 10 .
[0106] (6) Multiple test datasets were randomly selected and generated using code, and the results are as follows: Figure 11 As shown.
[0107] Implementing this embodiment allows for the addition of a vulnerability risk element to terminal assets and the expansion of the dataset when a larger dataset needs to be created. This can be achieved without altering the existing business dataset relationships; simply add the vulnerability information of the terminal assets as a core element to the existing yml file and expand the current element values. This significantly reduces the difficulty of expanding the dataset.
[0108] In this embodiment, the technical problem to be solved by this method is as follows:
[0109] (1) When constructing the configuration file dataset, first sort out the data set to increase the scalability of data elements;
[0110] (2) To sort out the hierarchical constraint relationships between data elements, increase the business rationality of the dataset, and solve the problem of the lack of correlation between data elements;
[0111] (3) By setting the value range of data elements, the problem of the value range of data elements being too random can be avoided.
[0112] In this embodiment, the method focuses on the organization and representation of relationships within the dataset to facilitate program interpretation. Specifically, when organizing the element set, the method not only adopts the general principles of high cohesion and low coupling, but also organizes the constraints between elements within the set using set and hierarchical relationships. Furthermore, it represents the element set and its hierarchical constraints using YAML for easy program interpretation. Finally, it represents the element set and its values using a directory format, facilitating random sampling during business test dataset organization to ensure the relevance, rationality, and diversity of business data.
[0113] In this embodiment, the subject executing the method can be a computing device such as a computer or server, and no limitation is made in this embodiment.
[0114] In this embodiment, the subject executing the method can also be a smart device such as a smartphone or tablet, and no limitation is made in this embodiment.
[0115] As can be seen, the method for constructing the business dataset described in this embodiment can reduce the investment in constructing invalid and fragmented data during actual terminal security posture business testing. This solves the problem of the lack of correlation between data elements, allowing the test dataset to focus on the risk elements of the managed terminal assets, thereby ensuring the rationality, effectiveness, and diversity of the dataset construction and facilitating the rapid presentation of business results. Furthermore, the dataset construction scheme in this method has excellent scalability; the dataset can be further divided into sub-datasets to enhance its diversity, and more datasets can be added to it, providing more favorable data support for data analysis.
[0116] Example 2
[0117] Please refer to Figure 2 , Figure 2 This is a schematic diagram of a device for constructing a business dataset provided in this embodiment. Figure 2 As shown, the apparatus for constructing this business dataset includes:
[0118] Acquisition unit 210 is used to acquire business data;
[0119] Extraction unit 220 is used to extract data elements from business data;
[0120] The sorting unit 230 is used to sort the data elements to obtain an element set;
[0121] Division unit 240 is used to perform hierarchical division of the element set to obtain the element hierarchical division result;
[0122] Transformation unit 250 is used to convert feature hierarchy partitioning results into a dataset configuration file;
[0123] Construction unit 260 is used to construct a business dataset based on the dataset configuration file and business data.
[0124] As an optional implementation, the sorting unit 230 includes:
[0125] The identification subunit 231 is used to identify special elements in the data elements that require special processing.
[0126] Sub-unit 232 is used to remove special features from data features to obtain the features to be processed;
[0127] Sub-unit 233 is used to sort out the elements to be processed and obtain the element set.
[0128] As an optional implementation, the sorting subunit 233 is specifically used to sort the elements to be processed into an element set based on the principle of high cohesion and low coupling.
[0129] As an optional implementation, the partitioning unit 240 includes:
[0130] Sub-unit 241 is used to divide the element set into multiple element levels based on the association relationship and association order of the element set;
[0131] Merging sub-unit 242 is used to merge each element level to obtain the element level division result.
[0132] As an optional implementation, the conversion unit 250 is specifically used to convert the feature hierarchy partitioning results into a dataset configuration file using the YAML language.
[0133] In this embodiment, the explanation of the device for constructing the business dataset can be referred to the description in Embodiment 1, and will not be repeated here.
[0134] As can be seen, the business dataset construction apparatus described in this embodiment can reduce the investment in constructing invalid and fragmented data during actual terminal security posture business testing, thereby solving the problem of unrelated data elements. This allows the test dataset to focus on the risk elements of the managed terminal assets, ensuring the rationality, effectiveness, and diversity of the dataset construction, and facilitating the rapid presentation of business results. Furthermore, the dataset construction scheme in this method has excellent scalability; the dataset can be further divided into sub-datasets to enhance its diversity, and more datasets can be added to it, providing stronger data support for data analysis.
[0135] This application provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor runs the computer program to enable the electronic device to execute the method for constructing a business dataset in embodiment 1 of this application.
[0136] This application provides a computer-readable storage medium storing computer program instructions. When these computer program instructions are read and executed by a processor, the method for constructing a business dataset as described in Embodiment 1 of this application is performed.
[0137] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0138] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0139] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0140] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0141] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0142] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for constructing a business dataset, characterized in that, The method includes: Acquire business data; the business data includes terminal data, threat data, and vulnerability data; Extract data elements from the business data; the data elements include threat IP, threat region, victim asset IP, victim asset name, victim security domain, attack type, attack details, attack stage, severity level, and attack result; The data elements are sorted to obtain an element set; The element set is hierarchically divided to obtain the element hierarchy division result; Convert the feature hierarchy partitioning results into a dataset configuration file; Based on the dataset configuration file and the business data, construct the business dataset; The step of performing a hierarchical division of the element set to obtain the element hierarchical division result includes: Based on the association relationships and association order of the element set, the element set is divided into multiple element levels according to the sequential constraints. Each of the aforementioned element levels is aggregated and merged to obtain the element level division result.
2. The method for constructing a business dataset according to claim 1, characterized in that, The step of sorting the data elements to obtain an element set includes: Identify special elements among the data elements that require special processing; The special elements are removed from the data elements to obtain the elements to be processed. The elements to be processed are sorted out to obtain an element set.
3. The method for constructing a business dataset according to claim 2, characterized in that, The step of sorting the elements to be processed to obtain an element set includes: Based on the principle of high cohesion and low coupling, the elements to be processed are organized into an element set.
4. The method for constructing a business dataset according to claim 1, characterized in that, The step of converting the feature hierarchy partitioning results into a dataset configuration file includes: The feature hierarchy division results are converted into a dataset configuration file using the YAML language.
5. A device for constructing a business dataset, characterized in that, The apparatus for constructing the business dataset includes: An acquisition unit is used to acquire business data; the business data includes terminal data, threat data, and vulnerability data. The extraction unit is used to extract data elements from the business data; the data elements include threat IP, threat region, victim asset IP, victim asset name, victim security domain, attack type, attack details, attack stage, severity level, and attack result; The sorting unit is used to sort the data elements to obtain an element set; A partitioning unit is used to perform hierarchical partitioning of the element set to obtain element hierarchical partitioning results; A conversion unit is used to convert the feature hierarchy partitioning results into a dataset configuration file. A construction unit is used to construct a business dataset based on the dataset configuration file and the business data; The division units include: The sub-unit division is used to divide the element set into multiple element levels according to the association relationship and association order of the element set and the sequential constraints. Merge sub-units, used to merge sets of each feature level to obtain the feature level division result.
6. The apparatus for constructing a business dataset according to claim 5, characterized in that, The combing unit includes: The identification subunit is used to identify special elements among the data elements that require special processing. The elimination sub-unit is used to remove the special element from the data elements to obtain the element to be processed. The sorting sub-unit is used to sort the elements to be processed to obtain an element set.
7. The apparatus for constructing a business dataset according to claim 6, characterized in that, The sorting subunit is specifically used to sort the elements to be processed into an element set based on the principle of high cohesion and low coupling.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory being used to store a computer program, and the processor running the computer program to cause the electronic device to perform the method for constructing a business dataset according to any one of claims 1 to 4.
9. A readable storage medium, characterized in that, The readable storage medium stores computer program instructions, which, when read and executed by a processor, perform the method for constructing the business dataset as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Hierarchical information system attack defense capability comprehensive evaluation system and method
CN109918914A