Method, device, equipment and medium for importing asset information into graph database
By using a preset classification framework to categorize asset information and establish connecting edges, the customization problem when importing enterprise asset information into a graph database is solved, improving import efficiency and applicability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QI-ANXIN LEGENDSEC INFORMATION TECH (BEIJING) INC
- Filing Date
- 2023-03-02
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, importing enterprise asset information into graph databases requires customized development based on the specific circumstances of each enterprise, resulting in poor classification universality and low efficiency.
A pre-defined classification framework based on the asset classification requirements of multiple IT entities is adopted to classify the asset information set according to basic classification units, and connection edges between nodes are created based on the relationships between asset data.
It improves the efficiency of importing asset information into graph databases, reduces the need for customized development, and is suitable for various types of IT entities.
Smart Images

Figure CN116383425B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, device and medium for importing asset information into a graph database. Background Technology
[0002] Corporate assets, as core information of a company, need to be effectively managed.
[0003] Currently, graph databases are primarily used for asset management in Internet Technology (IT) companies. Examples include Neo4j, JanusGraph, and NebulaGraph. To import an IT company's asset information into a graph database for management, the first step is to collect asset information from various parts of the company. Then, the collected asset information is categorized. Finally, the categorized asset information is imported into the nodes of the graph database, and connecting edges are established between nodes based on the relationships between the asset information within each node.
[0004] However, different companies have different actual situations, and their asset situations also vary. In the process of classifying company asset information, the classification criteria for each company's asset information need to be determined separately based on the actual situation of each company's asset information. For example, Company A's asset information includes asset information 1, asset information 2, and asset information 3. Analysis of the specific content of Company A's asset information reveals that Company A's asset information types include type a and type b. Asset information 1 and asset information 2 belong to type a, and asset information 3 belongs to type b. Therefore, when importing Company A's asset information into the graph database, asset information 1 and asset information 2 are imported into a certain node in the graph database, and asset information 3 is imported into another node in the graph database. Company B's asset information includes asset information 3, asset information 4, and asset information 5. Analysis of the specific content of Company B's asset information reveals that Company B's asset information types include type b, type c, and type d. Asset information 3 belongs to type b, asset information 4 belongs to type c, and asset information 5 belongs to type d. Therefore, when importing company B's asset information into the graph database, asset information 3 is imported into one node, asset information 4 into another node, and asset information 5 into yet another node. It is evident that current enterprise asset information classification requires first determining the asset types involved in each company based on its specific asset content, and then classifying all of the company's asset information based on these determined asset types. Because different companies have significantly different needs for asset information classification, the universality of enterprise asset information classification is poor. Consequently, when importing enterprise asset information into the graph database, customized node development is required based on the specific needs of each company. Summary of the Invention
[0005] The purpose of this application is to provide a method, apparatus, device, and medium for importing asset information into a graph database, so as to improve the efficiency of importing asset information into a graph database.
[0006] To address the aforementioned technical problems, this application provides the following technical solutions:
[0007] The first aspect of this application provides a method for importing asset information into a graph database. The method includes: receiving asset data from an IT entity to obtain an asset information set of the IT entity; classifying the asset information set according to basic classification units in a preset classification framework to obtain classified asset information corresponding to the basic classification units, wherein the basic classification units in the preset classification framework are indivisible classification units determined based on the asset classification needs of multiple IT entities, and the basic classification units correspond to nodes in the graph database; importing the classified asset information into the corresponding nodes of the graph database according to the basic classification units; determining the association relationship between each node based on the relationship between the asset data of the IT entity, and creating connection edges between the corresponding nodes based on the association relationship.
[0008] A second aspect of this application provides an apparatus for importing asset information into a graph database. The apparatus includes: a receiving module for receiving asset data of an IT entity to obtain an asset information set of the IT entity; a classification module for classifying the asset information set according to basic classification units in a preset classification framework to obtain classified asset information corresponding to the basic classification units, wherein the basic classification units in the preset classification framework are indivisible classification units determined based on the asset classification needs of multiple IT entities, and the basic classification units correspond to nodes in the graph database; a node import module for importing the classified asset information into the corresponding nodes of the graph database according to the basic classification units; and a connection edge import module for determining the association relationship between each node based on the relationship between the asset data of the IT entity, and creating connection edges between the corresponding nodes based on the association relationship.
[0009] A third aspect of this application provides an electronic device, comprising: a processor, a memory, and a bus; wherein the processor and the memory communicate with each other via the bus; and the processor is used to invoke program instructions in the memory to execute the method of the first aspect.
[0010] A fourth aspect of this application provides a computer-readable storage medium, comprising: a stored program; wherein, when the program is executed, it controls the device on which the storage medium is located to perform the method of the first aspect.
[0011] Compared to existing technologies, the method for importing asset information into a graph database provided in the first aspect of this application, after obtaining the asset information set of an IT entity, classifies it using basic classification units in a preset classification framework. Since the basic classification units in the preset classification framework are determined based on the asset classification needs of multiple IT entities, they are applicable to various types of IT entities. Therefore, when classifying the asset information set of an IT entity, it is not necessary to classify and summarize the asset information set of each IT entity; classification can be performed directly, improving the classification efficiency of asset information and thus improving the import efficiency of IT entity asset information into the graph database.
[0012] The apparatus for importing asset information into a graph database provided in the second aspect of this application, the electronic device provided in the third aspect, and the computer-readable storage medium provided in the fourth aspect have the same or similar beneficial effects as the method for importing asset information into a graph database provided in the first aspect. Attached Figure Description
[0013] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this application are illustrated by way of example and not limitation, with the same or corresponding reference numerals denoteing the same or corresponding parts, wherein:
[0014] Figure 1 This is a schematic diagram of the architecture for importing asset information into the graph database in an embodiment of this application;
[0015] Figure 2 This is a flowchart illustrating the method for importing asset information into a graph database in an embodiment of this application.
[0016] Figure 3 This is a schematic diagram illustrating the process of extracting asset information in an embodiment of this application;
[0017] Figure 4 This is a schematic diagram illustrating the process of integrating equipment asset information in an embodiment of this application;
[0018] Figure 5 This is a schematic diagram illustrating the process of integrating network interface card, process, port, software, and vulnerability asset information in an embodiment of this application.
[0019] Figure 6 This is a schematic diagram illustrating the process of importing asset information into the graph database in an embodiment of this application;
[0020] Figure 7 This is a schematic diagram of the device for importing asset information into the graph database in the embodiments of this application;
[0021] Figure 8 This is a schematic diagram of the structure of the electronic device in the embodiments of this application. Detailed Implementation
[0022] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0023] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains.
[0024] Currently, due to the flexible query language of graph databases, complex correlation analysis of enterprise assets can be easily achieved. Therefore, many enterprises are importing their asset information into graph databases for management. However, the specific content of asset information varies greatly among different enterprises. Each enterprise needs to determine the categorizable asset types based on its own asset information, then classify the asset information according to the classification types, and finally import the classified asset information into the nodes of the graph database, as well as add connecting edges to each node based on the relationships between the classified asset information. It is evident that the current method of importing enterprise assets into graph databases requires customized development for each enterprise, reducing the efficiency of importing asset information into graph databases.
[0025] Through in-depth research, the inventors discovered that the main reason for the low efficiency of importing enterprise asset information into graph databases is the wide variety of specific asset information from different enterprises. If a comprehensive classification standard could be provided, allowing each enterprise to categorize its asset information according to this standard, the step of summarizing enterprise asset types would be eliminated, thereby improving the efficiency of importing enterprise assets into graph databases. Furthermore, importing enterprise assets into graph databases would no longer require customized development; a single classification method could be adopted by various IT enterprises, saving costs associated with importing enterprise assets into graph databases.
[0026] In view of this, embodiments of this application provide a method, apparatus, device, and medium for importing asset information into a graph database. A preset classification framework is determined based on the asset classification needs of multiple IT entities. This framework includes multiple indivisible basic classification units. Thus, the asset information of any IT entity to be imported into the graph database does not need further classification and summarization; it can be classified according to the basic classification units in the preset classification framework. The classified asset information is then imported into the corresponding nodes in the graph database, and connection edges are established between the corresponding nodes based on the relationships between the classified asset information, thereby improving the efficiency of importing IT enterprise asset information into the graph database.
[0027] In order to clearly explain the method, apparatus, device and medium for importing asset information into the graph database provided in the embodiments of this application, the architecture in which asset information is imported into the graph database will be described first.
[0028] Figure 1 This is a schematic diagram of the architecture for importing asset information into the graph database in this application embodiment. See also... Figure 1 As shown, this architecture may include: a data input layer, a data extraction layer, a data fusion layer, a knowledge reasoning layer, a RESTful API, and a front-end.
[0029] The data input layer is primarily used to interface with enterprise asset information. Specifically, it can interface with enterprise asset data in RDBMS (e.g., MySQL, PostgreSQL) and NoSQL databases (e.g., MongoDB, ElasticSearch).
[0030] The data extraction layer is primarily used to break down and standardize enterprise asset information (e.g., unifying the content format of heterogeneous fields (e.g., int / long conversion, time format conversion, field definition, etc.)). Specifically, through entity extraction, relationship extraction, attribute extraction, and extraction task management, combined with rule strategies, machine learning, and human intervention, the raw asset information is initially transformed into enterprise asset information with unified standards and input into the Data Material Library (RDBMS) for later use.
[0031] The data fusion layer is primarily used to normalize and automatically complete the data of the initially processed enterprise assets (e.g., determining the data insertion time, whether the Internet Protocol (IP) address is internal or external, whether it is IPv4 / IPv6, and whether it should be written to a graph database). It also generates a new Universally Unique Identifier (UUID) for each piece of data and writes it to the UUID mapping table, which in turn writes it to the formal database tables. Specifically, through entity ambiguity correction, relation error correction, entity linking, and knowledge merging, combined with rule-based strategies, machine learning, and human intervention, the initially processed asset information is transformed into multiple correct, non-duplicated asset records, which are then input into the Association Analysis Database (GregorDB) and the Data Fusion Database (RDBMS) for later use.
[0032] The knowledge reasoning layer is primarily used to analyze the relationships between various asset information within an enterprise. Specifically, it achieves diversified relationship analysis of various asset information through asset correlation analysis, asset management, and asset topology analysis.
[0033] RESTful APIs are primarily used to connect asset information with users. Specifically, they enable users to manage asset information through asset queries, asset updates / deletions, and vulnerability checks.
[0034] The front end is mainly used to display asset information to users in a flexible and intuitive way.
[0035] Based on this architecture, the method for importing asset information into a graph database provided in the embodiments of this application will be described in detail below.
[0036] Figure 2 This is a flowchart illustrating the method for importing asset information into a graph database in an embodiment of this application. See [link / reference]. Figure 1 As shown, the method may include:
[0037] S201: Receive asset data from the IT entity and obtain a set of asset information of the IT entity.
[0038] The IT entity here refers to the IT company or individual that needs to import asset data into the graph database for management. Of course, the IT entity can also be other organizations. The specific object referred to as the IT entity is not limited here, as long as it is relevant to the IT scenario.
[0039] When an IT entity needs to import its asset information into a graph database for management, it can expose the various storage objects containing the asset data to the execution entity of this embodiment. The execution entity then interfaces with these storage objects to export and aggregate the various asset data stored within them, forming an asset information set. Alternatively, the IT entity can proactively send all its asset data to the execution entity of this embodiment. The execution entity will then aggregate all received asset data to form an asset information set. The specific method of receiving asset data from the IT entity is not limited here.
[0040] S202: Classify the asset information set according to the basic classification units in the preset classification framework to obtain the classified asset information corresponding to the basic classification units.
[0041] The basic classification units in the preset classification framework are indivisible units determined based on the asset classification needs of multiple IT entities. These basic classification units correspond to nodes in the graph database.
[0042] To obtain a comprehensive asset classification applicable to various IT entities, the asset classification needs of multiple IT entities can be pre-calculated to obtain the asset classifications of each IT entity. These classifications are then integrated together, and each classification in the integrated classification is a classification unit. These classification units form the pre-defined classification framework.
[0043] The basic classification units in the preset classification framework are derived from the asset classification needs of multiple IT entities. Therefore, the basic classification units in the preset classification framework can be applied to the classification of various IT entities. For some IT entities, their asset information is relatively small, and the classification may not cover every basic classification unit in the preset classification framework. Therefore, the basic classification units that are not covered can be ignored.
[0044] For example, suppose a pre-defined classification framework includes basic classification units A, B, and C, and an IT entity possesses asset information 1, asset information 2, and asset information 3. During classification, since asset information 1 and asset information 2 belong to basic classification unit A, and asset information 3 belongs to basic classification unit C, asset information 1 and asset information 2 are classified into basic classification unit A, and asset information 3 is classified into basic classification unit C. Basic classification unit B is ignored.
[0045] In the specific process of classifying an asset information set according to the basic classification units in a preset classification framework, the asset information set can first be divided into different major categories according to each basic classification unit, and then each major category can be further divided into different categories of asset information according to the specific assets. Of course, the asset information set can also be directly divided into different major categories according to each basic classification unit, and these different major categories are the different categories of asset information.
[0046] For example, suppose a pre-defined classification framework includes basic classification unit A and basic classification unit B, and an IT entity possesses asset information 1, asset information 2, and asset information 3. Since asset information 1 and asset information 2 belong to basic classification unit A, and asset information 3 belongs to basic classification unit B, asset information 1 and asset information 2 can be grouped into one category, and asset information 3 into another. Then, asset information 1 and asset information 2 belonging to different specific assets can be further separated, resulting in three categories of asset information: asset information 1, asset information 2, and asset information 3. Alternatively, asset information 1 and asset information 2 can be directly grouped into one category, and asset information 3 into another, resulting in two categories: asset information 1 and asset information 2, and asset information 3. The specific classification method depends on the level of detail in the asset classification and the specific meaning of each node in the graph database; no specific limitations are imposed here.
[0047] S203: Import the classified asset information into the corresponding nodes of the graph database according to the basic classification units.
[0048] After obtaining the categorized asset information, i.e., the categorized asset information, the asset information for each category can be imported into a node of the graph database. Of course, it is also possible to import asset information from several categories into a node of the graph database, which needs to be determined according to the specific management requirements of the IT entity in the graph database, and is not limited here.
[0049] S204: Based on the relationships between asset data of the IT entity, determine the association between each node, and create connection edges between the corresponding nodes based on the association.
[0050] After importing the categorized asset information into the nodes of the graph database, the relationships between the asset information stored in each node can be determined, i.e., asset relationships. Then, based on these asset relationships, connection edges are created between the corresponding nodes in the graph database. When there is no relationship between two asset information pieces, there is no connection edge between their corresponding nodes. When there is a relationship between two asset information pieces, such as a subordinate relationship or an upstream / downstream relationship, a connection edge corresponding to that relationship needs to be established between the two corresponding nodes.
[0051] As can be seen from the above, the method for importing asset information into a graph database provided in this application, after obtaining the asset information set of an IT entity, classifies it through basic classification units in a preset classification framework. Since the basic classification units in the preset classification framework are determined based on the asset classification needs of multiple IT entities, they can be applied to various types of IT entities. Therefore, when classifying the asset information set of an IT entity, it is not necessary to classify and summarize the asset information set of each IT entity. Classification can be performed directly, improving the classification efficiency of asset information and thus improving the import efficiency of the asset information of IT entities into the graph database.
[0052] Furthermore, the basic classification units in the preset classification framework can include: devices, network interface cards (NICs), processes, ports, software, and vulnerabilities. These specific basic classification units can cover various asset information involved in the IT entity, and the types are not too numerous or complex, which can further improve the classification efficiency of asset information, thereby improving the efficiency of importing asset information into the graph database.
[0053] Specifically, step S202 above may include:
[0054] Step A: Classify the asset information set according to device, network card, process, port, software and vulnerability to obtain information on each device, network card, process, port, software and vulnerability.
[0055] In other words, after obtaining the asset information set, the asset information is first categorized by device, network interface card (NIC), process, port, software, and vulnerability, resulting in device-related information sets, NIC-related information sets, process-related information sets, port-related information sets, software-related information sets, and vulnerability-related information sets. Then, the device-related information set is further divided into different device information (i.e., individual device information) based on the specific device. The further categorization of NICs, processes, ports, software, and vulnerabilities follows the same method as for devices, and will not be elaborated here.
[0056] As can be seen from the above, classifying asset information sets according to devices, network cards, processes, ports, software, and vulnerabilities can cover various asset information types involved in the IT entity without causing too many types and resulting in complex classification. This ensures the comprehensiveness of asset information classification and also improves the classification efficiency of asset information, thereby improving the import efficiency of asset information into the graph database.
[0057] Furthermore, after classifying the asset information set according to devices, network cards, processes, ports, software, and vulnerabilities, the information of each device, network card, process, port, software, and vulnerability can be further classified to facilitate viewing of various detailed information.
[0058] Specifically, after step A above, the method may further include:
[0059] Step B1: Categorize each device information according to the device details items in the preset device list to obtain the device list.
[0060] The details of each device in the preset device list are determined based on the details of multiple devices.
[0061] In other words, detailed information about various devices is collected in advance, and then the collected detailed information is gathered together, cleaned, and deduplicated. The resulting detailed information is all the detailed information related to the device, that is, the detailed items of each device in the preset device list.
[0062] After obtaining the information for each device, you can continue to write the details for each device according to the items in the preset device list. This way, each piece of equipment information corresponds to a device list. The items and format of the device lists are identical across all device information entries, facilitating subsequent viewing of asset information.
[0063] In practical applications, the details of each device in the preset device list may include: device number, device name, device type, device version, device manufacturer, device serial number, device CPE, device description, CPU architecture, CPU description, whether it is domestically produced, hard drive model, hard drive capacity, motherboard number, memory model, memory capacity, asset responsible person number, asset confidentiality, asset integrity, asset availability, asset importance, asset risk level, asset risk value, intranet security domain, network number, unit number, computer room number, rack location, device form, device manufacturing time, device registration time, device online time, device offline time, device activation time, device deactivation time, device first discovery time, whether the device is online, whether the device is classified, communication device type, security device type, protected object number, security device business function, server type, whether it is a key asset, storage device type, storage method, storage medium, disk array mode, endpoint device type, whether it is a removable device, auto-incrementing ID, and asset responsible person name.
[0064] To provide a more intuitive understanding of the details of each device in the preset device list, the following table displays the details of each device in the preset device list.
[0065] Table 1
[0066]
[0067]
[0068]
[0069]
[0070]
[0071] Step B2: Categorize the information of each network card according to the details of each network card in the preset network card list to obtain a list of each network card.
[0072] The details of each network card in the preset device list are determined based on the details of multiple network cards.
[0073] The specific classification method for network card information is the same as that for device information. Please refer to the specific classification method for device information in step B1 above. Here, only the specific content of each network card detail item in the preset network card list is given.
[0074] The details of each network card in the preset network card list may include at least one of the following: network card number, network card name, network card manufacturer, network card IPv4, network card IPv6, network card MAC, network card enabled status, whether it is the default network card, the unique identifier of the device to which the network card belongs, and the auto-incrementing ID.
[0075] To provide a more intuitive understanding of the details of each network card in the preset network card list, the following is a list of the details of each network card in the preset network card list.
[0076] Table 2
[0077]
[0078] Step B3: Categorize the information of each process according to the process details items in the preset process list to obtain the process list.
[0079] The process details in the preset device list are determined based on the details of multiple processes.
[0080] The specific classification method for process information is the same as that for device information. Please refer to the specific classification method for device information in step B1 above. Here, only the specific content of each process detail item in the preset process list is given.
[0081] The process details in the preset process list can include at least one of the following: process ID, process name, process path, process running status, process description, process operation result code, process operation behavior code, process alarm type, process command line, process startup user ID, process working directory, process network connection, parent process ID, process digital signature, parent process path, process start time, process end time, whether it is a privileged process, process dynamic library list, process service ID, process software ID, process level, process asset unique identifier, process compiled ID, process corresponding device unique identifier, and auto-incrementing ID.
[0082] To provide a more intuitive understanding of the details of each process in the preset process list, the details of each process in the preset process list are displayed in a list format below.
[0083] Table 3
[0084]
[0085]
[0086] Step B4: Categorize each port information according to the port details items in the preset port list to obtain the port list.
[0087] The port details in the preset device ports are determined based on the details of multiple ports.
[0088] The specific classification method for port information is the same as that for device information. Please refer to the specific classification method for device information in step B1 above. Here, only the specific content of each port detail item in the preset port list is given.
[0089] The port details in the preset port list can include at least one of the following: auto-incrementing ID, unique port node identifier, port number, and unique identifier of the device to which the port belongs.
[0090] To provide a more intuitive understanding of the details of each port in the preset port list, the following table displays the details of each port in the preset port list.
[0091] Table 4
[0092] id int4 Auto-incrementing ID. asset_id varchar A unique identifier for a port node. port_num varchar Port number. dev_uuid varchar The unique identifier of the device to which the port belongs.
[0093] Step B5: Categorize each piece of software information according to the software details items in the preset software list to obtain the software list.
[0094] The software details items in the preset device list are determined based on the details of multiple software programs.
[0095] The specific classification method for software information is the same as that for device information. Please refer to the specific classification method for device information in step B1 above. Here, only the specific content of each software detail item in the preset software list is given.
[0096] The software details in the preset software list can include at least one of the following: software ID, Chinese name of the software, English name of the software, software type, software version, software description, software CPE, Chinese name of the vendor, English name of the vendor, software installation path, software license, software hash value, whether it is a key asset, whether it is open source, whether it can connect to the internet, first discovery time, creation time, update time, interface language, current software version, current software version installation time, first installed version of the software, first installed time of the software, unique identifier of the device to which the software belongs, and auto-incrementing ID.
[0097] To provide a more intuitive understanding of the details of each software item in the preset software list, the following is a list of the details of each software item in the preset software list.
[0098] Table 5
[0099]
[0100]
[0101] Step B6: Classify each vulnerability information according to the vulnerability details items in the preset vulnerability list to obtain each vulnerability list.
[0102] The vulnerability details in the preset device list are determined based on the details of multiple vulnerabilities.
[0103] The specific classification method for vulnerability information is the same as that for device information. Please refer to the specific classification method for device information in step B1 above. Here, only the specific content of each vulnerability detail item in the preset vulnerability list is given.
[0104] The vulnerability details in the preset vulnerability list can include at least one of the following: vulnerability number, vulnerability name, CVE number, QVD number, CNVD number, CNNVD number, confidentiality impact, integrity impact, availability impact, user authentication, permission requirements, user interaction, attack path, attack complexity, scope of impact, vulnerability severity level, vulnerability description, vulnerability type, creation time, unique identifier of the device to which the vulnerability belongs, and update time.
[0105] To provide a more intuitive understanding of the details of each vulnerability in the preset vulnerability list, the following table displays the details of each vulnerability in the preset vulnerability list.
[0106] Table 6
[0107]
[0108]
[0109] It should be noted that steps B1-B6 above can be executed synchronously or asynchronously. The specific order in which steps B1-B6 are executed is not specified here.
[0110] As can be seen from the above, after obtaining information on each device, network card, process, port, software, and vulnerability, the system further categorizes and refines this information according to the details of each device in the preset device list, each network card in the preset network card list, each process in the preset process list, each port in the preset port list, each software in the preset software list, and each vulnerability in the preset vulnerability list. This allows for a more comprehensive organization of information on each device, network card, process, port, software, and vulnerability, making it easier for users to view.
[0111] Furthermore, the specific content of each item in the details of each device in the preset device list, each network card in the preset network card list, each process in the preset process list, each port in the preset port list, each software in the preset software list, and each vulnerability in the preset vulnerability list is not imagined out of thin air, but is based on certain objective facts. These objective facts can be various documents and data related to asset information involved in the daily operations of multiple IT entities.
[0112] Specifically, prior to step B1 above, the method may further include:
[0113] Step C1: Obtain full information for multiple IT entities.
[0114] The multiple IT entities mentioned here can be other entities within the same IT industry as the IT entity whose asset information is to be imported. The "full data" can refer to all information related to the assets of the IT entity, including asset data. Obtaining full data from multiple IT entities allows for a more comprehensive and accurate determination of which asset types each entity needs to classify, thereby generating a large and comprehensive asset classification table applicable to all IT entities, improving asset classification efficiency, and consequently, asset import efficiency.
[0115] In practical applications, the full information of an IT entity can be obtained through various aspects of the IT entity.
[0116] Specifically, step C1 above may include:
[0117] Step C11: Obtain all information from multiple IT entities, including national standards, product information, industry specifications, asset scanning devices, traffic probes, asset management databases, and application scenarios, as the total information for multiple IT entities.
[0118] The national standards mentioned here may include: GB / T14885-2010 (for equipment), GB / T14885-2010 (for software), and GB / T35416-2017 (for software), etc.
[0119] By acquiring full information about the IT entity through multiple dimensions, levels, and methods, the content of the full information is enriched to maximize its richness, so as to generate more detailed preset lists for each specific item in the future.
[0120] Step C2: Classify all information to obtain category information.
[0121] The category information includes details of each device, network card, process, port, software, and vulnerability in the preset device list.
[0122] Since assets are categorized into six main types—devices, network interface cards (NICs), processes, ports, software, and vulnerabilities—it is necessary to extract device-related metrics, NIC-related metrics, process-related metrics, port-related metrics, software-related metrics, and vulnerability-related metrics from the full dataset. Then, the specific details of each of these categories can be determined.
[0123] In the process of extracting specific relevant indicators from the full dataset, taking a device as an example, we can first find device-related information from the full dataset, and then filter out the indicators from that information. This way, we can obtain the device-related indicators. The method for obtaining network interface cards (NICs), processes, ports, software, and vulnerabilities is the same as for devices, and will not be elaborated here.
[0124] After extracting the specific relevant indicators, since these indicators may come from different IT entities, and the asset information types of different IT entities may overlap, there will also be duplicate indicators within the specific relevant indicators, which need to be deduplicated. Taking devices as an example, deleting duplicate indicators from the device-related indicators and arranging them in order yields the details of each device in the preset device list. The generation method for network cards, processes, ports, software, and vulnerabilities is the same as for device details, and will not be repeated here.
[0125] As can be seen from the above, by acquiring relevant indicators of devices, network cards, processes, ports, software, and vulnerabilities from the full information of multiple IT entities, the classification of asset information can be more comprehensive and applicable to the asset information classification of various IT entities. During the transfer of asset information of IT entities to the graph database, it is no longer necessary to customize a classification scheme for each enterprise. A comprehensive classification table can be used to classify the asset information of various IT entities, thereby improving the efficiency of importing asset information into the graph database.
[0126] Furthermore, in the process of classifying asset information sets according to devices, network cards, processes, ports, software, and vulnerabilities, the order in which to classify them can be selected based on the hierarchical relationship between these devices, network cards, processes, ports, software, and vulnerabilities.
[0127] Specifically, step A above may include:
[0128] Step D1: Classify the asset information set according to different equipment to obtain information on each equipment and its corresponding ancillary asset information.
[0129] Among devices, network interface cards (NICs), processes, ports, software, and vulnerabilities, devices are the most fundamental. NICs, processes, ports, software, and vulnerabilities all operate on top of devices. Therefore, the asset information set is first divided according to different devices. Information related to devices is divided into information for each device, and the remaining information is treated as the corresponding ancillary asset information based on the device it belongs to.
[0130] Step D2: Classify the corresponding ancillary asset information of each device according to different network cards, processes, ports, software and vulnerabilities to obtain the corresponding network card information, process information, port information, software information and vulnerability information for each device.
[0131] After obtaining the corresponding ancillary asset information for each device, it can be categorized according to different network cards, processes, ports, software, and vulnerabilities. In this way, each device information will have corresponding network card information, process information, port information, software information, and vulnerability information.
[0132] The order in which network cards, processes, ports, software, and vulnerabilities are categorized can be arbitrary or based on hierarchical relationships. For example, categorize by network cards and software first, then by processes, ports, and vulnerabilities. Categorization can be done individually or simultaneously. There is no specific restriction on the order of categorization for network cards, processes, ports, software, and vulnerabilities.
[0133] As can be seen from the above, the asset information is first categorized by device, and then by network interface card (NIC), process, port, software, and vulnerability. Since the device serves as the carrier for the operation of NICs, processes, ports, software, and vulnerabilities, and is the foundation, this categorization method not only divides the asset information set into six categories—device, NIC, process, port, software, and vulnerability—but also clarifies the relationship between NICs, processes, ports, software, and vulnerabilities and their respective devices. This makes the categorized asset relationships clearer and facilitates their import into the nodes of the graph database, thus improving the management of the graph database.
[0134] Furthermore, when classifying the asset information set by device, since the asset information set contains asset data collected from various parts of the IT entity, duplicate data may be stored. Therefore, it is necessary to merge duplicate device information.
[0135] Specifically, step A above may include:
[0136] Step E1: Obtain the unique device identification code for each asset in the asset information set.
[0137] The unique identifier here is the Universally Unique Identifier, or UUID. Generally, when storing information, a UUID is generated for each piece of information to facilitate information management and distinguish it from other information. For storing device information, a device UUID is generated.
[0138] Step E2: Merge asset information with the same unique identification code to obtain merged information for each device.
[0139] The presence of identical unique identifiers indicates that multiple pieces of asset (equipment) information pertain to the same asset (equipment), but are stored in different locations. Therefore, unique identifiers can be used to merge identical asset (equipment) information, allowing all information belonging to the same asset (equipment) stored in different locations to be integrated.
[0140] In the specific process of integrating device information, asset information with the same unique device identifier is merged. The merged asset information constitutes the information of a single device. Other information in the asset information besides the device information can be considered as supplementary asset information for the corresponding device information, in preparation for subsequent merging of network interface card, process, port, software, and vulnerability information.
[0141] Step E3: Calculate the similarity of device information with different device unique identifiers or without device unique identifiers, and merge asset information with similarity greater than the preset similarity to obtain the merged device information.
[0142] Different device unique identifiers do not necessarily mean that multiple pieces of device information do not belong to the same device. It is possible that information about the same device was stored in different locations, generating unique identifiers for each. Therefore, further investigation is needed to determine whether multiple pieces of device information with different unique identifiers belong to the same device. Furthermore, it is even more difficult to determine whether device information without a unique identifier belongs to the same device as other device information, and further investigation is required.
[0143] In the process of calculating the similarity of multiple device asset information, the similarity can be calculated based on all or part of the details of each device in the preset device list.
[0144] Specifically, step E3 above may include:
[0145] Step E31: Calculate the similarity of asset information with different device unique identifiers or without device unique identifiers pairwise according to the calculation method of each device detail item in the preset device list and the weight configured for each device detail item.
[0146] The specific content of each device detail item in the preset device list is different, and the similarity calculation method is also different. Furthermore, the importance of each device detail item in measuring device similarity is also different. Therefore, it is necessary to configure the calculation method and weight for each device detail item.
[0147] The specific configurations for each device detail item in the preset device list are shown in the table below.
[0148] Table 7
[0149]
[0150]
[0151]
[0152] Accordingly, after configuring the details of each device in the preset device list, you can continue to configure the calculation method and weight for the details of network cards, processes, ports, software, and vulnerabilities.
[0153] The specific configurations for each network card detail item in the preset network card list are shown in the table below.
[0154] Table 8
[0155]
[0156]
[0157] The specific configurations for each process detail item in the preset process list are shown in the table below.
[0158] Table 9
[0159]
[0160]
[0161] The specific configurations for each port detail item in the preset port list are shown in the table below.
[0162] Table 10
[0163] id int4 none 0 asset_id varchar none 0 port_num varchar Exact match 1 dev_uuid varchar none 0
[0164] The specific configurations for each software detail item in the preset software list are shown in the table below.
[0165] Table 11
[0166]
[0167]
[0168] The specific configurations for each vulnerability detail item in the preset vulnerability list are shown in the table below.
[0169] Table 12
[0170]
[0171]
[0172] By calculating the similarity of device, network card, process, port, software, and vulnerability information according to the calculation method and weight configured for each detail item in Table 7-12 above, information belonging to the same device, network card, process, port, software, and vulnerability can be merged.
[0173] The specific calculation formulas for devices, network cards, processes, ports, software, and vulnerabilities are as follows:
[0174]
[0175] Wherein, similarity refers to the similarity between two devices, network cards, processes, ports, software, or vulnerability information, and score i The weight is the score for each device, network card, process, port, software, or vulnerability detail item. i The weight of each device, network interface card, process, port, software, or vulnerability detail item is denoted by 'i', which represents the weight of each device, network interface card, process, port, software, or vulnerability detail item.
[0176] As can be seen from the above, the unique identification code of the device can be used to merge the device information of the same device, avoiding the duplication of device information. This not only saves the space occupied by device information in each node of the graph database, but also makes the device information of each node in the graph database more concise and clear, which facilitates the management and use of device information in the graph database.
[0177] Furthermore, after obtaining the merged information of each device, a new unique identifier can be generated for each device to distinguish it from the merged information of other devices, and it can continue to be used after being imported into the graph database.
[0178] Specifically, after steps E2 and E3 above, the method may further include:
[0179] Step F1: Generate new unique identifiers for each merged device and add the new unique identifiers to the corresponding device information.
[0180] In other words, a new unique identifier is generated for each of the merged device information entries. Furthermore, a correspondence is established between the new unique identifier and the corresponding device information, so that a specific device information entry can be located using the new unique identifier later.
[0181] Step F2: Establish the correspondence between the new unique identifier and the corresponding device's unique identifier to obtain a unique identifier mapping table, so as to establish connection edges for each node in the graph database according to the unique identifier mapping table.
[0182] In the merged device information, there will also be a unique identifier for that device before classification. This unique identifier has already been associated with the unique identifiers of network cards, processes, ports, software, and vulnerabilities existing in that device. Therefore, it is necessary to establish a new unique identifier and the original unique identifier (i.e., the original unique identifier of the device) again. In this way, after various asset information is imported into the corresponding nodes of the graph database, the original unique identifier of the device can be obtained through the new unique identifier in the node. Through the original unique identifier of the device, it can be known which network cards, processes, ports, software, and vulnerabilities the device is associated with (the details of network cards, processes, ports, software, and vulnerabilities all contain the unique identifier of the device to which they belong), thereby realizing the establishment of the association between the nodes in the graph database.
[0183] The creation of a new unique identifier after merging network cards, processes, ports, software, and vulnerabilities is the same as that for devices, and will not be elaborated here.
[0184] As can be seen from the above, by generating new unique identifiers for each merged device information and establishing a correspondence between the new unique identifiers and the previous unique identifiers of the corresponding device information, we can not only provide convenience for subsequent device information management, but also ensure that the association between devices and various network cards, processes, ports, software and vulnerabilities is not overlooked, thereby improving the accuracy of establishing connection edges between nodes in the graph database.
[0185] Finally, the method for importing asset information into a graph database provided in this application embodiment will be described again from the processes of asset information extraction, fusion, and importing asset information into the graph database.
[0186] Figure 3 This is a schematic diagram illustrating the process of extracting asset information in an embodiment of this application. See also... Figure 3 As shown, asset data is extracted from data sources 1, 2, and 3 respectively, and then integrated (structured data: cleaning, transformation, and mapping to corresponding columns; semi-structured data: regular expression / machine learning algorithms; unstructured data: regular expression / machine learning algorithms) and stored in the data access library. Then, asset data is extracted according to data quality standards (e.g., encoding specifications (unified UTF-8 encoding), underscore specifications (_, —, -, etc.), full-width / half-width constraints, space constraints, invisible character constraints (CR, LF, CRLF, etc.), key field integrity constraints (UUID, foreign keys not null, etc.), case sensitivity constraints, etc.) and its compliance with quality standards is assessed (whether data is missing, whether transcoding errors have occurred, etc.). If the data meets the quality standards, it is stored in the data material library; if it does not meet the quality standards, manual intervention is used to standardize the processing of problematic asset data, determine if it can be repaired, if so, repair it, and store it in the data material library; if it cannot be repaired, it is discarded.
[0187] Figure 4 This is a schematic diagram illustrating the process of integrating device asset information in an embodiment of this application. See [link / reference] Figure 4 As shown, the process queries the data library to see if there is any new device data. If not, the process ends. If there is, the new device asset data is extracted, and duplicate data from the same source (same UUID) is normalized and deduplicated. A new UUID is generated for the duplicate data. Duplicate data from different sources (actually the same device) is also normalized and deduplicated and written into the data fusion table, which is then stored in the data fusion library. A new UUID is generated for the duplicate data and written into the device UUID mapping table (new UUID: original UUID), finally obtaining the device UUID mapping table.
[0188] Figure 5 This is a schematic diagram illustrating the process of integrating network interface card, process, port, software, and vulnerability asset information in an embodiment of this application. See also... Figure 5 As shown, the process queries the data asset library for newly added network interface cards (NICs), processes, ports, software, or vulnerability data. If none are found, the process ends. If so, the newly added NIC, process, port, software, or vulnerability asset data is extracted. Data with the same origin (same UUID) is normalized and deduplicated, and new UUIDs are generated for duplicate data. Data with different origins (actually the same device) is also normalized and deduplicated. Then, the device UUID mapping table is read to find the UUID of the device to which the NIC, process, port, software, or vulnerability belongs. The device UUID is updated to the latest UUID. The NIC, process, port, software, or vulnerability asset data is grouped according to the device UUID, and normalized within each group. This data is then written into the NIC, process, port, software, or vulnerability UUID mapping table, resulting in the NIC, process, port, software, or vulnerability UUID mapping table. Finally, the data is written into the data fusion library, ending the process.
[0189] Figure 6 This is a schematic diagram illustrating the process of importing asset information into the graph database in an embodiment of this application. See [link / reference] Figure 6 As shown, the process retrieves a UUID mapping table for devices, network interface cards (NICs), processes, ports, software, and vulnerabilities from the data fusion library, and extracts node data (information on each device, NIC, process, port, software, and vulnerability) and relationship data (the associations between devices, NICs, processes, ports, software, and vulnerabilities). The extracted node data is then imported into the corresponding nodes in the graph database. The corresponding normalized UUID is then found based on the UUID mapping table for devices, NICs, processes, ports, software, and vulnerabilities, and imported into the graph database along with the extracted relationship data, establishing connections between the nodes in the graph database.
[0190] This concludes the description of the method for importing asset information into a graph database provided in this application.
[0191] Based on the same inventive concept, as an implementation of the above method, this application also provides an apparatus for importing asset information into a graph database. Figure 7 This is a schematic diagram of the device for importing asset information into the graph database in an embodiment of this application. See also... Figure 7 As shown, the device may include:
[0192] The receiving module 701 is used to receive asset data of the IT entity and obtain the asset information set of the IT entity;
[0193] The classification module 702 is used to classify the asset information set according to the basic classification units in the preset classification framework to obtain the classified asset information corresponding to the basic classification units. The basic classification units in the preset classification framework are indivisible classification units determined based on the asset classification needs of multiple IT entities, and the basic classification units correspond to the nodes in the graph database.
[0194] The node import module 703 is used to import the classified asset information into the corresponding nodes of the graph database according to the basic classification units;
[0195] The connection edge import module 704 is used to determine the association relationship between each node based on the relationship between the asset data of the IT entity, and to create connection edges between the corresponding nodes based on the association relationship.
[0196] In this embodiment of the application, the basic classification unit includes: device, network card, process, port, software, and vulnerability. The classification module is specifically used to classify the asset information set according to device, network card, process, port, software, and vulnerability to obtain information on each device, each network card, each process, each port, each software, and each vulnerability.
[0197] In this embodiment, the device further includes: a subdivision module, configured to classify each device information according to each device detail item in a preset device list to obtain a device list, wherein each device detail item in the preset device list is determined based on detail information of multiple devices; classify each network interface card (NIC) information according to each NIC detail item in a preset NIC list to obtain a NIC list, wherein each NIC detail item in the preset NIC list is determined based on detail information of multiple NICs; classify each process information according to each process detail item in a preset process list to obtain a process list, wherein each process detail item in the preset process list is based on... Detailed information of multiple processes is determined; each port information is classified according to the port details in a preset port list to obtain a port list, wherein each port details in the preset port list is determined based on the detailed information of multiple ports; each software information is classified according to the software details in a preset software list to obtain a software list, wherein each software details in the preset software list is determined based on the detailed information of multiple software; each vulnerability information is classified according to the vulnerability details in a preset vulnerability list to obtain a vulnerability list, wherein each vulnerability details in the preset vulnerability list is determined based on the detailed information of multiple vulnerabilities.
[0198] In this embodiment of the application, the device further includes: an establishment module, used to acquire full information of multiple IT entities, the full information including asset data; classifying the full information to obtain category information, the category information including details of each device, each network card, each process, each port, each software, and each vulnerability in a preset device list.
[0199] In this embodiment of the application, the establishment module is specifically used to acquire all information related to multiple IT entities, including national standards, product information, industry specifications, asset scanning devices, traffic probes, asset management databases, and application scenarios of multiple IT entities, as the full information of multiple IT entities.
[0200] In this embodiment of the application, the classification module is specifically used to classify the asset information set according to different devices to obtain information about each device and its corresponding associated asset information; and to classify the associated asset information of each device according to different network cards, different processes, different ports, different software and different vulnerabilities to obtain information about each network card, each process, each port, each software and each vulnerability for each device.
[0201] In this embodiment of the application, the classification module is specifically used to obtain the unique device identification code of each asset information in the asset information set; merge asset information with the same unique device identification code to obtain merged device information; perform similarity calculation on device information with different unique device identification codes or no unique device identification code, and merge asset information with similarity greater than a preset similarity to obtain merged device information.
[0202] In this embodiment of the application, the classification module is specifically used to perform similarity calculations on device information with different device unique identifiers or without device unique identifiers, according to the corresponding calculation method of each device detail item in the preset device list and the weight configured for each device detail item.
[0203] In this embodiment of the application, the calculation method includes: calculating the Jaccard distance, performing a perfect match, generating a TF-IDF vector, and calculating the cosine distance.
[0204] In this embodiment of the application, the device further includes: a generation module, used to generate new unique identifiers for each merged device, and add the new unique identifiers to the corresponding device information; establish a correspondence between the new unique identifiers and the unique identifiers of the corresponding devices to obtain a unique identifier mapping table, so as to establish connection edges for each node in the graph database according to the unique identifier mapping table.
[0205] It should be noted that the description of the above device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0206] Based on the same inventive concept, embodiments of this application also provide an electronic device. Figure 8 This is a schematic diagram of the electronic device in an embodiment of this application. See also... Figure 8 As shown, the electronic device may include: a processor 801, a memory 802, and a bus 803; wherein the processor 801 and the memory 802 communicate with each other through the bus 803; the processor 801 is used to call program instructions in the memory 802 to execute the methods in one or more of the above embodiments.
[0207] It should be noted that the descriptions of the above electronic device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the electronic device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0208] Based on the same inventive concept, embodiments of this application also provide a computer-readable storage medium, which may include: a stored program; wherein, when the program is running, it controls the device where the storage medium is located to execute the methods in one or more of the above embodiments.
[0209] It should be noted that the descriptions of the storage medium embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0210] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for importing asset information into a graph database, characterized in that, The method includes: Receive asset data from the IT entity to obtain a set of asset information for the IT entity; The asset information set is classified according to the basic classification units in the preset classification framework to obtain the classified asset information corresponding to the basic classification units. The basic classification units in the preset classification framework are indivisible classification units determined based on the asset classification needs of multiple IT entities. The basic classification units correspond to nodes in the graph database. When classifying, the basic classification units in the preset classification framework that cannot be covered are ignored. The classified asset information is imported into the corresponding nodes of the graph database according to the basic classification units; Based on the relationships between the asset data of the IT entity, determine the association relationships between the nodes, and create connection edges between the corresponding nodes based on the association relationships; The basic classification units include: devices, network interface cards (NICs), processes, ports, software, and vulnerabilities. The process of classifying the asset information set according to the basic classification units in the preset classification framework to obtain the classified asset information corresponding to each basic classification unit includes: The asset information set is classified according to device, network card, process, port, software and vulnerability to obtain information on each device, network card, process, port, software and vulnerability. The process of classifying the asset information set according to device, network interface card (NIC), process, port, software, and vulnerability to obtain information on each device, NIC, process, port, software, and vulnerability includes: Obtain the unique device identification code for each asset in the asset information set; Merge asset information with the same unique identification code to obtain merged information for each device. The similarity of equipment information with different or no unique equipment identification codes is calculated, and asset information with similarity greater than the preset similarity is merged to obtain the merged equipment information. The step of calculating the similarity of device information with different or non-existent unique device identifiers includes: Device information with different device unique identifiers or without device unique identifiers is used to calculate similarity pairwise according to the calculation method of each device detail item in the preset device list and the weight configured for each device detail item; Each detail item in the device information, network card information, process information, port information, software information, and vulnerability information is configured with its own corresponding similarity calculation method and weight; the calculation method includes: calculating Jaccard distance, complete matching, and generating TF-IDF vectors and calculating cosine distance.
2. The method according to claim 1, characterized in that, After obtaining information on each device, each network card, each process, each port, each software, and each vulnerability, the method further includes: Each device information is categorized according to the device details items in the preset device list to obtain each device list, wherein each device details item in the preset device list is determined based on the details information of multiple devices; Each network interface card (NIC) information is categorized according to the details of each NIC in a preset NIC list to obtain a list of each NIC. The details of each NIC in the preset NIC list are determined based on the details of multiple NICs. Each process information is categorized according to the process details items in the preset process list to obtain a process list, wherein each process details item in the preset process list is determined based on the details information of multiple processes; Each port information is categorized according to the port details in the preset port list to obtain a port list, wherein each port details in the preset port list is determined based on the details of multiple ports; Each piece of software information is categorized according to the software details items in a preset software list to obtain each software list, wherein each software details item in the preset software list is determined based on the details information of multiple software programs; Each vulnerability information is categorized according to the vulnerability details in a preset vulnerability list to obtain a vulnerability list. The vulnerability details in the preset vulnerability list are determined based on the details of multiple vulnerabilities.
3. The method according to claim 2, characterized in that, Before categorizing each device information according to the device details items in the preset device list, the method further includes: Obtain full information from multiple IT entities, including asset data; The full set of information is categorized to obtain category information, which includes details of each device, network card, process, port, software, and vulnerability in the preset device list.
4. The method according to claim 3, characterized in that, The acquisition of full information from multiple IT entities includes: Acquire all information from multiple IT entities, including national standards, product data, industry specifications, asset scanning devices, traffic probes, asset management databases, and application scenarios, as comprehensive information for all IT entities.
5. The method according to any one of claims 1 to 4, characterized in that, The asset information set is categorized according to device, network interface card (NIC), process, port, software, and vulnerability to obtain information on each device, NIC, process, port, software, and vulnerability, including: The asset information set is categorized according to equipment to obtain information on each piece of equipment and its corresponding ancillary asset information; The corresponding ancillary asset information of each device is classified according to network card, process, port, software and vulnerability, resulting in the corresponding network card information, process information, port information, software information and vulnerability information for each device.
6. The method according to claim 1, characterized in that, After obtaining the merged information of each device, the method further includes: Generate a new unique identifier for each merged device, and add the new unique identifier to the corresponding device information; Establish a correspondence between the new unique identifier and the unique identifier of the corresponding device to obtain a unique identifier mapping table, so as to establish connection edges for each node in the graph database according to the unique identifier mapping table.
7. An apparatus for importing asset information into a graph database, characterized in that, The device includes: The receiving module is used to receive asset data from the IT entity and obtain a set of asset information of the IT entity. The classification module is used to classify the asset information set according to the basic classification units in the preset classification framework to obtain the classified asset information corresponding to the basic classification units. The basic classification units in the preset classification framework are indivisible classification units determined based on the asset classification needs of multiple IT entities. The basic classification units correspond to nodes in the graph database. When classifying, the basic classification units in the preset classification framework that cannot be covered are ignored. The node import module is used to import the classified asset information into the corresponding nodes of the graph database according to the basic classification units; The connection edge import module is used to determine the association relationship between each node based on the relationship between the asset data of the IT entity, and to create connection edges between the corresponding nodes based on the association relationship; The basic classification units include: devices, network cards, processes, ports, software, and vulnerabilities. The classification module is specifically used to classify the asset information set according to devices, network cards, processes, ports, software, and vulnerabilities to obtain information on each device, each network card, each process, each port, each software, and each vulnerability. Specifically, the classification module is used to obtain the unique device identification code of each asset information in the asset information set; merge asset information with the same unique device identification code to obtain merged device information; perform similarity calculation on device information with different unique device identification codes or no unique device identification code, and merge asset information with similarity greater than a preset similarity to obtain merged device information. Specifically, the classification module is used to calculate the similarity between device information with different device unique identification codes or without device unique identification codes, according to the calculation method of each device detail item in the preset device list and the weight configured for each device detail item; Each detail item in the device information, network card information, process information, port information, software information, and vulnerability information is configured with its own corresponding similarity calculation method and weight; the calculation method includes: calculating Jaccard distance, complete matching, and generating TF-IDF vectors and calculating cosine distance.
8. An electronic device, characterized in that, The electronic device includes: a processor, a memory, and a bus; wherein the processor and the memory communicate with each other via the bus; the processor is used to call program instructions in the memory to execute the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The storage medium includes: a stored program; wherein, when the program is executed, it controls the device where the storage medium is located to perform the method as described in any one of claims 1 to 6.