Method for constructing knowledge graph based on internet surveying data and model display method
By acquiring, classifying, encoding, and integrating Internet mapping data, a knowledge graph is constructed and interactively displayed, solving the problem of multi-level association modeling and dynamic updating of Internet assets, and realizing efficient and accurate network asset management and analysis.
Patent Information
- Application Number
- CN202510050988.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Existing technologies lack the ability to model multi-level relationships between internet assets, have not formed a unified construction and display process, lack semantic analysis and dynamic update capabilities, and have limited display functions and insufficient applicability.
By acquiring data from multiple internet mapping platforms, classifying and encoding the data, standardizing and integrating it, constructing a knowledge graph, and employing interactive graphical display technology, we can achieve this.
It enables efficient and accurate processing and multi-dimensional display of network asset data, supports dynamic updates and interactive analysis, and improves the system's scalability and user experience.
Smart Images

Figure CN119962647B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the application relates to the technical field of Internet mapping, and particularly relates to a knowledge graph construction and model display method based on Internet mapping data. BACKGROUND
[0002] With the rapid expansion of the Internet, collecting and constructing network models of Internet assets (such as IP addresses, ports, services, etc.) through mapping technology has become the basis for network management and security protection. On this basis, knowledge graph technology has been gradually applied to the efficient organization and analysis of network resources.
[0003] However, the existing knowledge graph of network assets is mostly a simple model based on logical topology. For example, some methods obtain Internet asset data based on public mapping platforms (such as Shodan and Censys), classify asset information using rules, and generate a logical connection relationship graph. These methods are usually limited to the construction of physical or logical connections, and lack the mining of semantic associations between data.
[0004] In addition, current knowledge graph technology is mainly concentrated in the fields of search engines and intelligent recommendations, and has few applications in network asset management. Existing attempts have mostly focused on single scenarios and have not constructed a complete graph model combined with Internet mapping data, nor have they formed a systematic construction scheme. SUMMARY
[0005] The embodiment of the application provides a knowledge graph construction and model display method based on Internet mapping data to solve the above problems.
[0006] In a first aspect, the embodiment of the application provides a knowledge graph construction and model display method based on Internet mapping data, comprising:
[0007] Obtaining mapping asset data from multiple Internet mapping platforms;
[0008] Classifying and encoding the mapping asset data;
[0009] Standardizing and fusing various data from different platforms;
[0010] Based on the standardized data, constructing a knowledge graph of Internet assets;
[0011] Using interactive graphical display technology to display the knowledge graph to the user.
[0012] In a second aspect, the embodiment of the application provides an electronic device, comprising:
[0013] One or more processors;
[0014] a memory storing one or more programs,
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for constructing and displaying a knowledge graph based on Internet mapping data according to any one of the embodiments.
[0016] In a third aspect, the embodiments of the present application also provide a computer readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for constructing and displaying a knowledge graph based on Internet mapping data according to any one of the embodiments.
[0017] In summary, the embodiments of the present application provide a method for constructing and displaying a knowledge graph based on Internet mapping data, which aims to construct a standardized mapping data processing and analysis process, from data collection to data display, each link is designed in a process, so that the processing of network asset data is more efficient, accurate, and has certain scalability. The method can achieve the following
[0018] Advantages:
[0019] 1. The standardized process improves the efficiency and consistency of data processing. By formulating a standardized process, the present application standardizes the processing of Internet mapping data from different sources. This process includes data collection, cleaning, coding, fusion, standardization, modeling, etc. steps, ensuring that each link can follow a unified standard, thereby improving the consistency and efficiency of data processing.
[0020] 2. Simplifies the data processing process and reduces the complexity of operation. Traditional mapping data processing usually requires complex data cleaning, format conversion and feature selection operations, and the process is complex and requires high manual intervention. The present application simplifies the operation steps by standardizing the data processing process, so that the operator can work according to the process steps, reducing manual intervention.
[0021] 3. Ensures the accuracy and consistency of data fusion. Internet mapping data from different sources usually has differences in format and structure. The present application can efficiently unify heterogeneous data through a standardized data fusion process, ensuring the accuracy of data fusion.
[0022] 4. Improve the scalability and flexibility of the system. The standardized process of the embodiments of the present application not only applies to the current Internet mapping data, but also can easily adapt to changes or expansion of future data sources. Through a unified processing framework, the system can flexibly access new data sources and process them, supporting the construction and display of large-scale network topology data.
[0023] 5. The interactivity and user experience of network topology display are improved. Through standardized and streamlined design, the data output by the embodiment of the present application can be used to build visual network topology display. Through interactive display, users can conveniently view and analyze network asset data.
[0024] 6. The timeliness and real-time updating capability of data are enhanced. Through standardized data processing flow, especially in data updating and incremental processing, the embodiment of the present application can timely reflect the changes of network environment and automatically update surveying and mapping data. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0026] Figure 1 is a flowchart of a knowledge graph construction and model display method based on Internet surveying and mapping data provided by the embodiment of the present application;
[0027] Figure 2 is a flow system diagram of a knowledge graph construction based on multi-source surveying and mapping data provided by the embodiment of the present application;
[0028] Figure 3 is a network asset data coding mode schematic diagram provided by the embodiment of the present application;
[0029] Figure 4 is a network asset knowledge graph data model schematic diagram provided by the embodiment of the present application;
[0030] Figure 5 is a schematic diagram of a standardized data template provided by the embodiment of the present application;
[0031] Figure 6 is a fusion framework schematic diagram based on multi-source surveying and mapping data provided by the embodiment of the present application;
[0032] Figure 7 is a structural schematic diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the scope protected by the present application.
[0034] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0035] In the description of the present application, it should be noted that, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connection" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0036] The embodiment of the present application provides a knowledge graph construction and model display method based on Internet surveying and mapping data. In order to illustrate the method, first, the deficiencies of the prior art in constructing a knowledge graph and model display based on Internet surveying and mapping data are described in more detail:
[0037] 1. Lack of multi-level association modeling capability
[0038] The logical network topology modeling is limited to nodes and their simple connections, without deeply mining the semantic association between nodes, resulting in single dimension of model information and insufficient semantic and relationship analysis between data.
[0039] 2. No unified construction and display process
[0040] The collection, fusion and analysis display of surveying and mapping data are usually processed separately, lacking a unified process framework, resulting in complex and inefficient technical application.
[0041] 3. Insufficient semantic analysis and dynamic updating capability
[0042] Existing solutions are mostly based on static data to build simple models, lack of support for dynamic data update and semantic analysis, and cannot realize real-time update and analysis of the correlation.
[0043] 4. The display function is single and has insufficient applicability
[0044] The display of logical topology is mostly static graphics, which is difficult to express the complexity of semantic relationship, and does not support multi-dimensional visual analysis and interactive operation.
[0045] In view of the above shortcomings, Figure 1 is a flowchart of a knowledge graph construction and model display method based on Internet mapping data provided by the embodiments of the present application. The method aims to form a standardized process running through data collection, fusion, modeling, knowledge graph construction to model dynamic display, and improve the analysis depth and display effect of Internet asset data. The method is executed by an electronic device, such as Figure 1 as shown, specifically comprising the following steps:
[0046] S110, acquiring mapping asset data from a plurality of Internet mapping platforms.
[0047] In combination with Figure 2 , this step corresponds to the data collection module, which automatically collects mapping asset data on the Internet through Internet mapping platforms (such as Shodan, Censys, Daydaymap, etc.), mainly including device IP address, port, operating system, service, application, device, location, autonomous domain, vulnerability, etc. Information. The collected data will be preprocessed according to the predefined rules to remove invalid data, and the data will be formatted to prepare for subsequent processing.
[0048] In a specific embodiment, the data collection is based on the API interface of the Internet mapping platform, and the network asset information is periodically grabbed by automatic script. The collected data includes device IP, open port, service type and other information, which are standardized and stored in the database for storage. Specifically, the public mapping platform interface (such as Shodan, Censys) is called to obtain the original data, and the ETL tool is used to perform data deduplication, attribute completion and formatting, etc. The result is stored in the database for subsequent processing.
[0049] S120, data classification and coding of the mapping asset data.
[0050] In combination with Figure 2 , this step corresponds to the data classification and coding module, and the main purpose of this module is to further classify, label and code the mapping asset data obtained from the data collection module. This process can provide standardized input for subsequent asset management, data analysis, mining and application. The specific functions include:
[0051] Data Classification: The collected mapping asset data is grouped or labeled according to predetermined classification criteria, such as device type (router, switch, server, etc.), service type (HTTP, FTP, SSH, etc.), geographic location (country, city, etc.), etc.
[0052] Encoding: By assigning a unique identifier to each type of data (such as device code, service code, etc.), the data can be quickly identified and located in subsequent processing and analysis.
[0053] Labeling and Tagging: According to the characteristics of the device, relevant labels or annotation information (such as the operator of the device, the protocol used, etc.) are generated to facilitate filtering, aggregation or modeling during data analysis.
[0054] In a specific embodiment, asset data can be classified according to network resource classification and coding principles and the hierarchical characteristics of the elements themselves; the coding refers to the coding method in the geographic information element classification and coding (GB / T 2013923-2022) to extend network asset information element coding, such as Figure 3 .
[0055] For example, based on the network space mapping data model structure of the current mainstream network mapping platform (censys, shodan, daydaymap, etc.), ontology modeling can be performed to construct a general ontology model, such as Figure 4 , from which multiple data types are extracted, including ASN, location, device, service, operating system and vulnerability; according to the characteristics of network assets, each data type is subdivided into multiple levels; according to each data type and the subdivided type, each data in the mapping asset data is encoded; and other information in each data except the encoding information is added to the encoded data as a label.
[0056] S130, standardize and fuse various types of data from different platforms.
[0057] In combination with Figure 2 , this step corresponds to the data fusion and standardization module, which converts the collected Internet mapping data from different sources into a unified format and cleans the data to ensure the uniformity and accuracy of the data. The specific steps include:
[0058] Data Standardization: Through standardization algorithms (such as JSON, CSV, etc.), the data fields and units are unified.
[0059] Data deduplication and consistency processing: merge duplicate node information, eliminate redundant data, and ensure the uniqueness of node information in the knowledge graph.
[0060] Multi-source data fusion: match and fuse data from different mapping platforms, match data through a matching rule engine, and generate a complete set of network asset information. The matching rule engine defines the mode of data matching and the solution to conflicts.
[0061] In a specific embodiment, for the collected data from different platforms, a rule-based matching algorithm can be used for data fusion. The implementation process includes: defining a standardized format template for the target data (such as Figure 5 shown); processing data from different mapping platforms into a unified format through a standardization algorithm; using a matching rule engine to fuse different mapping data into a set of standard data; storing the fusion result data in a database to support subsequent graph construction.
[0062] S140, based on the standardized data, constructing a knowledge graph of Internet assets.
[0063] In combination Figure 2 , this step corresponds to the knowledge graph construction module. Based on data fusion, a graph database or graph structure data model is used to model Internet assets and construct a knowledge graph. The core idea of the knowledge graph is to construct a network structure through nodes and edges.
[0064] Node modeling: different categories of data are represented by different nodes according to data classification and coding.
[0065] Relationship modeling: through the analysis of the relevance of data, the logical connections or semantic associations between different nodes (such as IP and service, device and operating system, etc.) are identified, and the associations between them are represented by edges in the graph.
[0066] Graph expansion and update: supports dynamic update function, can update the graph incrementally according to new mapping data or network topology changes.
[0067] In a specific embodiment, the construction of the knowledge graph includes: data model construction based on fusion result data; the categories of graph nodes are based on data classification and coding, and different categories are modeled by nodes; relationship modeling according to the relationship between nodes; defining a graph database management data interface to support the addition, deletion, modification and query of graph model data; using an incremental update interface to regularly update graph data, so that it can reflect the latest data of the network.
[0068] S150, using interactive graphical display technology to display the knowledge graph to the user.
[0069] In combination Figure 2 , this step corresponds to the visualization display and analysis module, which uses interactive graphical display technology to display the constructed knowledge graph to the user through a graphical interface. This display module supports:
[0070] Visualization of nodes and edges: Each network asset node and its associated edges are presented in a graphical form, and users can scale, drag, and rotate the graph as needed to visually view the network structure.
[0071] Multi-dimensional display: Users can choose different types of nodes for topology display to analyze in different dimensions.
[0072] Relationship analysis: Through graph analysis technology, key nodes and high-risk connection paths in the network are extracted to help users identify potential risk points in the network.
[0073] In a specific embodiment, the constructed knowledge graph can be displayed through a visualization tool (such as Gephi, D3.js), supporting node scaling, dragging, and multi-dimensional attribute display. Through the graphical interface, users can conveniently view the entire network structure and conduct in-depth analysis. The implementation process includes: using Gephi as the front-end module of the graph, defining the graph database query interface, supporting data query results to be loaded into Gephi; using the visualization and analysis capabilities of Gephi to display and analyze the graph; extending the related functions of Gephi to support more flexible custom operations.
[0074] The entire process described above can also be understood in combination with the fusion framework shown in Figure 6
[0075] In summary, the embodiment of the present application provides a knowledge graph construction and model display method based on Internet surveying and mapping data, which can achieve the following beneficial effects:
[0076] 1. A knowledge graph construction method based on Internet surveying and mapping data is proposed. Traditional network topology modeling methods usually only rely on static physical or logical connection information, ignoring the deep semantic association between network assets. The embodiment of the present application fuses Internet surveying and mapping data (such as IP address, port, service, device type, etc.) and combines knowledge graph technology to construct a knowledge graph model that can deeply mine the association between network assets. This model not only describes the simple connection between nodes, but also carries the semantic relationship (such as device address, device vulnerability, etc.) between nodes through the relationship edge. Compared with traditional topology graphs, the knowledge graph of the present application can more comprehensively and dynamically express the multi-dimensional relationship between network assets, providing more abundant information sources for subsequent network analysis, optimization, and security detection.
[0077] 2. A unified network mapping data graph construction process and framework is provided. Most network topology construction schemes are currently handled separately, lacking a unified process system. The embodiments of the present application propose a unified process system from data collection, cleaning, classification and coding, standardization, fusion, graph construction to visualization, ensuring seamless connection and efficient execution between each link. This standardized process not only simplifies data processing and graph construction, but also improves the efficiency and maintainability of the entire system, and is suitable for various Internet mapping data sources and network analysis needs.
[0078] 3. Multi-source data fusion and standardization technology is provided. In Internet mapping data, the data formats and structures from different platforms (such as Shodan, Censys, etc.) differ greatly, bringing great challenges to data processing. The embodiments of the present application propose a fusion matching algorithm based on self-defined conflict resolution, which automatically matches and integrates information from different data sources through a rule engine, and unifies the data format to ensure data consistency and accuracy; in addition, the embodiments of the present application also propose a set of standardized template data to maximize the compatibility of data from different mapping platforms. This technology can effectively solve the heterogeneity between data from different mapping platforms, ensure that the constructed knowledge graph data is complete and has no redundancy, and improve the system's processing capacity for diversified data.
[0079] 4. A unified network mapping data classification and coding system is provided. Different mapping platforms and tools use different classification standards and coding methods, resulting in confusion and duplication of data in classification, retrieval and analysis, making it difficult to form a unified knowledge system. The embodiments of the present application propose a unified classification and coding system based on the classification and coding principles of network resources and the hierarchical structure characteristics of the elements themselves, combining multi-level labels and multi-dimensional attributes to classify and label network mapping data. By formulating strict coding rules and classification standards, combining automatic classification algorithms and manual review mechanisms, mapping data from different platforms is classified and arranged and uniquely identified, ensuring clear classification logic and unified coding rules. This system can significantly improve the organization and management efficiency of network mapping data, avoiding the data integration obstacles caused by inconsistent classification standards of different platforms. Through unique coding, the independence and traceability of data nodes in the knowledge graph are ensured, while reducing the existence of duplicate data and redundant information, improving the data retrieval and analysis performance of the system.
[0080] 5. This paper proposes a knowledge graph model that supports dynamic updates and incremental construction. Traditional network topology diagrams are mostly static displays, failing to reflect changes in the network environment in a timely manner. Changes in network assets (such as the addition, removal, and status changes of devices) can cause the topology diagram to become outdated, affecting the accuracy of network management. This application proposes a knowledge graph mechanism that supports dynamic updates and incremental construction. Based on incremental data processing technology, the system can automatically update the knowledge graph according to new mapping data or network changes, ensuring the timeliness of the graph. This dynamic update function allows the knowledge graph to reflect changes in the network environment in real time, ensuring that the management and monitoring of network assets are always consistent with the actual situation, making it particularly suitable for the rapidly changing Internet environment.
[0081] 6. This application provides a multi-dimensional visualization and interactive analysis technology for network topology maps. Traditional network topology map display methods often lack flexibility and cannot effectively present complex network asset relationships. This application's embodiments employ advanced graphical visualization technologies (D3.js, Gephi, etc.) to enable users to view network topology based on different dimensions (such as device type, service type, geographical location, etc.) through interactive graphical display technology. Furthermore, it supports dynamic interaction (such as zooming, dragging, filtering, annotation, etc.), allowing users to easily explore key nodes, paths, or potential risks in the network. This technology not only improves the operability of network data display but also enables users to deeply analyze and explore complex network relationships, improving the accuracy and efficiency of network asset management.
[0082] Furthermore, based on the above embodiments, this embodiment refines the data standardization and fusion process. Optionally, the standardization and fusion of various types of data from different platforms may include the following steps:
[0083] Step 1: Extract multiple data types from the data ontology model used to build the knowledge graph. For example... Figure 4 As shown, the various data types include ASN, location, device, service, operating system, and vulnerability.
[0084] Step 2: Based on the constructed data model, design a unified data mapping template, such as... Figure 5 As shown, data from different sources is converted into a standard format. Specifically, fields are mapped according to the data format defined in the data template to ensure consistency in field naming, data type, and data unit. Additionally, auxiliary databases, such as IP address databases, device fingerprint databases, and vulnerability databases, are used to complete the data information. The IP address database is used to complete the model's location information; the device fingerprint database and matching rule algorithms are used to complete the model's device information; and the vulnerability database data is used to complete the model's CVES information.
[0085] Step three, construct a three-dimensional measurement matrix with each data type, mapping platform and data quality evaluation index as the dimension. After standardization of data of different mapping platforms, all contain information of the above table dimensions: including autonomous system (ASN) information, location (Location) information, device (Device) information, service / component (Sevices) information, operating system (OS) information, vulnerability (CVES) information. This embodiment is aimed at the above data types, and the quality of data records is comprehensively judged from five aspects of data reliability, data accuracy, data integrity, data consistency and data timeliness. Specifically, the process can include the following steps:
[0086] First, a three-dimensional measurement empty matrix is constructed with each data type, mapping platform and data quality evaluation index as the dimension. Exemplarily, the three-dimensional measurement matrix takes the mapping platform as one dimension, and each mapping platform includes a two-dimensional measurement table as shown in Table 1:
[0087] Table 1
[0088] Data Type Reliability Accuracy Integrity Consistency Timeliness Confidence Value ASN Location Device Sevices OS CVES
[0089] After obtaining the above three-dimensional measurement matrix, each element in the table can be valued according to experience, or the following method can be used to value the elements:
[0090] Step A, according to the usage rate of each mapping platform, an element value is given to each element of the mapping platform. Optionally, according to the usage rate of each mapping platform, each element in the mapping platform is initially valued, wherein the higher the usage rate of the mapping platform, the larger the initial element value of all elements in Table 1 corresponding to the platform.
[0091] Step B, match the ASN and location from each mapping platform with the known data in the existing address library, and according to the matching rate, another element value is given to the accuracy dimension of the ASN and location of each mapping platform. For the two data types of ASN and Location, the accuracy of the two can be calculated according to the existing IP address library. The larger the proportion of the number of ASN information consistent with the existing address library in the total data amount in a certain platform, the higher the accuracy of the ASN corresponding to the platform, and the Location is similar.
[0092] Step C, match the device from each mapping platform with the known data in the existing fingerprint library, and give another element value to the accuracy dimension of the device of each mapping platform according to the matching rate. For the Device data type, the accuracy of the Device can be calculated according to the existing fingerprint library. The greater the proportion of the number of Device information consistent with the Device information in the existing fingerprint library in the total data amount in a certain platform, the higher the accuracy corresponding to the Device of the platform.
[0093] Step D, give another element value to the integrity dimension of each data type of each mapping platform according to the number of overlapping fields between each data type data from each mapping platform and the data standardization template. In this embodiment, the completeness of the data is judged by the fields in the data standardization template. For the same data type (such as Device), the more the number of repeated fields included in the Device of a certain platform and the fields in the standardization template, the higher the completeness corresponding to the Device of the platform. The other data types are similar.
[0094] In order to distinguish and describe, the element value given according to the use rate of the mapping platform in step A is referred to as the first element value, the element value given for accuracy in steps B and C is referred to as the second element value, and the element value given for completeness in step D is referred to as the third element value. The remaining data quality indicators can also be given another element value other than the first element value according to the calculation rule, which is referred to as the fourth element value. For example, consistency can be assigned by cross-validation of different mapping platforms, and timeliness can be assigned by the update date of the data.
[0095] Step E, calculate the final value of each element in the three-dimensional measurement matrix according to the first element value, the second element value and the third element value. Optionally, for accuracy, multiply the first element value by the second element value to obtain the final element value; for completeness, multiply the first element value by the third element value to obtain the final element value; and for the remaining three data evaluation quality, multiply the first element value by the corresponding fourth element value to obtain the final element value.
[0096] Step four, obtain the confidence value of each data type according to the three-dimensional measurement matrix. After the above calculation, the data indicator element of each data type of each platform is obtained by weighted average, and the confidence value of each data type is obtained, that is, the confidence analysis is completed.
[0097] Step five, according to the confidence value of each data type, build the fusion mode configuration template of each data type, wherein the fusion mode in the template includes confidence priority, aggregation, time priority, and priority of a certain mapping platform. Specifically, the data fusion process focuses on the processing of data conflicts. Therefore, this embodiment establishes a set of data fusion solutions to solve conflicts, which includes the following fusion methods:
[0098] 1. Confidence priority. The above confidence analysis has been analyzed through multiple data quality indicators, and different weights are given to different dimensions to realize the confidence scoring of different object attributes of the data model. Therefore, the attribute object with high confidence can be selected as the target value for processing. In addition, if the user has his own standard for data confidence, he can also modify the existing confidence calculation result or directly give a value. We also choose the best one.
[0099] 2. Aggregation. In the aggregation method, the data is not selected, but the objects from different sources are aggregated, that is, the fields from different sources are combined into one data, and the repeated information is removed. Users can also configure whether to perform aggregation processing. For example, for port information from different platforms, they may all exist, so the aggregation processing method can be selected.
[0100] 3. Time priority. The time priority principle states that data objects from different platforms can be selected according to the update time of the record, and the most recently updated record is selected as the final result data. Because the timeliness of Internet data is strong, the device IP may have been replaced with another IP a long time ago.
[0101] 4. Select a certain platform priority. If the user is familiar with data from different platforms, he can configure the required platform data as the final result based on experience.
[0102] Since a single fusion method may not be suitable for all data conflict processing, this embodiment builds a fusion mode configuration template, as shown in Table 2, which can configure different conflict processing methods for different data types by assigning values in the template to achieve the best fusion result.
[0103] Table 2 Custom fusion mode configuration template
[0104]
[0105] In a specific embodiment, when assigning values to the above template, the data types can be classified again according to the confidence distribution law of the same data type in different mapping platforms; and then according to the classification results, build the fusion mode configuration template of each data type.
[0106] Specifically, the standard deviation of the confidence values of the same data type from different mapping platforms can be calculated first, and the data type is divided into two categories of aggregated fusion and non-aggregated fusion according to the size of the standard deviation. Taking ASN as an example, the standard deviation of the confidence values of ASN under different mapping platforms is calculated, if the standard deviation is greater than a set threshold, ASN is divided into the aggregated fusion category, otherwise, ASN is divided into the non-aggregated fusion category, and the rest of the data types are similar. For the aggregated fusion category data, the fusion method is set to aggregated fusion, and for the non-aggregated fusion category data, the next step is entered.
[0107] In the next step, according to the significance of the important data quality indicators of the same data type in the three-dimensional measurement matrix in the dimension of different mapping platforms, the data type of the non-aggregated fusion category is divided into a certain mapping platform priority and a confidence value priority. Taking Device as an example, assuming that it is judged to be a non-aggregated fusion category, it is judged whether the accuracy and integrity of its important data indicators (i.e. high weight in weighted average) are significantly better than the performance of other platforms, if so, Device is configured to be platform priority; otherwise, Device is configured to be confidence value priority.
[0108] Step six, according to the fusion method of each data type, configure the template, and fuse each type of data from different forms. Exemplarily, finally for the network device (Device) information in the data model, a certain platform is selected as priority; for the Services object, it is selected by default through the confidence value; and for the vulnerability information, it is selected by the principle of time priority.
[0109] Figure 7 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1. Figure 7 As shown in the figure, the device includes a processor 60, a memory 61, an input device 62 and an output device 63; the number of processors 60 in the device can be one or more, Figure 7 The processor 60 in the device, the memory 61, the input device 62 and the output device 63 can be connected through a bus or other means, Figure 7 Taking connection through a bus as an example.
[0110] The memory 61 is a kind of computer readable storage medium, which can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the knowledge graph construction and model display method based on internet mapping data in the embodiment of the present application. The processor 60 executes the software program, instruction and module stored in the memory 61, thereby performing various functions and data processing of the device, i.e. implementing the above-mentioned knowledge graph construction and model display method based on internet mapping data.
[0111] The memory 61 can include a program storage area that stores an operating system, application programs required for at least one function, and a data storage area that stores data created according to the use of the terminal, etc. In addition, the memory 61 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-volatile solid state memory device. In some examples, the memory 61 can further include a memory disposed remotely with respect to the processor 60, which can be connected to the device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0112] The input device 62 can be used to receive input digital or character information, and to generate key signal input related to user settings and function control of the device. The output device 63 can include a display device such as a display screen.
[0113] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method for constructing a knowledge graph based on Internet surveying data and model display according to any embodiment.
[0114] The computer storage medium of the embodiment of the present application can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples (non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, device or apparatus.
[0115] A computer readable signal medium can include a propagated data signal with computer executable code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport programming code.
[0116] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0117] Computer program code for carrying out operations for aspects of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0118] Finally, it should be noted that the above-described embodiments are merely intended to illustrate the technical solutions of the present application, rather than limit the technical solutions of the present application; even though the above-described embodiments of the present application have been described in detail, those skilled in the art should understand that: they can still modify the technical solutions recorded in the above-described embodiments, or make equivalent replacement to some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present application.
Claims
1. An Internet mapping data-based knowledge graph construction and model display method, characterized in that, The method comprises the following steps: obtaining mapping asset data from multiple Internet mapping platforms; classifying and encoding the mapping asset data; extracting multiple data types from a data ontology model for constructing a knowledge graph, wherein the multiple data types include ASN, location, device, service, operating system, and vulnerability; constructing a three-dimensional measurement empty matrix with dimensions of each data type, mapping platform, and data quality evaluation index, wherein the data quality evaluation index includes reliability, accuracy, completeness, consistency, and timeliness; assigning a first element value to the matrix element of each mapping platform according to the usage rate of each mapping platform; matching the ASN and location from each mapping platform with known data in an existing address library, and assigning a second element value to the accuracy of the ASN and location of each mapping platform according to the matching rate; matching the device from each mapping platform with known data in an existing fingerprint library, and assigning a second element value to the device accuracy of each mapping platform according to the matching rate; assigning a third element value to the completeness of each data type of each mapping platform according to the number of overlapping fields between the data type data from each mapping platform and the data standardization template; calculating the final value of each element in the three-dimensional measurement matrix according to the first, second, and third element values; obtaining the confidence value of each data type according to the three-dimensional measurement matrix; constructing a fusion mode configuration template for each data type according to the confidence value of each data type, wherein the fusion mode in the template includes confidence priority, aggregation, time priority, and priority of a certain mapping platform; fusing data of different forms according to the fusion mode configuration template for each data type; constructing a knowledge graph of Internet assets based on standardized data; using interactive graphical display technology to display the knowledge graph to users.
2. The method of claim 1, wherein, The method of classifying and encoding the mapping asset data comprises the following steps: extracting multiple data types from a data ontology model for constructing a knowledge graph, wherein the multiple data types include ASN, location, device, service, operating system, and vulnerability; performing multi-level subdivision on each data type according to network asset characteristics; encoding each data in the mapping asset data according to the data type and multi-level subdivision type; adding other information in each data except the encoded information to the encoded data as a label.
3. The method of claim 1, wherein, The method of constructing a fusion mode configuration template for each data type according to the confidence value of each data type comprises the following steps: performing secondary classification on each data type according to the confidence distribution law of the same data type between different mapping platforms; constructing a fusion mode configuration template for each data type according to the classification result.
4. The method of claim 3, wherein, The method of performing secondary classification on each data type according to the confidence distribution law of the same data type between different mapping platforms comprises the following steps: calculating the standard deviation of the confidence value of the same data type data from different mapping platforms, and dividing the data type into two categories of aggregated fusion and non-aggregated fusion according to the size of the standard deviation. According to the significance of the same data type in the three-dimensional measurement matrix in different mapping platform dimensions, the data type of the non-aggregated fusion type is divided into a certain mapping platform priority and a confidence priority.
5. The method of claim 1, wherein, The knowledge graph of the Internet asset is constructed based on the standardized data, including: According to the data ontology model, node modeling is performed based on data classification and coding into different categories; Relationship modeling is performed according to the relationship between nodes; A graph database management data interface is defined to support the increase, deletion, modification and query of the graph model data; The graph data is updated regularly using an incremental update interface, so that the updated graph data can reflect the latest data of the network.
6. The method of claim 1, wherein, The knowledge graph is displayed to the user using an interactive graphical display technology, including: Each network asset node and its associated edge is presented to the user in the form of a graph, and the user can zoom in, drag and rotate the graph and switch the presentation mode; In response to the user's selection of a certain type of node, the selected type of node is topologically displayed to provide different dimensional analysis; Through graph analysis technology, key nodes and high-risk connection paths in the network are extracted and presented to help the user identify potential risk points in the network.
7. An electronic device, comprising: It includes: One or more processors; Memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the knowledge graph construction and model display method based on Internet mapping data according to any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which is executed by a processor to implement the knowledge graph construction and model display method based on Internet mapping data according to any one of claims 1-6.
Citation Information
Patent Citations
Network condition situation dynamic drawing system and method fusing data quality multi-dimensional evaluation
CN112732781A
Knowledge graph construction method and platform and computer storage medium
CN116881476A