A domain name system risk analysis method, device and related equipment
Patent Information
- Application Number
- CN202611304089.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
然而,由于各告警列表彼此独立,运维人员只能逐个查看和处理每个告警列表,风险处理效率不高
[0020]第二方面提供的域名系统风险分析装置对应于第一方面提供的域名系统风险分析方法,故第二方面中任意一种实现方式所具有的技术效果,可参见上述第一方面中相应实现方式所具有的技术效果的相关之处描述,在此不做赘述。
Smart Images

Figure CN122824518A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer network technology, and in particular to a method, apparatus and related equipment for domain name system risk analysis. Background Technology
[0002] The Domain Name System (DNS) is a crucial infrastructure of the internet, used to map domain names to Internet Protocol (IP) addresses. Domain name resolution typically relies on the correct configuration of multiple components, including parent domain delegation, authoritative servers, glue records, and alias records. If any of these components is misconfigured, it can lead to domain name resolution failures or reduced service availability. Therefore, DNS risk detection is necessary to identify potential configuration problems. Currently, DNS risk detection is typically achieved through configuration checking tools. These tools use various checking rules to examine the configuration of individual domain names and generate a separate list of configuration error alerts for each domain.
[0003] In real-world scenarios, it's often necessary to monitor multiple domains simultaneously. Since configuration monitoring tools generate separate alert lists for each domain, multiple domains result in a large number of independent alert lists. However, because these alert lists are independent, operations personnel can only view and handle each alert list individually, leading to inefficient risk management. Summary of the Invention
[0004] This application provides a domain name system risk analysis method to determine the scope of impact of the same configuration defect, thereby providing a basis for decision-making in subsequent risk handling and improving the efficiency of risk handling. In addition, this application also provides a corresponding domain name system risk analysis device, computing equipment, computer-readable storage medium, and computer program product.
[0005] Firstly, this application provides a method for domain name system risk analysis. This method can be executed by a domain name system risk analysis device or a computing device with data processing capabilities. Specifically, the domain name system risk analysis device acquires a directed graph constructed based on domain name resolution data. This directed graph includes various types of nodes and various types of edges, with labels indicating the type of association between different nodes. Nodes include nodes indicating domain names recorded in the domain name resolution data, and nodes indicating other types of information associated with the domain names. Then, the domain name system risk analysis device examines the directed graph to obtain a starting node, wherein the risk of the starting node causing domain name resolution failure when participating in domain name resolution is higher than a risk threshold. Next, based on the labels of the edges connected to the starting node, the domain name system risk analysis device traverses the directed graph from the starting node to obtain multiple target nodes affected by the starting node. The number of hops between each target node and the starting node does not exceed a hop count threshold, which is determined based on the edge labels.
[0006] Thus, by acquiring a labeled directed graph, the domain names and their associated objects, originally scattered throughout the domain name resolution data, are organized into a graph structure that can express relationships through labels. This provides a foundation for subsequently detecting the range of domain names affected by the same configuration defect. Since the starting node obtained is selected based on the degree of risk, the traversal does not start randomly from any node, but specifically from the starting node where the resolution risk has been confirmed to exist. This effectively improves the accuracy of the risk analysis results. Because different types of edges correspond to different hop count thresholds, the traversal does not indiscriminately spread along all associated edges, but proceeds layer by layer according to the actual propagation level of the relationships in domain name resolution, thereby improving the rationality of the risk analysis results. Therefore, other nodes affected by the same node with resolution risk can be identified together, allowing the scope of the same configuration defect to be centrally determined. This eliminates the need for operations personnel to sift through numerous independent alarm lists to check multiple alarms caused by the same defect, thus avoiding duplicate processing of the same configuration defect. This provides a basis for decision-making in subsequent risk handling and improves risk handling efficiency.
[0007] In one possible implementation, the DNS risk analysis device determines a hop count threshold based on the labels of the edges connected to the starting node. Then, starting from the starting node, the device traverses the directed graph according to the hop count threshold to obtain nodes whose hop count to the starting node does not exceed the threshold. Thus, the traversal process is no longer performed unrestricted along all associated edges, but is limited by the hop count threshold, excluding irrelevant nodes and improving the accuracy of determining nodes affected by the starting node.
[0008] In one possible implementation, the edges connected to the starting node include a first edge and a second edge, which are of different types. The DNS risk analysis device determines a first hop count threshold based on the label of the first edge connected to the starting node. The DNS risk analysis device also determines a second hop count threshold based on the label of the second edge connected to the starting node. Then, starting from the starting node, the DNS risk analysis device traverses the directed graph along the direction indicated by the first edge according to the first hop count threshold to obtain a first target node. Furthermore, starting from the starting node, the DNS risk analysis device traverses the directed graph along the direction indicated by the second edge according to the second hop count threshold to obtain a second target node. Multiple target nodes include both first and second target nodes. In this way, different types of edges are traversed separately according to their corresponding hop count thresholds, thus differentiating the scope of influence of different types of associations, rather than using a uniform traversal range, thereby further improving the accuracy of determining the nodes affected by the starting node.
[0009] In one possible implementation, the Domain Name System (DNS) risk analysis device examines a directed graph to obtain multiple candidate nodes and the type of risk each candidate node poses that could lead to DNS resolution failure. Then, based on the risk type and mapping relationship corresponding to each candidate node, the DNS risk analysis device determines the failure risk value of each candidate node. Next, based on the failure risk value of each candidate node, the DNS risk analysis device filters nodes whose failure risk values exceed a risk threshold from the multiple candidate nodes to obtain the starting node. In this way, the starting node no longer relies on subjective human judgment but is determined through an objective comparison of the failure risk value and the risk threshold, allowing the risk level of the starting node to be quantified and compared, thereby improving the accuracy and objectivity of starting node identification.
[0010] In one possible implementation, the Domain Name System (DNS) risk analysis device detects a directed graph to obtain multiple candidate nodes, including at least one of the following three methods: The first method involves the DNS risk analysis device detecting the topology of the directed graph to obtain multiple candidate nodes, including nodes in the topology that cause domain name resolution failure. The second method involves the DNS risk analysis device detecting the semantics of the nodes in the directed graph to obtain multiple candidate nodes, including nodes whose semantics do not conform to the DNS protocol and nodes with anomalous configuration semantics. Nodes with anomalous configuration semantics have an increased risk of causing domain name resolution failure due to their configuration semantics. The third method involves the DNS risk analysis device detecting the graph features of the directed graph to obtain multiple candidate nodes. Graph features include the position, connection method, or attribute distribution of nodes in the directed graph. The quantization value corresponding to each candidate node exceeds a quantization threshold, and the quantization value corresponding to the candidate node is determined based on the graph features of the candidate node.
[0011] In this way, by detecting topology, node semantics, and graph features, nodes with parsing risks can be identified from the graph structure level, configuration level, and infrastructure level, respectively, so that risks at different levels are covered, thereby improving the comprehensiveness of risk identification.
[0012] In one possible implementation, multiple candidate nodes include a first candidate node. When the first candidate node corresponds to a single risk, the domain name system risk analysis device queries the basic risk value corresponding to the type of the single risk according to the mapping relationship, and uses the basic risk value corresponding to the type of the single risk as the failure risk value of the first candidate node. When the first candidate node corresponds to at least two risks, the domain name system risk analysis device queries the basic risk value corresponding to the type of each of the at least two risks according to the mapping relationship, and fuses the basic risk values corresponding to the types of the at least two risks according to a preset risk fusion rule to determine the failure risk value of the first candidate node. For example, the preset risk fusion rule may be to take the maximum value among the basic risk values corresponding to the types of the at least two risks, or it may be a weighted sum based on the weights of each risk type.
[0013] In this way, different failure risk value determination methods are used for single risks and multiple risks, which allows the failure risk value of a single risk node to be determined directly, and also allows the failure risk value of multiple risk nodes to be determined comprehensively, thereby improving the accuracy and rationality of failure risk value determination.
[0014] Secondly, this application provides a domain name system risk analysis device. The device includes an acquisition module, a detection module, and a traversal module. The acquisition module acquires a directed graph constructed based on domain name resolution data. The directed graph includes various types of nodes and various types of edges. Edges have labels indicating the type of association between different nodes. Nodes include nodes indicating domain names recorded in the domain name resolution data, and nodes indicating other types of information associated with the domain names. The detection module detects the directed graph to obtain the starting node, whose risk of causing domain name resolution failure when participating in domain name resolution exceeds a risk threshold. The traversal module, based on the labels of the edges connected to the starting node, traverses the directed graph from the starting node to obtain multiple target nodes affected by the starting node. The number of hops between each target node and the starting node does not exceed a hop count threshold, which is determined based on the edge labels.
[0015] In one possible implementation, the traversal module is specifically used to determine a hop count threshold based on the labels of the edges connected to the starting node. Then, starting from the starting node, the traversal module traverses the directed graph according to the hop count threshold to obtain nodes whose hop count from the starting node does not exceed the hop count threshold.
[0016] In one possible implementation, the traversal module is specifically used to determine a first hop count threshold based on the label of the first edge connected to the starting node when the edges connected to the starting node include a first edge and a second edge. The traversal module then determines a second hop count threshold based on the label of the second edge connected to the starting node. Then, starting from the starting node, the traversal module traverses the directed graph along the direction indicated by the first edge according to the first hop count threshold to obtain a first target node. Furthermore, starting from the starting node, the traversal module traverses the directed graph along the direction indicated by the second edge according to the second hop count threshold to obtain a second target node. The multiple target nodes include the first target node and the second target node.
[0017] In one possible implementation, the detection module is specifically used to detect the directed graph, obtaining multiple candidate nodes and the type of risk that each candidate node might cause domain name resolution failure when participating in domain name resolution. Then, based on the risk type and mapping relationship corresponding to each candidate node, the detection module determines the failure risk value of each candidate node. Next, based on the failure risk value of each candidate node, the detection module filters out nodes from the multiple candidate nodes whose failure risk values exceed a risk threshold to obtain the starting node.
[0018] In one possible implementation, the detection module is specifically used to detect the topological structure of the directed graph to obtain multiple candidate nodes. Alternatively, the detection module is specifically used to detect the semantics of the nodes in the directed graph to obtain multiple candidate nodes. Or, the detection module is specifically used to detect the graph features of the directed graph to obtain multiple candidate nodes.
[0019] In one possible implementation, the detection module is specifically used to query the basic risk value corresponding to the type of each of the at least two risks according to the mapping relationship when multiple candidate nodes include a first candidate node and the first candidate node corresponds to at least two risks, and to determine the failure risk value of the first candidate node according to the basic risk value corresponding to the type of the at least two risks.
[0020] The domain name system risk analysis device provided in the second aspect corresponds to the domain name system risk analysis method provided in the first aspect. Therefore, the technical effects of any implementation method in the second aspect can be found in the relevant descriptions of the technical effects of the corresponding implementation methods in the first aspect, and will not be repeated here.
[0021] Thirdly, this application provides a computing device. The computing device includes a processor and a memory, the memory storing a computer program. When the processor executes the computer program, it implements the domain name system risk analysis method provided in any possible embodiment of the first aspect.
[0022] Fourthly, this application provides a computer-readable storage medium. This computer-readable storage medium stores a computer program, which, when executed by a processor, implements the Domain Name System risk analysis method provided in any possible embodiment of the first aspect.
[0023] Fifthly, this application provides a computer program product. This computer program product includes a computer program that, when executed by a processor, implements the Domain Name System risk analysis method provided in any possible implementation of the first aspect.
[0024] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0025] Figure 1 A schematic diagram of the structure of a domain name system risk analysis system provided in this application embodiment; Figure 2 A schematic flowchart of a domain name system risk analysis method provided in this application embodiment; Figure 3 This application provides a schematic diagram of a tagged DNS resolution directed graph construction. Figure 4This application provides a schematic diagram illustrating the traversal process and the output of a set of risk nodes in an embodiment. Figure 5 This is a schematic diagram of a candidate node detection and starting node screening process provided in an embodiment of this application; Figure 6 A schematic diagram of a root cause attribution and evidence chain generation process provided in this application embodiment; Figure 7 This application provides a schematic diagram of a fault scenario simulation and governance sequencing process. Figure 8 This application provides a schematic diagram of the overall process of a domain name system risk analysis method. Figure 9 A schematic diagram of a domain name system risk analysis device provided in this application embodiment; Figure 10 This is a schematic diagram of the hardware structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0026] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a method of distinction used in describing objects with the same attributes in the embodiments of this application.
[0027] To make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the embodiments of this application will be further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0028] See Figure 1 , Figure 1 This is a schematic diagram of the structure of a domain name system risk analysis system provided in an embodiment of this application. Figure 1 As shown, the Domain Name System risk analysis system may include DNS data device 101, graph data device 102, and DNS risk analysis device 103.
[0029] DNS data device 101 can be implemented by one or more computing devices with data processing capabilities. For example, DNS data device 101 can be implemented by a server, server cluster, cloud computing node, or other computing device. DNS data device 101 is used to store domain name resolution data. Domain name resolution data can come from offline zone file resolution, real-time DNS queries, passive DNS data, or the result of a fusion of multiple data sources. DNS data device 101 can be deployed independently or as a data node in a DNS management platform, security operation platform, or other network management system.
[0030] Graph data device 102 can be implemented by a database server, graph database cluster, storage server, or other computing device with graph data storage and management capabilities. Graph data device 102 may include a memory for storing labeled directed graphs and a data interface for accessing the graph data. When graph data device 102 is implemented by multiple computing devices, the multiple computing devices can form a graph database cluster. Graph data device 102 is used to store directed graphs constructed based on domain name resolution data. The directed graph includes various types of nodes and various types of edges, with labels indicating the type of association between different nodes.
[0031] The DNS risk analysis device 103 can be implemented by a separately deployed processor, controller, programmable logic device, or other hardware with data processing capabilities, and this application does not limit this. The DNS risk analysis device 103 is used to obtain a directed graph constructed based on domain name resolution data, detect the directed graph to obtain the starting node, and determine the risk level of the starting node causing domain name resolution failure when participating in domain name resolution, which is higher than the risk threshold. Based on the labels of the edges connected to the starting node, the device traverses the directed graph starting from the starting node to obtain multiple target nodes affected by the starting node.
[0032] exist Figure 1 In the implementation shown, DNS data device 101 can provide domain name resolution data to graph data device 102. Graph data device 102 can construct a labeled directed graph based on the domain name resolution data. Alternatively, other devices can construct a labeled directed graph based on the domain name resolution data and then provide it to graph data device 102 for storage. DNS risk analysis device 103 can obtain the labeled directed graph from graph data device 102 and perform the aforementioned detection and traversal operations.
[0033] In one implementation, the DNS data device 101, the graph data device 102, and the DNS risk analysis device 103 can be deployed in different computing devices. When the three devices are deployed separately, the DNS data device 101, the graph data device 102, and the DNS risk analysis device 103 can interact with each other through a communication network or a data interface.
[0034] In another implementation, at least two of the DNS data device 101, graph data device 102, and DNS risk analysis device 103 can be integrated into the same computing device or the same server cluster. The integrated devices can interact with each other through local data access, shared storage space, or program calls. In a distributed deployment, the DNS data device 101, graph data device 102, and DNS risk analysis device 103 can also collaborate through databases, message queues, graph computing services, or program interfaces.
[0035] therefore, Figure 1 The device division in this document is primarily used to illustrate different data sources and data processing functions, and is not intended to limit the DNS risk analysis system to include three independent physical devices. DNS data device 101, graph data device 102, and DNS risk analysis device 103 can be configured separately according to the actual deployment environment, and their respective functions can also be integrated into one or more computing devices.
[0036] The DNS risk analysis device 103 is also used to output the starting node and the target node as a set of interconnected risk nodes. The output can be in the form of an application programming interface (API) result, a report file, or a graphical interface.
[0037] In one possible implementation, the DNS risk analysis device 103 may include a processor and a memory. The processor is used to execute a computer program corresponding to the Domain Name System risk analysis method. The memory is used to store the computer program and the data used or generated during the execution of the method. The DNS risk analysis device 103 may also include a communication interface for data interaction with the DNS data device 101 and the graph data device 102.
[0038] It should be noted that, Figure 1 The system architecture shown is merely an example to illustrate a deployment scenario applicable to the domain name system risk analysis method provided in this application, and does not constitute a limitation on the scope of protection of this application. For example, in other possible implementations, the domain name system risk analysis system may also include an output device, a display device, or a terminal device for displaying a set of risk nodes to operations and maintenance personnel.
[0039] In one possible implementation, the target size of the labeled directed graph can be in the tens of thousands. The graph data device 102 can construct the labeled directed graph using an offline batch processing method, without incremental updates. A single full construction can be completed in minutes.
[0040] During the domain name resolution process, a domain name is usually not isolated, but depends on the delegation management of its parent domain, the resolution service of authoritative servers, and possibly the resolution results of alias targets. These entities can serve multiple domain names simultaneously, thus different domain names are associated with each other through a common parent domain, authoritative server, or alias target.
[0041] However, relevant DNS configuration checking tools analyze DNS configurations on a per-domain basis, checking each domain separately. For each domain checked, the tool generates a separate list of configuration error alerts, without reflecting the aforementioned relationships between domains. When a configuration flaw affects multiple domains, the flaw will appear in the alert lists of different domains, creating multiple independent alerts.
[0042] Because the alert lists are independent of each other, operations and maintenance (O&M) personnel can only learn about the configuration problems of each domain from the alert lists themselves. They cannot determine whether these alerts are caused by the same configuration defect, and therefore cannot determine which domains are specifically affected by a configuration defect. In this situation, O&M personnel can only check and handle these alerts one by one, and it is only after handling a large number of alerts that they may discover that these alerts actually originate from the same configuration defect. In this process, O&M personnel repeatedly handle the same defect, increasing unnecessary O&M workload and resulting in low risk handling efficiency.
[0043] Based on this, this application provides a method for domain name system risk analysis. This method uniformly represents the resolution dependencies between domain names and their associated objects recorded in domain name resolution data using a labeled directed graph. On this basis, the directed graph is inspected to identify the starting node whose risk level of causing domain name resolution failure exceeds a risk threshold when participating in domain name resolution. Then, based on the labels of the edges connected to the starting node, a traversal is performed starting from the starting node to obtain multiple target nodes affected by the starting node. Through this method, a node with resolution risk can be associated with other nodes with resolution risk along the resolution dependencies in the directed graph, thereby determining the scope of the configuration defect's impact, providing a decision-making basis for subsequent risk handling, and improving risk handling efficiency.
[0044] based on Figure 1 The domain name system risk analysis system shown below, combined with... Figure 2 The main process of the domain name system risk analysis method provided in the embodiments of this application will be described.
[0045] Figure 2 This is a schematic flowchart illustrating a domain name system risk analysis method provided in an embodiment of this application. The method can be... Figure 1 The DNS risk analysis device 103 shown can be used for this purpose, but it can also be executed by other computing devices with data processing capabilities. For ease of explanation, the following description uses the DNS risk analysis device 103 as an example. Figure 2 As shown, the method may include the following steps.
[0046] S201: DNS risk analysis device 103 acquires a directed graph constructed based on domain name resolution data.
[0047] Directed graphs include various types of nodes and various types of edges. Edges have labels that indicate the type of relationship between different nodes. Nodes include those indicating domain names recorded in domain name resolution data, and those indicating other types of information associated with domain names.
[0048] In this embodiment, domain name resolution data refers to data used to describe the domain names and their associated objects involved in the domain name resolution process. Domain name resolution data can originate from offline zone file resolution, real-time DNS queries, passive DNS data, or a fusion of multiple data sources. This application does not limit the specific source of the domain name resolution data.
[0049] For example, each domain name resolution data may include fields such as domain name, parent domain, list of authoritative name servers (NS), canonical name (CNAME) target, and address record.
[0050] In DNS, the parent domain refers to the parent domain of a domain in the domain hierarchy. For example, the parent domain of the domain "www.example.com" is "example.com", and the parent domain of the domain "example.com" is "com". NS refers to the servers that provide DNS resolution services for a domain. A domain can specify one or more authoritative servers through an NS list, and these authoritative servers are responsible for responding to DNS resolution queries for that domain and its subdomains. A CNAME record is a type of DNS record, also known as a canonical name record or alias record. A CNAME record is used to point one domain to another. When a domain has a CNAME record configured, the DNS resolution result does not directly return an address record, but rather the target domain pointed to by the CNAME record. After obtaining the target domain, the resolver needs to continue resolving the target domain until it obtains the final address record. An address record is an address mapping record corresponding to a domain or authoritative server, and can include A records (Arecord) and AAAA records (AAAA record). A records are used to map domain names to Internet Protocol version 4 (IPv4) addresses, and AAAA records are used to map domain names to Internet Protocol version 6 (IPv6) addresses. In this application, A records and AAAA records can be collectively referred to as A / AAAA address records.
[0051] Nodes are the basic elements in a graph, used to represent domain names recorded in domain name resolution data, as well as other types of objects associated with those domain names. A node must include at least a domain name node and nodes containing other types of information associated with that domain name.
[0052] For example, nodes may include the following types: domain name node, used to indicate the domain name recorded in the domain name resolution data; parent domain node, used to indicate the parent domain corresponding to the domain name recorded in the domain name resolution data; authoritative server node, used to indicate the authoritative server hostname that provides resolution services for the domain name; alias target node, used to indicate the target domain name that the domain name points to through the CNAME record; and address node, used to indicate the A record or AAAA record address corresponding to the domain name or authoritative server.
[0053] In cases where the same name may play multiple roles in domain name resolution data, such as "ns1.example.com" being both resolved as a domain name and acting as an authoritative server providing resolution services for other domain names, in this embodiment, the DNS risk analysis device 103 creates a node using the domain name string as the primary key. This node can be pointed to by other domain name nodes through an authoritative server edge, and it can also have a parent domain delegation edge pointing to its parent domain. This conforms to the semantics of the same name representing the same entity in the DNS resolution system.
[0054] In this context, delegation refers to a parent domain entrusting the management of its child domain's DNS resolution to a designated authoritative server. For example, .com delegates the management of example.com to the authoritative server ns1.example.com, which is responsible for providing DNS resolution for all domains under example.com. If ns1.example.com has a configuration defect, all child domains that rely on example.com for DNS resolution may be affected. In this embodiment, this delegation relationship is represented in the graph by a parent domain delegation edge pointing from a child domain to the parent domain.
[0055] See Figure 3 , Figure 3 This is a schematic diagram illustrating the construction of a tagged DNS resolution directed graph, as provided in an embodiment of this application. Figure 3As shown, the DNS risk analysis device 103 extracts the domain name node, parent domain node, authoritative server node, CNAME target node, and address node from the domain name resolution data. Simultaneously, the DNS risk analysis device 103 generates parent domain delegation edges, authoritative server edges, and alias edges based on the fields in the domain name resolution data. Each edge carries a label indicating the type of the corresponding relationship. For example, a parent domain delegation edge represents a resolution dependency relationship where a child domain points to a parent domain; an authoritative server edge represents a resolution dependency relationship where a domain name points to an authoritative server hostname; and an alias edge represents a resolution dependency relationship where a domain name points to a CNAME target. Thus, the DNS risk analysis device 103 constructs a labeled directed DNS resolution graph, providing the graph data foundation for subsequent detection and traversal.
[0056] Edge labels are used to indicate the type of association between different nodes. Association types can include parent domain delegation, authoritative server, alias, etc.
[0057] Specifically, the DNS risk analysis device 103 can generate edges according to the following rules: if the domain name has a parent domain field, a parent domain delegation edge is generated pointing from the domain name to the parent domain; if the domain name has an NS list, an authoritative server edge is generated pointing from the domain name to each NS hostname; if the domain name has a CNAME target, an alias edge is generated pointing from the domain name to the CNAME target. Each edge can carry a corresponding label, which indicates the type of association corresponding to the edge. For example, the label for a parent domain delegation edge can be "parent", the label for an authoritative server edge can be "ns", and the label for an alias edge can be "cname".
[0058] For example, edges can be stored in the form of an edge table. A row in the edge table can include three fields: start point, end point, and edge label. Table 1 provides an example of a labeled edge table.
[0059] Table 1. An example of a labeled edge table
[0060] In this way, the DNS risk analysis device 103 organizes the domain names and their associated objects, which were originally scattered in the domain name resolution data, into a graph structure that can express the relationships through labels, using a labeled directed graph. This provides a foundation for subsequent detection of the range of domain names affected by the same configuration defect. Compared with the method of analyzing only a single domain name or a single record, a directed graph can express the relationship between a domain name and objects such as authoritative servers, parent domains, and alias targets. This allows subsequent risk analysis to identify interconnected risk nodes along the relationships without omitting other nodes associated with the starting node.
[0061] It should be noted that the direction of the edges in this embodiment can adopt the resolution dependency direction. For example, the parent domain delegation edge points the child domain to the parent domain, the authoritative server edge points the domain name to the NS hostname, and the alias edge points the domain name to the CNAME target. When traversing, the DNS risk analysis device 103 can traverse in the opposite direction to the edge direction to determine the nodes affected by the starting node, or it can pre-convert the edge direction to the service provision direction before traversing. This application does not limit the specific storage direction of the edges.
[0062] S202: DNS risk analysis device 103 detects the directed graph and obtains the starting node.
[0063] The starting node is a node whose risk of causing domain name resolution failure when participating in domain name resolution exceeds a risk threshold. The DNS risk analysis device 103 can first detect the directed graph to obtain multiple candidate nodes, and the type of risk that each candidate node will cause domain name resolution failure when participating in domain name resolution, thereby determining the starting node.
[0064] In this context, "risk type" refers to the specific risk category that a candidate node may pose, potentially leading to domain name resolution failure. "Basic risk value" is a pre-defined numerical value corresponding to the risk type, used to characterize the severity of that risk type on domain name resolution. "Failure risk value" is a numerical value determined for a candidate node based on its risk type and basic risk value, used to characterize the overall risk level of that candidate node causing domain name resolution failure when participating in domain name resolution. The failure risk value can be understood as the quantitative basis for selecting a candidate node as the starting node.
[0065] Then, the DNS risk analysis device 103 determines the failure risk value based on the basic risk value corresponding to the risk type, and takes the candidate node whose failure risk value exceeds the risk threshold as the starting node.
[0066] DNS risk analysis device 103 detects multiple candidate nodes in a directed graph, which may include any one or more of the following: topology detection, node semantic detection, and graph feature detection.
[0067] Topology detection is used to identify risks associated with the connections between nodes in a directed graph. The DNS risk analysis device 103 can detect whether a closed-loop structure consisting of nodes and relationships exists in the directed graph. For example, the DNS risk analysis device 103 can detect strongly connected components (SCCs) in the directed graph. If multiple nodes have mutually reachable directed paths, these nodes constitute a strongly connected component. Nodes in a closed-loop structure, due to their mutual dependencies, may cause resolution failures when participating in domain name resolution.
[0068] For example, in a parent domain delegation subgraph, if a strongly connected component exists, then that strongly connected component can correspond to a circular delegation. A circular delegation indicates that the parent domain delegation relationship forms a closed loop, and the resolver cannot determine the final authoritative server within this closed loop, leading to resolution failure.
[0069] For example, in a resolution dependency subgraph containing parent domain delegation edges, authoritative server edges, and alias edges, if a strongly connected component exists, then that strongly connected component can correspond to an authoritative dependency cycle. An authoritative dependency cycle can be understood as a closed loop of resolution dependencies formed by parent domain delegation relationships, authoritative server relationships, and / or alias relationships. The DNS risk analysis device 103 uses nodes in the closed loop structure as candidate nodes.
[0070] In one possible implementation, the DNS risk analysis device 103 can also detect other topology anomalies. For example, multiple parent conflicts: In a parent domain delegation subgraph, if a node in the same subdomain has two or more parent domain nodes, the DNS risk analysis device 103 determines it as a multiple parent conflict. This detection is a pure topology check, meaning it only judges based on the number of parent domain nodes and does not compare delegation content. Depth overflow: The DNS risk analysis device 103 detects whether the delegation chain depth of a node exceeds a preset threshold. The default threshold can be 10 layers, but it can also be configured as needed. Hanging authorization: The DNS risk analysis device 103 detects whether the node pointed to by the parent domain delegation edge exists in the directed graph. If it does not exist, the edge is a hanging authorization, and the related node is considered a candidate node.
[0071] Node semantic detection is used to identify nodes whose semantics do not conform to the Domain Name System (DNS) protocol, as well as nodes with anomalous configuration semantics. Nodes whose semantics do not conform to the DNS protocol refer to nodes whose configurations do not meet the basic requirements stipulated by the DNS protocol. Nodes with anomalous configuration semantics refer to nodes whose configurations appear to meet some requirements, but whose configuration content is abnormal, resulting in insufficient actual fault tolerance.
[0072] Nodes whose semantics do not conform to the Domain Name System (DNS) protocol can include missing glue records, insufficient NS redundancy, and CNAME rings. A missing glue record means that the authoritative server node lacks an A / AAAA address record. A glue record is an A / AAAA address record provided by the parent zone for the authoritative servers within its jurisdiction. If an authoritative server node lacks an A / AAAA address record, the resolver cannot determine the address of that authoritative server, leading to resolution failure. Insufficient NS redundancy means that the number of authoritative servers associated with a domain name node is less than a preset threshold, such as less than 2. A CNAME ring refers to a closed loop formed by alias chains.
[0073] Nodes in abnormal configuration semantics can include pseudo-redundancy. Pseudo-redundancy refers to the apparent configuration of multiple NSs or addresses, but these NSs ultimately fall within the same IP address, IP prefix, provider, or fault domain, resulting in insufficient actual fault tolerance. For example, a domain name may be configured with three authoritative servers "ns1.example.com", "ns2.example.com", and "ns3.example.com", but if all three authoritative servers point to the same IP address, then the domain name actually has only one authoritative server, creating a hidden single point of failure.
[0074] In one possible implementation, the DNS risk analysis device 103 can also detect static bailiwicks. A bailiwick can be understood as whether a name is within a delegated zone or its jurisdiction according to rules. Specifically, the DNS risk analysis device 103 can detect whether the authoritative server hostname is equal to the corresponding zone, or whether it is a subdomain of that zone. For example, for the "example.com" zone, the authoritative server "ns1.example.com" satisfies the suffix matching rules and is within its bailiwick; while "ns1.cloudflare.net" does not, and the DNS risk analysis device 103 determines it to be a static bailiwick. This detection is a static preflight signal and does not necessarily mean that the resolution will fail, because many resolvers can eventually obtain A / AAAA records through other paths. Therefore, the failure risk value corresponding to this anomaly can be low, for example, set to 0.1.
[0075] Graph feature detection is used to identify risks related to the location, connectivity, or attribute distribution of nodes in a directed graph. Graph features can include node degree, betweenness centrality, parent domain dependency, NS redundancy, subgraph entropy, connected components, service provider identifier concentration, or IP prefix concentration. The DNS risk analysis device 103 can calculate the quantized values of the graph features corresponding to candidate nodes and select nodes whose quantized values exceed the quantization threshold as candidate nodes.
[0076] In this context, node degree refers to the number of edges connected to a node, representing the degree of direct association of that node in a directed graph. Betweenness centrality refers to the proportion of shortest paths in a directed graph that pass through a given node, representing the degree to which that node acts as a bridge node. Parent domain dependency refers to the number of parent domain nodes a domain node depends on through parent domain delegation edges, representing the dependency of that domain node at the parent domain level. NS redundancy refers to the number of authoritative server nodes associated with a domain node, representing the redundancy of that domain node at the DNS resolution service level. Subgraph entropy refers to the uniformity of node or edge distribution within a subgraph of a directed graph, representing the structural complexity of that subgraph. Connected components refer to the largest set of nodes that are mutually reachable in a directed graph, representing the connectivity of the directed graph. Service provider identity concentration refers to the degree to which multiple authoritative server nodes belong to the same service provider, representing the concentration risk of DNS resolution services at the provider level. IP prefix concentration refers to the degree to which multiple address nodes belong to the same IP prefix, representing the concentration risk at the address distribution level.
[0077] In one possible implementation, the DNS risk analysis device 103 can detect service provider identification concentration. For example, the DNS risk analysis device 103 can obtain the hostnames of all authoritative servers associated with the domain name node and determine whether these authoritative servers belong to the same service provider based on hostname keywords. If all authoritative servers belong to the same provider, the DNS risk analysis device 103 outputs a risk of provider concentration. As another example, the DNS risk analysis device 103 can detect whether the IP addresses resolved by all authoritative servers belong to the same IP address or the same / 24 subnet. If so, the DNS risk analysis device 103 outputs a risk of hidden single point of failure.
[0078] Graph feature detection can employ Boolean conditional judgment, which determines whether a node meets preset structural risk conditions (yes or no), rather than calculating continuous risk scores. A node is considered a candidate node if it meets any of the conditions in Table 2. The judgment rules in Table 2 are illustrative examples; in practice, other structural risk conditions can be set based on the specific circumstances of the domain name resolution data.
[0079] Table 2 shows examples of feature detection conditions.
[0080] In this context, the authoritative server subgraph refers to a subgraph within a directed graph consisting of authoritative server edges and the domain name nodes and authoritative server nodes connected by those edges. When a directed cycle exists in the authoritative server subgraph, it indicates a circular dependency between the domain name and the authoritative server, which may lead to resolution failure.
[0081] The above judgment conditions are all illustrative examples. The DNS risk analysis device 103 may also use other graph features or other judgment methods, which are not limited in this application.
[0082] S203: DNS risk analysis device 103, based on the labels of the edges connected to the starting node, traverses the directed graph starting from the starting node to obtain multiple target nodes affected by the starting node.
[0083] In this embodiment, the starting node is a node with a risk level higher than the risk threshold. To determine the other nodes affected by the starting node, the DNS risk analysis device 103 can traverse the directed graph based on the association relationships. Here, the traversal is not performed indiscriminately along all edges connected to the starting node, but rather the corresponding hop count threshold is determined based on the label carried on each edge to limit the scope of influence of different association relationships.
[0084] The hop count threshold is used to limit the number of layers that can be extended outward from the starting node along a certain type of association. Different types of associations have different propagation characteristics in domain name resolution, so different edge labels can correspond to different hop count thresholds. For example, Table 3 shows the correspondence between different edge labels and hop count thresholds.
[0085] Table 3 Example of the correspondence between side labels and hop count thresholds
[0086] RFC stands for Request for Comments, which describes the technical specifications of the Domain Name System (DNS) protocol.
[0087] The hop count threshold for parent domain delegation edges is set to one level because in domain name resolution, the delegation of child domains from the parent domain is done level by level. For example, .com delegates to example.com, and example.com then delegates to www.example.com. When a resolution risk occurs on .com, it directly affects example.com, without bypassing example.com to directly affect www.example.com. Therefore, traversing one level along the parent domain delegation edge is sufficient to cover the directly affected area. Continuing to traverse downwards would cross the delegation boundary, causing the affected area to be unreasonably expanded.
[0088] The hop count threshold for alias edges is set to 10 levels because alias chains can have multiple levels. For example, www.example.com points to web.example.com via an alias, and web.example.com in turn points to real.example.com via an alias. When a resolution risk occurs in one of the alias targets, all alias nodes pointing to it may be affected. According to the relevant restrictions in the RFC, the maximum hop count for alias chains is 10 levels, so the hop count threshold for traversing along alias edges is set to 10 levels. In this way, potentially affected nodes in the alias chain can be covered, avoiding infinite traversal.
[0089] It should be noted that the hop count thresholds in Table 3 are configurable example values. For example, in actual deployments, if the alias chain depth is usually short, the hop count threshold for alias edges can be adjusted to a smaller value. This application does not limit the specific value.
[0090] In one possible implementation, the DNS risk analysis device 103 can determine a uniform hop count threshold based on the labels of the edges connected to the starting node. For each edge connected to the starting node, the DNS risk analysis device 103 no longer distinguishes edge types, but instead traverses them all according to the uniform hop count threshold. For example, the DNS risk analysis device 103 can set the uniform hop count threshold to 3 levels. Starting from the starting node, it traverses each edge pointing to the starting node level by level, and includes nodes whose hop count to the starting node does not exceed 3 levels as traversed nodes. Edges exceeding 3 levels are not further expanded.
[0091] Thus, by using a uniform hop count threshold during the traversal process, the complexity of traversal rule configuration can be reduced. In scenarios where it is not necessary to finely distinguish the propagation levels of different relationships, the affected nodes can be determined with less configuration.
[0092] This application does not limit whether a differentiated hop count threshold or a uniform hop count threshold is used. For ease of explanation, the traversal process of this embodiment will be described in detail below, mainly taking the determination of the hop count threshold based on the edge labels as an example.
[0093] Before traversing, the DNS risk analysis device 103 can first obtain the labels of the edges connected to the starting node and determine the hop count threshold corresponding to each edge based on the labels. Then, the DNS risk analysis device 103 starts from the starting node and traverses layer by layer along the direction of the edges pointing to the starting node.
[0094] Specifically, the DNS risk analysis device 103 can maintain a traversal queue. First, it sets the starting node as the level 0 node and adds it to the traversal queue. Then, the DNS risk analysis device 103 repeats the following process until the traversal queue is empty: The first step is to take a node from the traversal queue and use it as the current node.
[0095] The second step is to find the edge in the directed graph that points to the current node, that is, to find the edge that depends on the current node, and to get the label on the edge.
[0096] The third step is to determine the hop count threshold corresponding to the edge type for each edge pointing to the current node based on the label on the edge.
[0097] The fourth step is to determine whether the cumulative number of hops traversed along this edge type from the starting node exceeds the hop count threshold corresponding to this edge type.
[0098] Fifth step: If the number of hops has not exceeded the threshold, add the node on the other side of the edge, i.e. the node that depends on the current node, to the traversal queue and mark the node as the affected node.
[0099] Step 6: If the number of jumps has exceeded the threshold, stop traversing deeper along the edge type.
[0100] The hop count threshold is accumulated separately for each edge type. That is, for edges of the same type, the number of layers accumulated from the starting node along that edge type cannot exceed the hop count threshold corresponding to that edge type. For example, if the hop count threshold for an alias edge is 10 layers, then a maximum of 10 layers can be traversed continuously from the starting node along the alias edge; while the hop count threshold for parent domain delegation edges and authoritative server edges is 1 layer, meaning only 1 layer can be traversed from the starting node along these two types of edges, and further expansion into deeper layers is not allowed.
[0101] It should be noted that during the traversal, nodes added to the traversal queue via authoritative server edges or parent domain delegation edges will not be used as the current node to continue searching for its predecessor node after being marked as affected nodes, since the hop count threshold for these two types of edges is 1 level. Nodes added to the traversal queue via alias edges, however, can continue to be used as the current node, and other edges pointing to that node can be searched, with the search continuing according to the hop count threshold for each edge.
[0102] For example, the traversal process can be implemented using the following pseudocode: impacted := {seed} queue := [seed] while queue not empty and depth < D_max: v := pop(queue) for each predecessor u of v via allowed_edge_types: if edge(u,v) satisfies hop / layer rules: impacted.add(u); push(u) The return impacted. Where `impacted` is the set of affected nodes, `seed` is the starting node, `queue` is the traversal queue, `depth` is the current traversal depth, and `D_max` is the maximum traversal depth. `predecessor u of v` indicates that node u is the predecessor of node v, meaning there exists an edge from u to v, and node u depends on node v. `allowed_edge_types` specifies the allowed edge types based on edge labels, and `hop / layer rules` are the hop count threshold rules determined based on edge labels.
[0103] After the traversal is complete, the DNS risk analysis device 103 can use the affected nodes obtained from the traversal as target nodes.
[0104] In one possible implementation, after obtaining multiple target nodes, the DNS risk analysis device 103 can also calculate the impact scale and impact ratio of the target node set. The impact scale is the number of target nodes in the target node set. The impact ratio is the ratio of the number of target nodes in the target node set to the total number of resolution objects that depend on the starting node. The total number of resolution objects that depend on the starting node can be determined based on the number of nodes in the directed graph that are associated with the starting node.
[0105] In this embodiment, the traversal determines the topology influence range solely based on the hop count threshold, without further attenuating the failure risk value hop-by-hop along the propagation path. In other words, the failure risk value of the target node is determined by its own detection results and is not related to the hop count between the target node and the starting node.
[0106] To better understand the above traversal process and the selection of target nodes, the following section combines... Figure 4 To illustrate this with a more concrete example, see [link to example]. Figure 4 , Figure 4 This is a schematic diagram illustrating a traversal process and the output of a risk node set, provided as an embodiment of this application. In this example, the traversal path of the DNS risk analysis device 103 is as follows: Figure 4 As shown.
[0107] Assuming the starting node is the authoritative name server node ns1.example.com, this node is at risk of missing glue records.
[0108] During the traversal, the DNS risk analysis device 103 first uses ns1.example.com as the current node and searches for edges pointing to ns1.example.com. These include: the edge from www.example.com to the authoritative server of ns1.example.com, the edge from mail.example.com to the authoritative server of ns1.example.com, and the edge from alias1.example.com to ns1.example.com via an alias edge, meaning alias1.example.com is an alias of ns1.example.com.
[0109] Since the hop count threshold for the authoritative server edge is 1 level, the DNS risk analysis device 103 marks www.example.com and mail.example.com as affected nodes, but does not continue to expand outward from these two nodes. Since the hop count threshold for the alias edge is 10 levels, the DNS risk analysis device 103 marks alias1.example.com as an affected node, and uses alias1.example.com as the new current node to continue searching for edges pointing to alias1.example.com.
[0110] Next, DNS risk analysis device 103, using alias1.example.com as the current node, finds edges pointing to alias1.example.com, including: alias2.example.com pointing to alias1.example.com via an alias edge, and sub.example.com pointing to alias1.example.com via an authoritative server edge. Since the hop count threshold for the alias edge is 10 levels, DNS risk analysis device 103 adds alias2.example.com to the traversal queue along the alias edge and marks it as an affected node. Since the hop count threshold for the authoritative server edge is 1 level, DNS risk analysis device 103 marks sub.example.com as an affected node along the authoritative server edge, but does not continue expanding outwards from this node.
[0111] Similarly, the DNS risk analysis device 103 continues with alias2.example.com as the current node, finding edges pointing to alias2.example.com, including alias3.example.com pointing to alias2.example.com via alias edges. Since the hop count threshold corresponding to alias edges is 10 levels, the DNS risk analysis device 103 adds alias3.example.com to the traversal queue along the alias edges and marks it as an affected node. Assuming alias3.example.com is no longer the target of other nodes pointing to via alias edges, the traversal ends.
[0112] After the traversal is complete, the DNS risk analysis device 103 takes all the nodes obtained and marked as affected nodes as target nodes. Therefore, the target node set includes www.example.com, mail.example.com, alias1.example.com, alias2.example.com, alias3.example.com, and sub.example.com.
[0113] Thus, by determining the hop count threshold based on the edge labels and traversing layer by layer according to the hop count threshold, the traversal range is not unrestricted and spreads along all associated edges, but is limited by the hop count threshold corresponding to each type of edge. Compared with indiscriminate traversal, this method can avoid including irrelevant nodes in the target node, thereby improving the accuracy of determining nodes affected by the starting node and enhancing the rationality of the risk analysis results.
[0114] In one possible implementation, after obtaining the target node, the DNS risk analysis device 103 can also output the starting node and the target node as a set of interconnected risk nodes. The output format can be an application programming interface (API) result, a report file, or a graphical interface.
[0115] For example, the DNS risk analysis device 103 can output the set of risk nodes to other systems for further processing via API; it can also output the set of risk nodes via report files for operation and maintenance personnel to view; and it can also highlight the starting node, target node and the associated edges between them through a graphical interface to intuitively display the scope of impact.
[0116] In one possible implementation, the DNS risk analysis device 103 can first determine the failure risk value of candidate nodes to determine the starting node. When the DNS risk analysis device 103 detects multiple candidate nodes in a directed graph, each candidate node is associated with a risk type. The risk type is the type of risk that a candidate node will cause domain name resolution failure when participating in domain name resolution. Based on the risk type, the DNS risk analysis device 103 can determine the failure risk value of each candidate node according to a mapping relationship. The mapping relationship indicates the correspondence between risk types and basic risk values.
[0117] For example, Table 4 shows the mapping relationship between risk type and basic risk value.
[0118] Table 4. Examples of Mapping Risk Types to Base Risk Values
[0119] It should be noted that the basic risk values in Table 4 are configurable example values, which can be set based on reliability engineering practices and the impact of the DNS protocol, or adjusted according to the actual deployment environment. This application does not limit the specific values of the basic risk values.
[0120] For a single risk, the DNS risk analysis device 103 can directly use the basic risk value corresponding to the single risk type as the failure risk value of the candidate node. For multiple risks, the DNS risk analysis device 103 can query the basic risk values corresponding to each of the multiple risk types and determine the failure risk value of the candidate node based on the basic risk values corresponding to the multiple risk types. For example, the DNS risk analysis device 103 can take the maximum value among the basic risk values corresponding to the multiple risk types as the failure risk value. Other fusion methods, such as weighted fusion, can also be used as alternative implementation methods.
[0121] After obtaining the failure risk value for each candidate node, the DNS risk analysis device 103 can identify candidate nodes whose failure risk values exceed a risk threshold as starting nodes. The risk threshold is a preset value used to distinguish between nodes with higher risk levels and nodes with lower risk levels. For example, the risk threshold can be set to 0.3. Thus, candidate nodes with failure risk values exceeding 0.3 are identified by the DNS risk analysis device 103 as starting nodes.
[0122] In this way, by objectively comparing the failure risk value with the risk threshold, the starting node is determined based on quantitative values rather than relying on subjective human judgment, thereby improving the accuracy and objectivity of starting node identification. This transforms risk analysis from manually judging multiple alarm items to automatically determining nodes with analytical risks based on numerical comparison.
[0123] See Figure 5 , Figure 5This is a schematic diagram illustrating a candidate node detection and starting node selection process provided in an embodiment of this application. Figure 5 As shown, the DNS risk analysis device 103 can first perform topology detection, node semantic detection, and graph feature detection on the directed graph to obtain candidate nodes. Then, the DNS risk analysis device 103 determines the failure risk value based on the risk type corresponding to the candidate nodes. Finally, the DNS risk analysis device 103 selects candidate nodes whose failure risk values exceed the risk threshold as starting nodes.
[0124] In one possible implementation, the method may further include a root cause attribution step. Root cause attribution is used to determine the root cause type and contributing factor type when multiple test results coexist, thereby outputting interpretable risk evidence objects.
[0125] Specifically, the DNS risk analysis device 103 can acquire multiple detection results associated with the starting node, each detection result corresponding to a risk type. Then, the DNS risk analysis device 103 determines the root cause type and contributing factor type from the multiple risk types according to a preset risk type priority. The preset risk type priority can be determined based on the mandatory and repairable nature of the DNS protocol.
[0126] For example, root cause priorities, from highest to lowest, can be: missing glue records, CNAME cycles, authority dependency cycles, insufficient NS redundancy, insufficient or pseudo-redundancy of NSIP, static anomalies in jurisdictions, provider concentration, IP prefix concentration, or bridging vulnerabilities, etc. When the starting node is associated with multiple risks simultaneously, the protocol or topology item with the highest priority is designated as the root cause, and the remaining hit items are considered contribution factors. Phenomenon-related discoveries at the topology layer, such as reporting only the existence of a cycle without hitting higher-priority protocol items, can be written into the discovery type instead of being treated as a separate root cause.
[0127] After determining the root cause type and contributing factor type, the DNS risk analysis device 103 can calculate the confidence level. The confidence level is used to characterize the reliability of the root cause determination. For example, the confidence level can be calculated as follows: ; Where base is the base confidence level corresponding to the root cause type, root is the root cause type, k is the contribution factor decay coefficient, and num is the number of contribution factors. clamp is a function that restricts the result to between 0 and 1.
[0128] For example, Table 5 shows the correspondence between root cause types and baseline confidence levels.
[0129] Table 5 Examples of Root Cause Types and Baseline Confidence Levels
[0130] It should be noted that the baseline confidence levels in Table 5 are all exemplary values and can be configured according to the certainty of the actual detection rules. Here, k is the contribution factor decay coefficient, with an example value of 0.08, which can also be dynamically calibrated based on historical repair or verification results. The more contribution factors there are, the lower the certainty of root cause determination, and the lower the confidence level accordingly.
[0131] After obtaining the root cause type, contribution factor type, and confidence level, the DNS risk analysis device 103 can generate a risk evidence object. The risk evidence object may include the discovery type, the involved resolver, the scale of impact, the proportion of impact, the risk weight or failure risk value, the root cause, root cause evidence, contribution factors, evidence path, confidence level, and remediation recommendations. The evidence path consists of nodes and typed edges in a directed graph, and is associated with detection evidence that triggers topological, protocol, or structural risks.
[0132] See Figure 6 , Figure 6 This is a schematic diagram illustrating a root cause attribution and evidence chain generation process provided in an embodiment of this application. Figure 6 As shown, the DNS risk analysis device 103 first acquires multiple detection results associated with the starting node, each detection result corresponding to a risk type. Then, based on a preset risk type priority, the DNS risk analysis device 103 determines the root cause type and contributing factor type from the multiple risk types. Next, the DNS risk analysis device 103 calculates the confidence level based on the root cause type and contributing factor type. Afterward, the DNS risk analysis device 103 generates a risk evidence object. This risk evidence object may include root cause, contributing factor, confidence level, evidence path, etc.
[0133] In one possible implementation, when the starting node belongs to a closed-loop structure, the DNS risk analysis device 103 can also calculate the comprehensive failure risk value of the closed-loop structure. For authoritative dependency loops, the DNS risk analysis device 103 can first calculate the baseline failure risk value based on the independent event assumption: ; Where n is the number of nodes in the ring, P i Let be the failure risk value of the i-th node within the ring.
[0134] If common factors such as shared provider, shared Autonomous System number, and shared IP prefix exist within the ring, the DNS risk analysis device 103 can use a common cause failure model to correct the overestimation caused by the independence assumption. A simplified correction method is as follows: ; Where δ=1 indicates that at least two nodes in the ring hit the same common factor, otherwise δ=0; β∈[0,1] common factor strength, which can be configured or calibrated by the historical common factor failure rate.
[0135] For example, two nodes within the ring have failure risk values of P1=0.1 and P2=0.2, and belong to the same provider. Then P cb ≈0.28. When β=0.15 and δ=1, P c ≈1-(1-0.28)×0.85≈0.39. It can be seen that after common-cause correction, the risk value of loop failure is increased.
[0136] In one possible implementation, the method may further include failure scenario simulation and governance ranking. Failure scenario simulation is used to simulate the impact of candidate critical nodes failing on domain name resolution, and outputs a governance ranking based on the degree of impact.
[0137] Specifically, the DNS risk analysis device 103 can identify candidate critical nodes. Candidate critical nodes can be identified from nodes with protocol risks, authoritative server nodes, and nodes with high centrality. For example, the DNS risk analysis device 103 can obtain a set of candidate critical nodes by merging and deduplicating from the following sources: nodes with protocol risks and their associated authoritative server nodes; all nodes involved in authoritative server relationships; and the top-N nodes ranked by degree centrality, where N is configurable, for example, N=25. The set of candidate critical nodes after deduplication from these three sources can be set to a maximum of 50.
[0138] For each candidate critical node, the DNS risk analysis device 103 can simulate its failure. Specifically, the DNS risk analysis device 103 can mark the candidate critical node as failed in the directed graph, and starting from the candidate critical node, traverse the nodes that depend on it to obtain the set of affected nodes. After obtaining the set of affected nodes, the DNS risk analysis device 103 can calculate the impact scale and impact ratio of the affected node set. The impact scale is the number of affected nodes in the affected node set. The impact ratio is the ratio of the number of affected nodes in the affected node set to the total number of DNS resolution objects that depend on the candidate critical node. This simulation process is performed on the graph and does not involve changes to the production DNS environment. The simulation results are used for governance priority ranking, not for automatic remediation.
[0139] For each domain name node in the affected node set, the DNS risk analysis device 103 can calculate the number of remaining healthy authoritative servers. The basic definition of a healthy authoritative server can be: the domain name has at least one valid authoritative server dependency edge, and the corresponding authoritative server node was not marked as invalid in the simulation. Extended definitions can also include A / AAAA record reachability, whether the DNS response time is below a threshold, and whether the returned content is correct.
[0140] If the number of remaining healthy authoritative servers for an affected domain is 0 after the candidate critical node fails, the DNS risk analysis device 103 will determine the affected domain as a hard failure domain. If the number of remaining healthy authoritative servers is greater than 0, the DNS risk analysis device 103 will determine the affected domain as a downgraded domain.
[0141] Based on the number of hard failure domains and the number of downgraded domains, the DNS risk analysis device 103 can calculate a failure scenario score. For example, the failure scenario score can be calculated using the following formula: ; Where, num f num represents the number of hard failure domains. d P represents the number of downgraded domains, where Ratio is the ratio of the number of affected domains to the total number of domains that depend on the target domain. t This represents the failure risk value of a failed node or ring. α, β, γ, and η are weights. For example, α = 1.0, β = 0.4, γ = 0.3, and η = 0.2, with α being greater than β, making the hard failure weight higher than the degradation weight.
[0142] For example, DNS risk analysis device 103 simulates the failure of authoritative server node N, num f =120, num d =30, ratio=0.15, P t When the value is 0.25, the score is approximately 120 + 12 + 0.045 + 0.05 ≈ 132.1.
[0143] The DNS risk analysis device 103 can calculate failure scenario scores for multiple candidate critical nodes and output the sorted results from highest to lowest failure scenario scores. The sorted results can output Top-K candidate governance objects, where K can be set to 10 by default, or can be configured as needed.
[0144] See Figure 7 , Figure 7 This is a schematic diagram of a fault scenario simulation and governance sorting process provided in an embodiment of this application.
[0145] like Figure 7As shown, the DNS risk analysis device 103 first obtains a set of candidate critical nodes. Then, for each candidate critical node, the DNS risk analysis device 103 marks it as in a failed state and performs dependency propagation based on labeled edges to obtain a set of affected resolution objects. Next, the DNS risk analysis device 103 calculates the impact scale and impact ratio of the set of affected resolution objects. For domain name nodes in the set of affected resolution objects, the DNS risk analysis device 103 calculates the number of remaining healthy authoritative servers after the candidate critical node fails. The DNS risk analysis device 103 distinguishes between hard-failed domains and downgraded domains based on the number of remaining healthy authoritative servers. Afterwards, the DNS risk analysis device 103 calculates a failure scenario score based on the number of hard-failed domains, the number of downgraded domains, etc. Finally, the DNS risk analysis device 103 sorts the failure scenario scores and outputs Top-K candidate governance objects.
[0146] To facilitate understanding of the overall relationship between the various processing steps in the method provided in the embodiments of this application, the following is combined with... Figure 8 Please provide an explanation. See also... Figure 8 , Figure 8 This is a schematic diagram illustrating the overall process of a domain name system risk analysis method provided in an embodiment of this application. Figure 8 As shown, the DNS risk analysis device 103 can first construct a labeled DNS resolution directed graph from domain name resolution data. Then, based on the labeled DNS resolution directed graph, the DNS risk analysis device 103 performs topology detection, node semantic detection, and graph feature detection to obtain hierarchical risk results. The hierarchical risk results are used for root cause attribution and to determine risk weights or failure risk values. Risk weights or failure risk values are used for dependency propagation calculations, and the node information affected by the starting node obtained from the propagation calculations is fed back to root cause attribution. Risk weights or failure risk values are also used for fault scenario simulation, and fault scenario scoring and ranking results are obtained from the fault scenario simulation. The results of root cause attribution are used to generate risk evidence objects. The risk evidence objects and the fault scenario scoring and ranking results are jointly output as risk inference results.
[0147] See Figure 9 , Figure 9 This is a schematic diagram of a domain name system risk analysis device provided in an embodiment of this application. Figure 9 As shown, the Domain Name System Risk Analysis Device 900 may include an acquisition module 901, a detection module 902, and a traversal module 903.
[0148] The acquisition module 901 is used to acquire a directed graph constructed based on domain name resolution data. The directed graph includes various types of nodes and various types of edges. The edges have labels, which are used to indicate the type of association between different nodes. The nodes include nodes that indicate the domain names recorded in the domain name resolution data, as well as nodes that indicate other types of information associated with the domain names.
[0149] The detection module 902 is used to detect the directed graph and obtain the starting node. The risk of the starting node causing domain name resolution failure when it participates in domain name resolution is higher than the risk threshold.
[0150] The traversal module 903 is used to traverse the directed graph starting from the starting node, based on the labels of the edges connected to the starting node, to obtain multiple target nodes affected by the starting node. The number of hops between each target node and the starting node does not exceed a hop count threshold, which is determined based on the edge labels.
[0151] In this way, by having the acquisition module 901, the detection module 902, and the traversal module 903 work together, other nodes affected by a node with resolution risk can be identified together. This expands the risk analysis object from a single isolated risk node to a group of risk nodes associated along the resolution dependency relationship, providing a set of associated nodes for domain name system risk analysis.
[0152] In one possible implementation, the traversal module 903 is specifically used to determine the hop count threshold based on the label of the edge connected to the starting node; starting from the starting node, traversing the directed graph according to the hop count threshold to obtain nodes whose hop count to the starting node does not exceed the hop count threshold.
[0153] In one possible implementation, the traversal module 903 is specifically configured to, when the edges connected to the starting node include a first edge and a second edge, determine a first hop count threshold based on the label of the first edge connected to the starting node; determine a second hop count threshold based on the label of the second edge connected to the starting node; traverse the directed graph from the starting node along the direction indicated by the first edge according to the first hop count threshold to obtain a first target node; and traverse the directed graph from the starting node along the direction indicated by the second edge according to the second hop count threshold to obtain a second target node; the multiple target nodes include the first target node and the second target node.
[0154] In one possible implementation, the detection module 902 is specifically used to detect the directed graph, obtain multiple candidate nodes and the type of risk that each candidate node will cause domain name resolution failure when participating in domain name resolution; determine the failure risk value of each candidate node according to the risk type and mapping relationship corresponding to each candidate node; and select nodes whose failure risk value exceeds the risk threshold from multiple candidate nodes according to the failure risk value of each candidate node to obtain the starting node.
[0155] In one possible implementation, the detection module 902 is specifically used to detect the topological structure of the directed graph to obtain multiple candidate nodes; or, to detect the semantics of the nodes in the directed graph to obtain multiple candidate nodes; or, to detect the graph features of the directed graph to obtain multiple candidate nodes.
[0156] In one possible implementation, the detection module 902 is specifically used to query the basic risk value corresponding to the type of each of the at least two risks according to the mapping relationship when multiple candidate nodes include a first candidate node and the first candidate node corresponds to at least two risks, and determine the failure risk value of the first candidate node according to the basic risk value corresponding to the type of the at least two risks.
[0157] In one possible implementation, the domain name system risk analysis device 900 may further include a root cause attribution module, an evidence chain generation module, and a failure scenario simulation module, which are used to perform the aforementioned supplementary implementations of root cause attribution, evidence chain generation, and failure scenario simulation, respectively.
[0158] From another perspective of engineering function division, the related functions of the Domain Name System Risk Analysis Device 900 can also be implemented by a graph construction module, a topology detection module, a node semantic detection module, a graph feature detection module, a failure risk value determination module, a risk propagation module, a root cause attribution module, an evidence chain generation module, and a fault scenario simulation module. These functional modules can be deployed in the same service or distributed collaboratively through a database, message queue, graph computing service, or program interface. The above functional division can be further subdivided from the functions corresponding to the acquisition module 901, detection module 902, and traversal module 903, without requiring... Figure 9 The system allows for the setting of a corresponding number of independent modules, without requiring each functional module to be set up as an independent physical module.
[0159] In one possible implementation, the acquisition module 901, detection module 902, and traversal module 903 can be deployed in the same service, or they can be distributed and coordinated through a database, message queue, graph computing service, or program interface. In actual deployment, the functions of the above modules can be merged into a single processing unit, or the modules can be deployed in different processing units or computing devices.
[0160] In another possible implementation, the Domain Name System Risk Analysis Device 900 can also be implemented as a software module and deployed on a computing device, and executed by the processor in the computing device to implement the Domain Name System Risk Analysis Method provided in any of the foregoing embodiments.
[0161] Since the Domain Name System Risk Analysis Device 900 corresponds to the DNS Risk Analysis Device 103 in the aforementioned method embodiment, therefore, Figure 9 The process by which each module of the domain name system risk analysis device 900 executes the above method and the technical effects thereof can be found in the relevant descriptions in the foregoing method embodiments, and will not be repeated here.
[0162] Figure 9 The module division in this framework is a functional division, and does not require each module to be implemented by independent physical devices. Two or more modules can be integrated, and the functionality of one module can be implemented by multiple modules. The connections between modules indicate functional associations, and do not restrict the order of calls or the direction of data transfer.
[0163] See Figure 10 The diagram shows a hardware structure schematic of a computing device provided in an embodiment of this application. Figure 10 The computing device 1000 shown can be used to implement the DNS risk analysis device 103 mentioned in the above embodiments. For example... Figure 10 As shown, the computing device 1000 includes a bus 1001, a processor 1002, a memory 1003, and a communication interface 1004. The processor 1002, the memory 1003, and the communication interface 1004 communicate with each other via the bus 1001.
[0164] Bus 1001 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 1001 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0165] The processor 1002 may be a central processing unit (CPU). The memory 1003 may include volatile memory, such as random access memory (RAM); it may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD). The memory 1003 stores executable code, which the processor 1002 executes to implement the steps performed by the DNS risk analysis device 103 in the aforementioned domain name system risk analysis method.
[0166] This application also provides a computer-readable storage medium storing instructions that, when run on a computing device, cause the computing device to perform the steps executed by the DNS risk analysis device 103 in the above-described Domain Name System risk analysis method.
[0167] This application also provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they implement the steps performed by the DNS risk analysis device 103 in the above-described domain name system risk analysis method.
[0168] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for risk analysis of the Domain Name System, characterized in that, The method includes: Obtain a directed graph constructed based on domain name resolution data. The directed graph includes multiple types of nodes and multiple types of edges. The edges have labels, which are used to indicate the type of association between different nodes. The nodes include nodes used to indicate the domain names recorded in the domain name resolution data, and nodes used to indicate other types of information associated with the domain names. The directed graph is inspected to obtain the starting node. The risk of the starting node causing domain name resolution failure when it participates in domain name resolution is higher than the risk threshold. Based on the labels of the edges connected to the starting node, the directed graph is traversed starting from the starting node to obtain multiple target nodes affected by the starting node. The number of hops between each of the multiple target nodes and the starting node does not exceed a hop count threshold, which is determined based on the edge labels.
2. The method according to claim 1, characterized in that, The step of traversing the directed graph from the starting node based on the labels of the edges connected to the starting node to obtain multiple target nodes affected by the starting node includes: The hop count threshold is determined based on the label of the edge connected to the starting node; Starting from the starting node, the directed graph is traversed according to the hop count threshold to obtain nodes whose hop count to the starting node does not exceed the hop count threshold.
3. The method according to claim 1, characterized in that, The edges connected to the starting node include a first edge and a second edge, which are of different types. The step of traversing the directed graph from the starting node based on the labels of the edges connected to the starting node to obtain multiple target nodes affected by the starting node includes: Determine the first hop count threshold based on the label of the first side connected to the starting node; The second hop count threshold is determined based on the label of the second side connected to the starting node; Starting from the starting node, according to the first hop count threshold, the directed graph is traversed along the direction indicated by the first edge to obtain a first target node whose hop count from the starting node does not exceed the first hop count threshold; Starting from the starting node, according to the second hop count threshold, the directed graph is traversed along the direction indicated by the second edge to obtain a second target node whose hop count from the starting node does not exceed the second hop count threshold; The plurality of target nodes includes the first target node and the second target node.
4. The method according to claim 1, characterized in that, The step of detecting the directed graph to obtain the starting node includes: The directed graph is inspected to obtain multiple candidate nodes and the type of risk that each candidate node may cause domain name resolution failure when participating in domain name resolution. Based on the risk type and mapping relationship corresponding to each candidate node, the failure risk value of each candidate node is determined, wherein the mapping relationship indicates that the risk type corresponds to the basic risk value, and the failure risk value is determined based on the risk type of the candidate node and the basic risk value corresponding to the risk type. Based on the failure risk value of each candidate node, nodes whose failure risk value exceeds the risk threshold are selected from the multiple candidate nodes to obtain the starting node.
5. The method according to claim 4, characterized in that, The process of detecting the directed graph yields multiple candidate nodes, including: The topology of the directed graph is detected to obtain multiple candidate nodes, including nodes in the topology that cause domain name resolution failure. Alternatively, the semantics of the nodes in the directed graph can be detected to obtain multiple candidate nodes. The multiple candidate nodes include nodes whose semantics do not conform to the Domain Name System protocol and nodes with abnormal configuration semantics. The nodes with abnormal configuration semantics have an increased risk of domain name resolution failure due to the configuration semantics. Alternatively, the graph features of the directed graph can be detected to obtain multiple candidate nodes. The graph features include the position, connection method, or attribute distribution of the nodes in the directed graph. The quantization value corresponding to each candidate node exceeds the quantization threshold. The quantization value corresponding to the candidate node is determined based on the graph features of the candidate node.
6. The method according to claim 4, characterized in that, Based on the risk type and mapping relationship corresponding to each candidate node, the failure risk value of each candidate node is determined, including: The plurality of candidate nodes includes the first candidate node; When the first candidate node corresponds to a single risk, the basic risk value corresponding to the type of the single risk is queried according to the mapping relationship, and the failure risk value of the first candidate node is the basic risk value corresponding to the type of the single risk. When the first candidate node corresponds to at least two risks, the basic risk value corresponding to the type of each of the at least two risks is queried according to the mapping relationship, and the failure risk value of the first candidate node is determined according to the basic risk value corresponding to the type of the at least two risks.
7. A domain name system risk analysis device, characterized in that, include: The acquisition module is used to acquire a directed graph constructed based on domain name resolution data. The directed graph includes multiple types of nodes and multiple types of edges. The edges have labels, which are used to indicate the type of association between different nodes. The nodes include nodes that indicate the domain names recorded in the domain name resolution data, and nodes that indicate other types of information associated with the domain names. The detection module is used to detect the directed graph and obtain the starting node. The risk of the starting node causing domain name resolution failure when participating in domain name resolution is higher than the risk threshold. The traversal module is used to traverse the directed graph starting from the starting node based on the labels of the edges connected to the starting node, to obtain multiple target nodes affected by the starting node, wherein the number of hops between each of the multiple target nodes and the starting node does not exceed a hop count threshold, and the hop count threshold is determined based on the edge labels.
8. A computing device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method of any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.