DNS anomaly detection and domain name hijacking early warning method based on multi-source NLP
Through the fusion analysis of multi-source NLP intelligence collection and DNS anomaly detection, a threat intelligence knowledge map is built, which solves the timeliness and accuracy of domain name hijacking recognition in the existing technology, and realizes early identification and efficient protection of new attacks.
Patent Information
- Application Number
- CN202510633835.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing technology is difficult to identify complex and changeable domain name hijacking attacks in a timely manner, with high false alarm rates and missed response rates, lacking real-time early warning capabilities, and unable to effectively utilize external intelligence for dynamic correlation analysis.
Using multi-source NLP fusion analysis technology, through multi-source intelligence acquisition and NLP semantic analysis, a threat intelligence knowledge graph is built, real-time data fusion and cross-modal analysis are carried out, abnormal detection and early warning are carried out in combination with DNS records, and a self-learning closed-loop optimization model is established.
It realizes early identification of new domain name hijacking attacks, reduces false alarm rates and missed alarm rates, improves detection efficiency and early warning accuracy, and adapts to protection capabilities for complex scenarios.
Smart Images

Figure CN120389896A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, relates to DNS security detection, and specifically relates to a DNS anomaly detection and domain name hijacking warning method based on multi-source NLP. Background Technique
[0002] The Domain Name System (DNS) is used to map domain names and IP addresses to each other and is an important basic service of the Internet. Mining intelligence from network data refers to obtaining threat clues or attack information related to network security from public or private data sources, which is common in security forums, hidden networks (such as the dark web), social media, etc. Domain name hijacking refers to an attack behavior where, by tampering with DNS resolution records or other means, the domain name accessed by a user points to an illegal or malicious IP address. Anomaly detection refers to performing statistical analysis or machine learning analysis on the collected DNS resolution logs or records to identify suspicious events that are significantly different from the normal behavior pattern.
[0003] Regarding DNS security, the common practice in the industry currently is to use traditional DNS monitoring tools, such as passive DNS log analysis and simple black and white lists, for domain name security management. The process of traditional DNS monitoring includes: (1) collecting passive DNS data and counting information such as query volume and query source IP; (2) comparing with a simple malicious domain name blacklist or domain name reputation database; (3) if a match is found in the blacklist or DNS resolution volume anomaly is detected, it is prompted that there may be domain name hijacking or malicious domain name activities. This DNS security monitoring technology has the following disadvantages: (1) lack of external intelligence support: relying only on internal DNS logs and simple blacklists, it is impossible to timely perceive new or hidden domain name hijacking methods; (2) high false negative rate and false positive rate: the black and white lists are not updated in a timely manner and the rules are rough, easily leading to a large number of false positives or false negatives; (3) difficult to perform dynamic correlation analysis: traditional tools are mainly based on static matching and are often unable to handle complex domain name hijacking attacks, such as those through multi-level CNAME jumps or dark web instruction distribution.
[0004] Some security vendors are trying to correlate simple intelligence data with DNS logs for DNS monitoring. The process includes: (1) Regularly (such as daily or weekly) crawling public security intelligence information (such as some forums, CVE announcements, etc.) to obtain threat clues related to domain name attacks; (2) Matching the obtained intelligence, such as suspicious IPs, domain names, or attacker IDs, with internal DNS query data; (3) If a match occurs, the system gives an alarm, indicating that there may be domain name abuse or hijacking. However, this technical solution has the following disadvantages: (1) Single intelligence source and delayed update: Most only refer to a small amount of public data, and the update cycle is long; (2) Unable to perform in-depth semantic analysis: Lack of NLP (Natural Language Processing) processing ability, unable to extract key attack methods or potential risk warnings from massive texts; (3) Lack of real-time warning ability: Intelligence and DNS logs are usually stored separately and compared manually at regular intervals, unable to detect quickly emerging domain name hijacking attacks in a timely manner. Summary of the Invention
[0005] In today's Internet, DNS domain name hijacking has become a common and extremely harmful attack method. Existing technologies either rely solely on internal DNS monitoring or simply interface with a small amount of external intelligence, making it difficult to identify complex and ever-changing hijacking attack scenarios in a timely manner. Therefore, the present invention provides a DNS anomaly detection and domain name hijacking warning method based on multi-source NLP, which realizes the rapid discovery, accurate positioning, and timely warning of domain name hijacking and DNS anomalies, and reduces false positives and false negatives.
[0006] A DNS anomaly detection and domain name hijacking warning method based on multi-source NLP provided by the present invention includes the following steps:
[0007] Step 1: The multi-source intelligence collection and NLP semantic analysis module first conducts intelligent collection of multi-source heterogeneous intelligence, and then performs NLP semantic analysis and intelligence structuring. The intelligent collection of multi-source heterogeneous intelligence includes: setting a dynamic crawler policy driven by reinforcement learning, regularly crawling multi-modal data related to network security from preset websites, performing cross-modal fusion on the multi-modal data, and using TextGrad adversarial training to clean the text of the multi-modal data. The NLP semantic analysis and intelligence structuring include: using the fused features and the multi-modal data after text cleaning as input, on the one hand, identifying small-sample threat entities and malicious entities among them, and on the other hand, performing causality-driven topic clustering to identify the topics among them, constructing the evolution path between the topics, and adding them to the threat intelligence knowledge graph to infer the attack patterns of malicious entities. The topic refers to the entities involved in the threat attack pattern, including APT organization names, CVE numbers, malicious domain names, IP addresses, and attack tools; the nodes in the threat intelligence knowledge graph are different entities, and the edges are the relationships between entities, and the relationships include usage, attack, and belonging. Send the detected malicious entities, entity relationships, and threat confidence scores to the intelligence data analysis and hijacking determination module.
[0008] Step 2: The DNS record collection and anomaly detection module regularly obtains DNS query and response records for anomaly detection.
[0009] Step 3: The intelligence correlation analysis and hijacking determination module collects the detected malicious entities from the multi-source intelligence collection and NLP semantic analysis module, and collects the anomaly patterns detected within the set time window from the DNS record collection and anomaly detection module, calculates the correlation score between the malicious entities and the anomaly patterns. If the score exceeds the preset threshold, it is determined that there is a domain name hijacking or DNS poisoning behavior, and the determination result and recommended disposal instructions and defense actions are output to the domain name hijacking warning and disposal module; the disposal instructions include the list of domain names or / and IPs to be intercepted and the risk level, and the defense actions include DNS redirection, certificate revocation, and traffic cleaning.
[0010] Step 4: After receiving the output from the intelligence correlation analysis and hijacking determination module, the domain name hijacking warning and disposal module executes: when it receives a determination that a domain name or / and IP has a high risk, it triggers an automatic alarm mechanism; performs defense actions according to the set strategy; stores the log information related to abnormal events.
[0011] Step 5: The continuous learning and dynamic update module collects the feedback of the system alarm results, collects the labeled samples detected and output by the intelligence correlation analysis and hijacking determination module, synchronizes the collaborative defense and the log information related to abnormal events, and updates the multi-source intelligence collection and NLP semantic analysis module as well as the DNS anomaly detection model.
[0012] The advantages and positive effects of the present invention are as follows:
[0013] (1) The method of the present invention uses the NLP fusion analysis technology of multi-source heterogeneous intelligence to achieve the joint parsing of text, code snippets and images of the target crawled website, and adopts an adversarial text cleaning algorithm to clean the collected fusion text, solve the spelling mistakes and Unicode disguises deliberately implanted by attackers, and achieve the unified encoding of cross-language entities in the hyperbolic space (Poincaré sphere model), perform entity semantic alignment, and construct a threat intelligence knowledge graph; fuse the NLP intelligence and the real-time data of DNS anomaly detection to achieve the automatic mapping and comparison of multi-source intelligence and DNS records in the time and space dimensions, which can improve the detection efficiency of real-time DNS anomaly detection.
[0014] (2) The method of the present invention can realize the early identification of new domain name hijacking attacks. Through multi-source NLP intelligence mining, the system can timely discover the latest hijacking methods or target domain names that appear on the dark web and conduct early protection and control.
[0015] (3) The method of the present invention can significantly reduce the false alarm rate and missed alarm rate. Through the correlation analysis of external intelligence and internal DNS data, a large amount of irrelevant information can be filtered, the accuracy of alarms can be improved, and the interference to normal business can be reduced.
[0016] (4) The method of the present invention establishes a self-learning closed loop. By combining the real-time DNS anomaly detection results collected with intelligence updates, the model is continuously iteratively optimized to enhance the adaptability to various complex domain name hijacking scenarios.
[0017] (5) The method of the present invention supports multiple application scenarios and is applicable to various business scenarios such as internal DNS protection of enterprises, root DNS security monitoring of Internet service providers, and domain name risk perception of large-scale content delivery networks (CDNs) or cloud platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the implementation flowchart of the DNS record anomaly detection and domain name hijacking early warning method of the present invention;
[0019] Figure 2 is a schematic diagram of multi-source heterogeneous data collection by the multi-source intelligence collection and NLP semantic analysis module of the present invention;
[0020] Figure 3 is an iterative schematic diagram of the self-learning closed loop of the continuous learning and dynamic update module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0022] As Figure 1As shown in the figure, the embodiment of the present invention implements the proposed DNS anomaly detection and domain name hijacking warning method based on multi-source NLP into several functional modules and deploys them on application devices. The main functional modules include: multi-source intelligence collection and NLP semantic analysis module, intelligence data analysis and hijacking determination module, domain name hijacking warning and handling module, continuous learning and dynamic update module, DNS record collection and anomaly detection module, and security management platform. The implementation of the DNS anomaly detection and domain name hijacking warning method based on multi-source NLP in the embodiment of the present invention includes the following steps.
[0023] Step 1, the multi-source intelligence collection and NLP semantic analysis module first conducts intelligent collection of multi-source heterogeneous intelligence, and then conducts NLP semantic analysis and intelligence structuring processing. Among them, the intelligent collection of multi-source heterogeneous intelligence refers to regularly crawling multi-modal data related to network security from preset websites, such as Figure 2 shown, the preset crawl source set, which includes dark web markets and forums, social media platforms, vulnerability database websites, and hacker tool repository websites, etc. Crawl multi-modal data from these websites, perform cross-modal fusion and adversarial text cleaning on the data, and set a dynamic crawler strategy to crawl multi-modal data to obtain threat clues or attack information related to network security, including the following steps 1.1 - 1.3. The NLP semantic analysis and intelligence structuring processing includes small sample threat entity recognition and causality-driven topic clustering, including steps 1.4 - 1.5.
[0024] Step 1.1: Perform cross-modal fusion on the multi-modal data as follows:
[0025] F(x) = Transformer(DeBERTa(x image )||CodeBERT(x code ));
[0026] Among them, F(x) is the feature representation after fusing the crawled multi-modal data x, and x contains text and code data; Transformer is the attention mechanism model for cross-modal feature fusion; x_image is the text content in the screenshot of the crawled page image, and in the embodiment of the present invention, it is the text content extracted by OCR from the screenshot of the dark web commodity page; x_code is the code content of the crawled page, and in the embodiment of the present invention, it is the code snippet in the dark web page, such as the source code of the DNS hijacking tool; DeBERTa(x_image) is the image modality encoder, which extracts features by processing the OCR text of the picture through the DeBERTa model; CodeBERT(x_code) is the code modality encoder, which extracts features by parsing the code snippet through the CodeBERT model; || represents the vector concatenation operator.
[0027] In the embodiments of the present invention, the commodity pages of the dark web market are parsed through triple-modal joint encoding to identify mixed content (text + screenshots + code snippets) containing domain hijacking tools / services.
[0028] Step 1.2: Perform adversarial text cleaning. In the embodiments of the present invention, TextGrad adversarial training is introduced to enhance the robustness of detection. The objective function of the adversarial training is as follows:
[0029] min θ max ||δ||≤ε L(f θ (x + δ), y);
[0030] where θ represents the parameters of the text cleaning model; δ is the adversarial perturbation vector, that is, the noise added to the input text, ε is the perturbation intensity threshold, and ||δ|| ≤ ε is the L2 norm constraint of the perturbation vector; L is the loss function of the model on the input after adding the perturbation. In the embodiments of the present invention, the cross-entropy loss is calculated; f θ is the text cleaning model with parameters; x represents the original input text, and y represents the label of the clean text after cleaning x. Through adversarial text cleaning, the spelling mistakes deliberately implanted by the attacker and the Unicode disguise problem are solved to effectively defend against the homographs used by the attacker, such as the apple domain name disguised as "аррlе.com".
[0031] Step 1.3: Set the currently crawled website as the target website and set a dynamic crawler policy driven by reinforcement learning. The crawler scheduling policy is as follows:
[0032]
[0033] where the policy function π s (a|s) is the probability distribution of selecting action a in state s; Q(s, a) is the state-action value function, which is used to evaluate the benefit of action a in state s, τ is the temperature coefficient, which is used to control the balance between exploration and exploitation; s is the state vector. In the embodiments of the present invention, s includes 15-dimensional features such as the anti-crawling intensity of the target website and the page update frequency; the action a in the embodiments of the present invention is to select a page parsing path at a certain crawling frequency, and the action space is the set of different page parsing paths at different crawling frequencies; the denominator in the formula represents the sum of the exponential function exp(Q(s, a′) / τ for all possible actions a′. In the embodiments of the present invention, the optimal parsing path and crawling frequency are selected through the dynamic crawler policy driven by reinforcement learning.
[0034] Step 1.4: Perform small-sample threat entity recognition on the fused features and the cleaned input text.
[0035] a. Construct a dynamic prompt learning framework for identifying malicious domains. The dynamic prompt learning framework constructed in the embodiments of the present invention is a two-stage architecture including a template generator and a FLAN-T5 inference engine. Among them, the dynamic template generator automatically adjusts the prompt template according to the input text features, and the FLAN-T5 model performs the malicious domain name classification task based on the generated prompt, which can achieve high-accuracy malicious domain name identification in the zero-shot scenario.
[0036] b. Conduct contrastive learning on the dynamic prompt learning framework to enhance the recognition results of the dynamic prompt learning framework. The NT-Xent loss function L is used during contrastive learning NTXent as follows:
[0037]
[0038] where z i , z j is the feature vector of a positive sample pair, such as the embedding vectors of "DNS hijacking" and "DNS hijacking"; 2N is the total number of samples in the batch, which is the augmented data containing N pairs of positive samples; sim(z i , z j ) is the cosine similarity function, which calculates the similarity between samples z i , z j ; r is the temperature coefficient, which is used to scale the range of similarity values; 1_{k≠i} is the indicator function, which is 1 when k≠i and 0 otherwise. The NT-Xent loss function promotes the model to learn discriminative feature representations by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs.
[0039] When training the dynamic prompt learning framework, the present invention constructs positive sample pairs with different language representations of the same threat entity, such as the positive sample pair "DNS hijacking DNS hijacking", to conduct contrastive learning to improve the cross-language entity alignment ability.
[0040] Step 1.5: Discover the entities in the threat attack patterns involved by the causal discovery algorithm for the fused features and the cleaned text, and automatically construct the evolutionary path between topics, such as generating the evolutionary path of "vulnerability → attack tool development → attack activity" for malicious domain name attacks.
[0041] Construct a neural causal topic model as follows:
[0042]
[0043] Among them, p(w|d) represents the probability of generating a cybersecurity-related word w such as DNS, vulnerability, etc. from the crawled data document d such as threat intelligence text; z is a latent variable representing a topic, such as "attack activity", "tool development", etc.; T is the total number of preset topics; p(w|z=t) represents the generation probability of word w under topic t; p(z=t|d) represents the probability that document d belongs to topic t (constrained by the causal graph). The topic in the embodiment of the present invention refers to the entity involved in the threat attack pattern, including the name of the APT (Advanced Persistent Threat) organization, the CVE (Common Vulnerability Exposure) number, malicious domain names, IP addresses, attack tools, etc.
[0044] Construct a time-series topic drift detection model as follows:
[0045] Δ t = ||θ t - θ t-1 ||2 > α · mad({Δ t-k});
[0046] Among them, θ t , θ t-1 are the topic distribution vectors at time t and time t-1 respectively; Δ t represents the topic drift amount at time t; ||.||2 is used to calculate the L2 norm, which is used to measure the amplitude of distribution change; mad({Δ t-k}) represents the median absolute deviation of the change values within the historical window; α is the early warning sensitivity adjustment coefficient. For example, when α = 3, it is a significant anomaly. The time-series topic drift detection model determines whether to trigger an early warning by calculating the change of the topic distribution.
[0047] The present invention first extracts the involved topics from the input using the neural causal topic model, then uses the time-series topic drift detection model to detect the change of the topics, and finally determines the causal relationship between the topics through the causal discovery algorithm. When the topic distribution change exceeds α times the historical median absolute deviation, an early warning is triggered. For example, when it is detected that the intensity of a topic such as tool development increases by more than a certain threshold such as 200% within the time window [t-1, t], the associated CVE number is automatically retrieved, and the relationship edge of "potential attack tool → exploitation → vulnerability" is created in the threat intelligence knowledge graph.
[0048] Step 1.6: Construct a threat intelligence knowledge graph. In the threat intelligence knowledge graph of the present invention, the nodes are the entities involved in the threat attack pattern, including the name of the APT organization, the CVE number, malicious domain names, IP addresses, attack tool names, etc.; the edges are based on the actual relationships between these entities, such as use, attack, belong to, etc.; the relationships formed by two nodes and edges, for example, an APT organization uses a certain vulnerability, the vulnerability is exploited to attack a certain domain name, and the domain name resolves to a certain IP address, etc.
[0049] When constructing the threat intelligence knowledge graph, the present invention performs hyperbolic space entity alignment on the representation forms of different grammars and different encodings of node entities as follows:
[0050]
[0051] Among them, u and v are two different embedded vector representations of entities in the Poincaré ball model, such as the embedded vector representations of threat entities described in different languages such as Chinese, Russian, and English; ||u|| is the Euclidean norm of calculating vector u; arcosh is the inverse hyperbolic cosine function, calculating the hyperbolic space distance d H (u, v).
[0052] Realize the unified representation of Chinese / Russian / English multilingual threat entities in the Poincaré ball model.
[0053] Use the constructed threat intelligence knowledge graph for attack pattern reasoning: adopt the RotatE relationship reasoning model: h ° r ≈ t, and automatically complete the attack chain of "APT organization → use → vulnerability → attack → domain name"; among them, h is the head entity embedding, such as "APT organization", r is the relationship embedding, such as "use"; t is the tail entity embedding, such as "vulnerability"; ° represents the complex space rotation operator, and rotates the head entity h to near the tail entity t through the relationship r.
[0054] The construction data of the initial threat intelligence knowledge graph of the present invention comes from two parts: the externally collected multi-source heterogeneous intelligence and the internal DNS monitoring data (DNS logs and records, anomaly detection results, etc.). After NLP processing, entities and relationships are extracted to construct the threat intelligence knowledge graph, and the graph will be continuously updated as the collection and DNS monitoring data change.
[0055] Such as Figure 1 As shown, the multi-source intelligence collection and NLP semantic analysis module obtains the operation status of this module, including monitoring indicators such as data collection throughput, NLP model load rate, threat intelligence update frequency, etc., summarizes the currently detected threat intelligence, including the threat type distribution, high-risk domain name / IP list, intelligence source credibility evaluation results, etc. statistically in the time dimension, and outputs them to the security management platform. The multi-source intelligence collection and NLP semantic analysis module provides structured threat data, such as malicious entities, threat semantic tags, and threat confidence scores extracted from, to the intelligence correlation analysis and hijacking determination module. Malicious entities such as domain names, IPs, attack tool names, etc., threat semantic tags such as DNS hijacking, phishing attacks, etc., and the threat confidence score is a risk quantification value based on NLP semantic analysis, with a value range of 0-1.
[0056] Step 2, the DNS record collection and anomaly detection module periodically performs anomaly detection on DNS records.
[0057] The DNS record collection and anomaly detection module is deployed in the log collection systems such as the operator DNS, root DNS, and authoritative DNS, and obtains DNS query and response records in real time or in batches, such as different types of records like A, CNAME, NS, MX, AAAA, etc.; then uses time series analysis, statistical features, or machine learning models to perform anomaly detection on DNS records to detect abnormal situations such as sudden increase in suspicious traffic, abnormal resolution pointing, and uncommon TTL changes. Above, A represents that this record contains the IPv4 address of the domain; CNAME represents that this record maps a domain or subdomain to another domain; NS represents that this record describes the authoritative name server of the domain; MX represents that this record contains information about the email server responsible for receiving emails for a specific domain; AAAA represents that this record contains the IPv6 address of the domain.
[0058] In step three, the intelligence data analysis and hijacking determination module cross-domain associates malicious entities such as suspicious domain names / IPs / keywords extracted by the multi-source intelligence collection and NLP semantic analysis module with the anomaly detection results in the DNS record collection and anomaly detection module, and at the same time performs entity association scoring based on cosine similarity. The cross-domain here refers to "the heterogeneous data association between the external threat intelligence domain (text semantic space) and the internal DNS log domain (network traffic space)". The associated objects are: a. malicious entities in external intelligence, and domain names / IPs / attacker IDs; b. resolution records in DNS anomaly events, such as A / CNAME / NS, etc. If the matching degree or risk score exceeds the preset threshold, it is determined that there may be domain name hijacking or DNS poisoning behavior, and the hijacking warning process is entered.
[0059] The intelligence data analysis and hijacking determination module realizes the correlation analysis of external intelligence and internal DNS data by using time-space alignment, multi-modal feature fusion, causal effect calculation, and comprehensive reproduction scoring. The specific implementation is as follows:
[0060] (1) Time-space alignment: Establish a sliding window matching rule for the external intelligence timestamp \(t_i\) and the DNS anomaly timestamp \(t_d\);
[0061] \(\vert t_i - t_d\vert\leq\Delta t\);
[0062] Among them, \(t_i\) represents the timestamp of external intelligence, \(t_d\) represents the time when the DNS anomaly occurs, and \(\Delta t\) is the time window threshold. In the embodiment of the present invention, \(\Delta t\) is set to 24 hours here, that is, the abnormal patterns detected by the DNS data collected within 24 hours in the present invention.
[0063] (2) Multi-modal feature fusion: Concatenate the external intelligence threat vector \(v_i\) and the DNS anomaly pattern vector \(v_d\) to form a fusion vector.
[0064] v_{fused} = Concat(v_i, v_d);
[0065] Among them, Concat(·, ·) represents the vector concatenation operation; the generated fused vector v_{fused} contains the information of the two vectors. The threat vector v_i of external intelligence is the word vector of the detected malicious entity, and the DNS anomaly pattern vector v_d is the word vector of the detected DNS anomaly pattern.
[0066] (3) Causal effect calculation: Quantify the causal impact of external threats on DNS anomalies through counterfactual reasoning. The causal effect CausalEffect generated by external threats on DNS anomalies is calculated as follows:
[0067] CausalEffect = P(Anomaly|do(Threat = 1)) - P(Anomaly|do(Threat = 0));
[0068] Among them, Threat represents external threats, Anomaly represents DNS anomalies, P(Anomaly|do(Threat = 1)) represents the probability of DNS anomalies occurring under the condition of actively setting Threat to 1, and P(Anomaly|do(Threat = 0)) represents the probability of DNS anomalies occurring under the condition of actively setting Threat to 0. The calculated CausalEffect reflects the strength of the causal relationship between external threats and DNS anomalies.
[0069] (4) Considering both the causal effect and the similarity between vectors, calculate the risk score RiskScore of external threats as follows:
[0070] RiskScore = 0.7 × CausalEffect + 0.3 × cos(v_i, v_d);
[0071] Among them, cos(v_i, v_d) represents the cosine similarity between vectors v_i and v_d, which is used to measure the consistency of their directions. The coefficients 0.7 and 0.3 respectively reflect the weights of the causal effect and vector similarity in the risk score.
[0072] The Information Association Analysis and Hijacking Judgment Module outputs the risk event report and analysis performance indicators to the security management platform; outputs the labeled sample data (verified threat judgment results, difference analysis between model prediction and actual results) and feature distribution deviation (feature distribution change trend of external intelligence and DNS anomaly data) detection to the Continuous Learning and Dynamic Update Module; outputs the disposal instructions, including the list of domain names / IPs to be intercepted and the risk level, and recommended defense actions such as DNS redirection, certificate revocation, traffic cleaning, etc. to the Domain Name Hijacking Early Warning and Disposal Module. The risk event report includes the list of high-risk domain names / IPs with cross-source association results, the attack chain visualization map, and the risk level classification. The analysis performance indicators include performance data such as association analysis accuracy rate, false alarm / missed alarm rate, processing delay, etc.
[0073] Step 4, the Domain Name Hijacking Early Warning and Disposal Module executes: a) When the system determines that a certain domain name or IP has a relatively high risk, trigger the automatic alarm mechanism; b) According to the pre-set policy, execute defense actions such as traffic interception, blocking resolution, and manual secondary review; c) At the same time, store the abnormal event and related log information, and submit it to the security management platform for retention and traceability.
[0074] The functions implemented by the Cooperative Defense and Log Storage Module are: a) Cooperative Defense: Link with other security defense systems (such as firewalls, IDSs, etc.) to timely block malicious IPs or traffic and enhance the defense strength. b) Log Storage and Traceability: Store all alarm information and abnormal event logs to the security management platform to ensure that the information is traceable and support subsequent security analysis and reporting.
[0075] Step 5, the Continuous Learning and Dynamic Update Module executes: a) Collect the feedback of the system alarm results, such as confirming whether it is a false alarm, new changes in attack techniques, etc., and update the multi-source intelligence collection and NLP semantic analysis module and DNS anomaly detection model online or offline; b) Form a self-learning closed loop to continuously improve the ability to capture new types of domain name hijacking attacks.
[0076] Such as Figure 1 And Figure 3As shown in the figure, the continuous learning and dynamic update module performs continuous learning and dynamic iteration to obtain data: obtain labeled sample data and feature distribution offsets from the intelligence correlation analysis and hijacking determination module. The labeled sample data is the verified threat determination result and the difference analysis between the model prediction and the actual result. The feature distribution offset is the change trend of the feature distribution of external intelligence and DNS anomaly data; synchronize all alarm information and anomaly event logs from the collaborative defense and log storage module. Then, update the models involved in the multi-source intelligence collection and NLP semantic analysis module according to the obtained data, and update the DNS anomaly detection model in the DNS record collection and anomaly detection module, so that the system can quickly adapt to new hijacking methods. Verify the alarm results and feedback them to the intelligence data analysis and hijacking determination module.
[0077] The functions implemented by the security management platform in the method of the present invention are: a) Real-time monitoring and reporting: The platform monitors the real-time state of the system and provides a comprehensive security report for easy viewing and analysis by the administrator. b) Event analysis and backtracking: Support the backtracking and in-depth analysis of historical security events to help identify attack patterns and vulnerabilities. At the same time, the security management platform issues policies and receives status. The types of policies issued include: (1) Collection policy: Adjust the priority of intelligence sources, such as increasing the frequency of dark web crawling; set data cleaning rules, such as filtering low-trust sources. (2) Analysis policy: Dynamically adjust the risk determination threshold and time window threshold Δt, such as shortening from 24 hours to 12 hours. (3) Defense linkage policy: Define linkage rules with other security systems (firewall, WAF), such as setting the rule of "triggering firewall IP ban for high-risk domain names". The status data received includes: (1) Module running status: CPU / memory occupancy rate, queue backlog, service availability (heartbeat detection) of each module. (2) Security situation data: Real-time alarm quantity and level distribution, historical attack pattern statistics, such as DNS hijacking attack trend in the past 7 days, and system defense effect indicators such as interception success rate and false interception rate.
[0078] The following takes a specific scenario as an example to illustrate the working process of DNS anomaly detection and domain name hijacking warning in the method of the present invention:
[0079] Step 1, Multi-source intelligence collection and NLP parsing: The system grabs a post on a certain hacker forum discussing how to use similar domain names for phishing attacks. The NLP parsing extracts:
[0080] There are malicious domain names such as "logln.example.com" similar to "log1n" and potential attack IP "198.51.xxx.xxx".
[0081] Step 2, Perform DNS record anomaly detection:
[0082] a) In the DNS logs, recently there have been a large number of requests to resolve "logln.example.com", and the geographical distribution of the source IP addresses of the access is abnormally concentrated in certain high-risk regions;
[0083] b) The IP address pointed to by the resolution partially matches "198.51.xxx.xxx", which causes a preliminary warning from the system.
[0084] Step 3, Intelligence Correlation Analysis:
[0085] a) By comparing with the intelligence database and DNS anomaly rules, it is confirmed that there is a high-risk correlation between the domain name detected in Step 1 and the extracted malicious IP;
[0086] b) The system automatically calculates the risk score and judges that the possibility of it being a suspected phishing domain name or domain hijacking behavior is relatively high.
[0087] Step 4, Early Warning and Disposal:
[0088] a) The system immediately issues a high-level alert and automatically notifies the security administrator;
[0089] b) Subsequently, this domain name can be automatically added to the enterprise internal blacklist or DNS resolution interception can be performed to block user access;
[0090] c) After the administrator conducts a manual review and confirms that it is indeed a malicious counterfeit domain name, the system marks this intelligence entry and provides a positive example feedback to the model.
[0091] Generally speaking, various example embodiments of the present disclosure can be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software executed by a controller, microprocessor, or other computing device. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0092] Except for the technical features described in the specification, they are all known technologies to those skilled in the art. The present invention omits the description of well-known components and well-known technologies to avoid redundancy and unnecessarily limit the present invention. The implementation manners described in the above embodiments do not represent all implementation manners consistent with the present application. Based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative labor are still within the protection scope of the present invention.
Claims
1. A DNS anomaly detection and domain name hijacking warning method based on multi-source NLP, characterized in that, Including: Step 1: The multi-source intelligence collection and NLP semantic analysis module first conducts intelligent collection of multi-source heterogeneous intelligence, and then performs NLP semantic analysis and intelligence structuring processing. Among them, the intelligent collection of multi-source heterogeneous intelligence includes: setting a reinforcement learning-driven dynamic crawler strategy to regularly crawl multi-modal data related to network security from preset websites, performing cross-modal fusion on the multi-modal data, and using TextGrad adversarial training to clean the text of the multi-modal data. The NLP semantic analysis and intelligence structuring processing includes: taking the fused features and the multi-modal data after text cleaning as input. On the one hand, it conducts small-sample threat entity recognition to identify malicious entities among them. On the other hand, it conducts causality-driven topic clustering to identify topics among them, constructs the evolution path between topics, and adds them to the threat intelligence knowledge graph to infer the attack patterns of malicious entities. The topic refers to the entities involved in the threat attack pattern, including APT organization names, CVE numbers, malicious domain names, IP addresses, and attack tools. The nodes in the threat intelligence knowledge graph are different entities, and the edges are the relationships between entities. The relationships include use, attack, and belong to. Among them, APT represents an advanced persistent threat organization, CVE represents Common Vulnerabilities and Exposures, and NLP represents natural language processing technology; Send the detected malicious entities, entity relationships, and threat confidence scores to the intelligence data analysis and hijacking determination module; Step 2: The DNS record collection and anomaly detection module regularly obtains DNS query and response records for anomaly detection; Step 3: The intelligence correlation analysis and hijacking determination module collects the detected malicious entities from the multi-source intelligence collection and NLP semantic analysis module, and collects the anomaly patterns detected within the set time window from the DNS record collection and anomaly detection module, calculates the correlation score between the malicious entities and the anomaly patterns. If the score exceeds the preset threshold, it determines that there is a domain name hijacking or DNS poisoning behavior, and outputs the determination result and recommended disposal instructions and defense actions to the domain name hijacking warning and disposal module. Among them, the disposal instructions include the list of domain names or / and IPs to be intercepted and the risk level, and the defense actions include DNS redirection, certificate revocation, and traffic cleaning; Step 4: After receiving the output from the intelligence correlation analysis and hijacking determination module, the domain name hijacking warning and disposal module executes: when it receives a determination that the domain name or / and IP has a high risk, it triggers an automatic alarm mechanism; executes defense actions according to the set strategy; stores the logs of abnormal events; Step 5: The continuous learning and dynamic update module collects the feedback of the system alarm results, collects the labeled samples detected and output by the intelligence correlation analysis and hijacking determination module, synchronizes the logs related to collaborative defense and abnormal events, and updates the multi-source intelligence collection and NLP semantic analysis module as well as the DNS anomaly detection model.
2. The method according to claim 1, wherein In the said Step 1, the multi-source intelligence collection and NLP semantic analysis module performs cross-modal fusion on the multi-modal data, which is expressed as: F(x) = Transformer(DeBERTa(x image ) || CodeBERT(x code )); Among them, x is the crawled multimodal data, x_image is the text in the screenshot of the crawled page image, and x_code is the code of the crawled page; F(x) is the fused feature; Transformer is the attention mechanism model for cross-modal feature fusion; DeBERTa(x_image) represents extracting the features of x_image through the DeBERTa model; CodeBERT(x_code) represents parsing x_code through the CodeBERT model to extract features; || represents the vector concatenation operator.
3. The method according to claim 1, wherein In step 1, the multi-source intelligence collection and NLP semantic analysis module sets a dynamic crawler policy driven by reinforcement learning. The current website to be crawled is used as the target website, the features of the current target website are used as the state s, and the set of different page parsing paths under different crawling frequencies is used as the action space. The policy function is as follows: where, π s (a|s) is the probability distribution of selecting action a in state s; Q(s,a) is the state-action value function, which is used to evaluate the return of action a in state s, and τ is the temperature coefficient, which is used to control the balance between exploration and exploitation; the denominator in the formula represents the summation of the exponential function exp(Q(s,a′) / τ) for all possible actions a′; the optimal parsing path and crawling frequency of the target website are selected through the dynamic crawler policy driven by reinforcement learning.
4. The method according to claim 1, characterized in that, In step 1, when the multi-source intelligence collection and NLP semantic analysis module performs small-sample threat entity recognition, it uses the constructed dynamic prompt learning framework to identify malicious domains for the input; the dynamic prompt learning framework is a two-stage architecture including a template generator and a FLAN-T5 inference engine. The template generator automatically adjusts the prompt template according to the input text features, and the FLAN-T5 model performs malicious domain classification tasks based on the generated prompts; when training the dynamic prompt learning framework, positive sample pairs with different language representations of the same threat entity are constructed, and the NT-Xent loss function is used to perform contrastive learning on the dynamic prompt learning framework to enhance the recognition results of the dynamic prompt learning framework.
5. The method according to claim 1, characterized in that, In step 1, the multi-source intelligence collection and NLP semantic analysis module extracts the involved topics from the input using the neural causal topic model, and then uses the temporal topic drift detection model to detect topic changes. When the detected topic distribution change exceeds α times the historical median absolute deviation, an alarm is triggered, where α is the alarm sensitivity adjustment coefficient; finally, the causal discovery algorithm is used to determine the causal relationship between topics and construct the evolutionary path between topics.
6. The method according to claim 1 or 5, characterized in that In step 1, when the multi-source intelligence collection and NLP semantic analysis module constructs the threat intelligence knowledge graph, it performs hyperbolic space entity alignment on different encoding representations of the same entity to achieve unified representation of different encoding representations of the same entity in the Poincaré ball model; and based on the existing threat intelligence knowledge graph, the RotatE relation reasoning model is used to perform attack pattern reasoning to complete the attack chain.
7. The method according to claim 1, characterized in that In step 2, the DNS record collection and anomaly detection module is a log collection system deployed on the operator DNS, root DNS, and authoritative DNS, and uses temporal analysis, statistical features, or machine learning models to perform anomaly detection on DNS records.
8. The method according to claim 1, characterized in that In the said step 3, the intelligence correlation analysis and hijacking determination module calculates the correlation score between malicious entities and abnormal patterns, including: Let the timestamp of the malicious entity detected by the multi-source intelligence collection and NLP semantic analysis module be \(t_i\), and the word vector be \(v_i\); the occurrence time of the DNS abnormal pattern collected from the DNS record and the abnormal detection module be \(t_d\), and the word vector be \(v_d\). First, obtain the DNS abnormal pattern based on the matching rule of the set time window threshold \(\Delta t\), satisfying \(|t_i - t_d|\leq\Delta t\). Then, quantitatively calculate the causal effect of the malicious entity on the DNS abnormal pattern through counterfactual reasoning: CausalEffect = P(Anomaly|do(Threat = 1)) - P(Anomaly|do(Threat = 0)). Among them, Threat represents the malicious entity, Anomaly represents the DNS abnormal pattern, P(Anomaly|do(Threat = 1)) represents the probability of the DNS anomaly occurring under the condition of actively setting Threat to 1, and P(Anomaly|do(Threat = 0)) represents the probability of the DNS anomaly occurring under the condition of actively setting Threat to 0. Finally, comprehensively consider the causal effect and the similarity between vectors, and calculate the risk score RiskScore of the malicious entity as follows: RiskScore = 0.7×CausalEffect + 0.3×cos(v_i, v_d). Among them, cos(v_i, v_d) represents the cosine similarity between the calculated vectors \(v_i\) and \(v_d\), and the coefficients 0.7 and 0.3 are the weights of the causal effect and the vector similarity in the risk score respectively.
Citation Information
Patent Citations
Threat early warning and monitoring system and method based on big data analysis and deployment architecture
CN107196910A
Specific network behavior analysis method and system based on multi-source data fusion
CN114500122A
DNS malicious domain name detection system and method based on big data
CN117354024A
Natural language threat intelligence extraction and analysis method and system
CN118627516A
Detection of domain name system hijacking
US20180007088A1
Cited By
Method and system for detecting and tracing domain name resolution hijacking of embedded device
CN120880780A
Multi-modal AI model attack behavior detection and blocking system
CN120979775A
Multimodal ai model attack behavior detection and blocking system
CN120979775B
Electric power information network multi-source threat intelligence analysis method, system, device and medium
CN120979831A
Road accident and illegal parking detection early warning method and system based on intelligent traffic
CN121011078A