Domain name system alarm processing method and system

By receiving alarm messages in the domain name system and using the alarm identification model and case library to generate processing scripts, the problem of low alarm handling efficiency in the domain name system is solved, automatic alarm processing is realized, and the stability and efficiency of the system are improved.

CN120342841APending Publication Date: 2025-07-18CHINA INTERNET NETWORK INFORMATION CENTER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510729018.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the domain name system has low efficiency in handling alarms, mainly due to the long manual analysis and handling process, and operation and maintenance personnel need to deeply understand the system topology and protocol principles, resulting in low efficiency in handling alarms.

Method used

By receiving alarm messages from the domain name system, extracting target information and inputting alarm recognition models, identifying case scenarios of alarm messages, generating alarm processing scripts, and automatically generating disposal scripts using alarm recognition models and disposal command templates in the case library.

Benefits of technology

It realizes automatic identification and handling of alarm messages, reduces manual intervention, improves the stability and reliability of the domain name system, reduces the response time for handling alarms, and avoids the risk of incorrect operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342841A_ABST
    Figure CN120342841A_ABST
Patent Text Reader

Abstract

The invention provides a domain name system alarm processing method and system. The method comprises the following steps: receiving an alarm message of a domain name system; executing target information extraction on the alarm message, wherein the target information comprises an alarm level, an alarm type, system information, host information and an alarm text; inputting the target information into an alarm identification model so as to enable the alarm identification model to identify a case scene of an alarm message; under the condition that the case scene is recognized, an alarm processing script is generated based on the target information and a corresponding processing command template in a case library, and the case library comprises the basic information of the case, the case scene and the processing command template. According to the method, the alarm message can be identified and the corresponding disposal script can be automatically generated, so that the problem of low alarm disposal efficiency caused by long manual analysis disposal process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of domain name services, and particularly to a method and system for processing domain name system alarms. Background Art

[0002] Domain name service is an integral part of the Internet infrastructure. The domain name service constitutes a resolution network through recursive DNS nodes, authoritative DNS nodes, database systems, and distributed server clusters to provide domain name resolution services for users. In this scenario, alarm information will be generated due to hardware devices, software systems, and security incidents. The alarm information contains a large number of professional terms, and the alarm chain often presents a multi-layer architecture correlation.

[0003] The risk determination and production disposal of alarm messages still rely on the professional knowledge and skills of operation and maintenance personnel. The operation and maintenance personnel need to manually analyze the alarm text according to the alarm level, write disposal commands in combination with experience, and review and execute them.

[0004] However, the manual analysis and disposal process is long, resulting in low alarm disposal efficiency. Summary of the Invention

[0005] This application provides a method and system for processing domain name system alarms to solve the problem of low alarm disposal efficiency.

[0006] In a first aspect, this application provides a method for processing domain name system alarms, including:

[0007] Receiving an alarm message of the domain name system;

[0008] Performing target information extraction on the alarm message, where the target information includes alarm level, alarm type, system information, host information, and alarm text;

[0009] Inputting the target information into an alarm recognition model to enable the alarm recognition model to recognize the case scenario of the alarm message;

[0010] In the case of recognizing the case scenario, generating an alarm disposal script based on the target information and the corresponding disposal command template in the case library, where the case library includes the basic information of the case, the case scenario, and the disposal command template.

[0011] In some feasible embodiments, the method further includes:

[0012] In the case of not recognizing the case scenario, marking the alarm message as an unrecognized state and generating a new case;

[0013] Storing the new case in the case library, and recording the number of the new cases and the training cycle of the alarm recognition model;

[0014] If the quantity is greater than or equal to a quantity threshold, and / or, if the training period is greater than or equal to a preset period threshold, trigger the alarm recognition model to perform iterative training;

[0015] Extract the alarm text, alarm level, alarm type, and alarm time of the new case;

[0016] Based on the alarm text, alarm level, alarm type, and alarm time of the new case, update the parameters of the alarm recognition model to generate an updated alarm recognition model.

[0017] In some feasible embodiments, the generating of the alarm handling script includes:

[0018] In the case where the case scenario is not recognized, based on a template addition instruction, generate a handling command template, where the template addition instruction is an instruction generated based on the new case or by replacing dynamic variables in the executed handling command template;

[0019] Match a handling command template based on the case type, and based on the handling command template, generate an alarm handling script.

[0020] In some feasible embodiments, the handling command template includes dynamic variables, and the dynamic variables include a host name, a system identifier, and a time parameter;

[0021] The generating of the handling command template based on the template addition instruction includes:

[0022] Replace corresponding dynamic variables with the host information and system information to generate a command sequence;

[0023] Embed a pre-check instruction and a post-check instruction in the command sequence to generate a template addition instruction, where the pre-check instruction is used to verify execution environment parameters, and the post-check instruction is used to verify the command execution result and trigger an environment recovery operation.

[0024] In some feasible embodiments, the generating of the alarm handling script includes:

[0025] Obtain the number of the case scenario;

[0026] According to the number, query the corresponding alarm handling command group in the case library, where the alarm handling command group contains multiple handling command numbers arranged in the execution order;

[0027] Perform logical splicing on the handling command template generated based on the template addition instruction in the execution order to generate an alarm handling script.

[0028] In some feasible embodiments, the generating of the alarm handling script includes:

[0029] When the case scenario is recognized, obtain the case type of the case scenario;

[0030] Match a disposal command template based on the case type, and generate an alarm disposal script based on the disposal command template.

[0031] In some feasible embodiments, after generating the new case, it further includes:

[0032] Obtain the preset scenario of the new case;

[0033] According to the preset scenario, classify and store the new case in the training set of the case library, and the preset scenario at least includes a first scenario set and a second scenario set;

[0034] If the preset scenario corresponding to the new case is the first scenario set, obtain corpus data;

[0035] Supplement the corpus data to the target training set to generate a supplemented training set, and the target training set is the training set corresponding to the first scenario set;

[0036] Based on the supplemented training set, perform incremental training on the alarm recognition model to generate an alarm recognition model.

[0037] In some feasible embodiments, the generating of the alarm disposal script includes:

[0038] Read the automatic execution flag of the case scenario to generate review information according to the automatic execution flag;

[0039] When the review information is that manual review is required, generate a review request, and the review request represents performing a manual review on the alarm disposal script;

[0040] Obtain review information, and the review information is the information generated after performing a manual review based on the review request;

[0041] Generate an alarm disposal script based on the review information.

[0042] In some feasible embodiments, the method further includes:

[0043] Perform text feature extraction on the alarm text to generate word vectors through a constructed model, and the constructed model is one of a binary bag-of-words model, an N-gram model, and a recurrent neural network;

[0044] Perform feature splicing on the word vectors with the alarm level, alarm type, and alarm time to construct a multi-dimensional joint feature vector;

[0045] Using the random forest algorithm, train the multi-dimensional joint feature vector to generate the alarm recognition model.

[0046] In a second aspect, the present application provides a processing system for domain name system alarms, including:

[0047] A disposal unit for receiving the alarm message of the domain name system;

[0048] An information extraction unit for performing target information extraction on the alarm message, where the target information includes an alarm level, an alarm type, system information, host information, and an alarm text;

[0049] An identification unit for inputting the target information into the alarm recognition model to enable the alarm recognition model to identify the case scenario of the alarm message;

[0050] The disposal unit is further configured to, when the case scenario is recognized, generate an alarm disposal script based on the target information and the corresponding disposal command template in the case library, where the case library includes the basic information of the case, the case scenario, and the disposal command template.

[0051] As can be seen from the above technical solutions, the present application provides a method and a system for processing domain name system alarms. The method includes: receiving the alarm message of the domain name system; performing target information extraction on the alarm message, where the target information includes an alarm level, an alarm type, system information, host information, and an alarm text; inputting the target information into the alarm recognition model to enable the alarm recognition model to identify the case scenario of the alarm message; when the case scenario is recognized, generating an alarm disposal script based on the target information and the corresponding disposal command template in the case library, where the case library includes the basic information of the case, the case scenario, and the disposal command template. The method can identify the alarm message and automatically generate the corresponding disposal script to solve the problem of low alarm disposal efficiency caused by the long manual analysis and disposal process. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a flowchart of the method for processing domain name system alarms provided by the embodiment of the present application;

[0054] Figure 2 It is a flowchart of generating an alarm recognition model provided by the embodiment of the present application;

[0055] Figure 3 The update process of the schematic alarm recognition model provided by the embodiments of this application;

[0056] Figure 4 The schematic diagram of the case library work process provided by the embodiments of this application. Specific implementation manners

[0057] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all the implementation manners consistent with this application. They are only examples of systems and methods consistent with some aspects of this application detailed in the claims.

[0058] The operation system of the domain name service consists of recursive DNS nodes, authoritative DNS nodes, a distributed database, and a server cluster to form a multi-level resolution network, supporting a large number of domain name resolution requests. At the recursive DNS level, the nodes deployed globally need to process user query requests in real time and synchronize data with the authoritative DNS nodes. Abnormal node load or response delay will trigger capacity and performance alarms; the authoritative DNS nodes maintain the domain name record database, and events such as record tampering, zone transfer anomalies, or DNSSEC verification failures will generate security alarms; the distributed database cluster bears the domain name registration information, and its sharding mechanism and master-slave synchronization process may cause database-level alarms such as lock conflicts and replica delays; at the server cluster level, events such as hardware failures, software anomalies, and network interruptions will form a composite alarm chain.

[0059] In this scenario, a single business failure may trigger multi-level cascade alarms. For example, recursive node overload may simultaneously trigger CPU load alarms, parsing timeout alarms, and upstream node unreachable alarms. The alarm information contains a large number of professional terms and has cross-layer reference logic. For example, the failure of the authoritative node synchronization indirectly causes the recursive node parsing error, requiring the handling personnel to deeply understand the system topology and protocol principles.

[0060] The alarm classification system based on the rule engine classifies alarms through keyword matching and threshold judgment and pushes them to the manual processing interface. The operation and maintenance personnel need to parse the unstructured alarm text, extract entities such as the abnormal server IP and the affected domain name, analyze the root cause in combination with historical data and the topology map, and then write a handling script and execute it through a double-check process.

[0061] Manual analysis of alarm texts takes a long time. For the correlation analysis of complex alarm chains, data needs to be retrieved across systems, significantly extending the single-event handling cycle. The rule engine has a high error rate in identifying nested descriptions and professional terms, resulting in misclassification of alarms and duplicate dispatching. Manually written scripts carry risks such as variable binding errors and lack of permission control, and some operation and maintenance accidents are caused by abnormal script execution.

[0062] To solve the above problems, some enterprises introduce natural language processing technology to build an alarm parsing system. Through syntax analysis and knowledge graphs, entity extraction and fault matching are achieved, shortening the parsing time for a single alarm. However, the natural language processing model has insufficient recognition accuracy for some professional terms and cannot understand the context correlation of cross-alarm events, which also leads to low alarm handling efficiency.

[0063] To solve the problem of low alarm handling efficiency, some embodiments of this application provide a method for processing Domain Name System (DNS) alarms. The method can identify alarm messages and automatically generate corresponding handling scripts to solve the problem of long manual analysis and handling processes, which leads to low alarm handling efficiency.

[0064] As Figure 1 shown, the method includes the following steps:

[0065] S100: Receive the alarm message of the Domain Name System.

[0066] In the Domain Name System (DNS), alarm information is generated from abnormal states or potential risks monitored during system operation. For example, system resource overrun, service function abnormality, security events, configuration errors, data consistency abnormalities, etc. The alarm message contains metadata such as the business domain to which the domain name belongs, the authoritative DNS server identifier, and the recursive resolution path.

[0067] The alarm message of the Domain Name System can be collected in real time by a monitoring agent deployed in a server cluster, or received through an alarm receiving server as a receiving module and transmitted through a standardized interface. Domain Name System alarms involve specific scenarios such as domain name resolution failure, DNS query timeout, domain name hijacking detection, DNS cache pollution, etc. The alarm text contains fields such as domain names, resolution records, and DNS server IPs.

[0068] S200: Perform target information extraction on the alarm message.

[0069] After the alarm message is received, the information extraction unit extracts the alarm message. The information extraction unit extracts target information such as alarm level, alarm type, system information, host information, time information, alarm order number, and alarm text.

[0070] In some embodiments, the information extraction unit may be an information extraction model, which is a natural language processing model based on deep learning and is used to extract structured fields from unstructured alarm texts.

[0071] For example, the information extraction model is an architecture that combines a bidirectional long short-term memory network and a conditional random field, and is supervised and trained through the labeled historical data of alarm messages. The training data can be processed alarm cases to ensure that the model is adapted to the business scenario.

[0072] In some embodiments, first, word segmentation and entity annotation are performed on the alarm information, then context semantic parsing is performed, and finally structured data is output.

[0073] In an example, using a domain dictionary, for example, a predefined DNS glossary, word segmentation is performed on the alarm text, entity types are labeled, and the association relationships in the text are identified through an attention mechanism. For example, the alarm text is "The A record resolution delay of the authoritative DNS server 192.168.1.1 for the domain name example.com exceeds 200 ms". The information extraction model associates "192.168.1.1" with the "authoritative DNS server" and extracts "A record" and "200 ms" as key parameters. The extracted entities and target information are integrated into a JSON object and stored in the database.

[0074] For the target information, the alarm levels are divided into three categories: critical, major, and minor; the alarm types include three categories: capacity alarm, performance alarm, and function alarm, and each category contains several sub-alarm types; the system information and host information may include the IP address, device model, operating system version, etc. of the relevant devices in the domain name system; the alarm text is the main body of the alarm message.

[0075] Compared with other text matching and extraction methods, the information extraction unit can process more complex alarm messages. For complex structures such as alarm nesting, reference, and ellipsis, the extracted information is more accurate, and at the same time, the understanding of text semantics and text context information is more sufficient, and the effect is better when extracting the content of alarm texts.

[0076] S300: Input the target information into the alarm recognition model to enable the alarm recognition model to identify the case scenario of the alarm message.

[0077] The case scenario can be matched through the alarm recognition model, replacing manual judgment. In some embodiments, the method further includes constructing an alarm recognition model, as Figure 2 shown, the specific steps S210 - S230.

[0078] S210: Perform text feature extraction on the alarm text to generate word vectors through a constructed model.

[0079] The above-mentioned text feature extraction refers to extracting the mathematical representation characterizing the semantic and structural features from the alarm text. Specifically, it is achieved by constructing a model to generate word vectors.

[0080] In some embodiments, the construction model is one of a bigram bag-of-words model, an n-gram model, and a recurrent neural network.

[0081] The bigram bag-of-words model divides the alarm text into a word sequence and then counts the co-occurrence frequency of adjacent two words to generate a sparse vector. For example, the alarm text "DNS server response timeout" is tokenized into ["DNS", "server", "response", "timeout"], and the bigram bag-of-words model converts it into a vector representation of ["DNS-server", "server-response", "response-timeout"].

[0082] The alarm text contains fixed phrases, such as "parsing failed" and "cache pollution". The bigram model can effectively capture such local context relationships and avoid the defect of ignoring word order in a single bag-of-words model. Although the feature dimension of the bigram bag-of-words model is higher than that of a single bag-of-words model, it is lower than that of a higher-order n-gram model. In the scenario of high concurrency of alarm messages, the storage and calculation costs of the sparse vector are controllable. Without a complex model architecture or large-scale training data, features can be directly generated through statistical methods, which is convenient for quickly integrating into the existing alarm processing system.

[0083] The n-gram model extends the bigram bag-of-words model to combinations of N consecutive words. For example, the trigram model converts "DNS server response timeout" into ["DNS-server-response", "server-response-timeout"] to capture longer-distance context dependencies and improve the representation ability for complex alarm patterns, such as multi-factor associated anomalies.

[0084] The recurrent neural network dynamically extracts the temporal features of the text through sequence modeling to generate a dense vector. For example, the recurrent neural network can identify the causal relationship in "DNS server recursive resolution timeout due to high load". The recurrent neural network supports variable-length text input, captures global semantics and long-distance dependencies, and is suitable for scenarios where there are implicit logics (such as conditional clauses) in the alarm text.

[0085] When computing resources are limited, the bigram bag-of-words model achieves key feature extraction at a low cost, balancing efficiency and effectiveness. When computing resources are sufficient, the n-gram model or the recurrent neural network can be used to improve the classification accuracy by increasing the context window or performing deep semantic analysis.

[0086] S220: Concatenate the word vector with the alarm level, alarm type, and alarm time to construct a multi-dimensional joint feature vector.

[0087] The multi-dimensional combined feature vector is composed of a word vector concatenated with the alarm level, alarm type, and alarm time. Exemplarily, the alarm text is converted into a vector with a fixed dimension by constructing a model. Among the alarm levels, the three categories of emergency, major, and minor are encoded as one-hot vectors. Among the alarm types, types such as capacity alarm, performance alarm, and function alarm are numerically mapped. Among the alarm times, periodic features such as the hour and day of the week of the timestamp are extracted and normalized to the interval [0, 1].

[0088] Then, the word vector and the encoded metadata are concatenated by dimension. For example, a 256-dimensional word vector and 5-dimensional metadata, that is, the alarm level, alarm type, and alarm time are concatenated into a 261-dimensional combined vector.

[0089] Through the multi-dimensional features of the text and metadata, the limitations of a single data source are compensated. For example, the text "DNS response latency" combined with the "performance alarm" type can match the "recursive resolution timeout" scenario.

[0090] S230: Use the random forest algorithm to perform training on the multi-dimensional combined feature vector to generate an alarm recognition model.

[0091] The random forest algorithm consists of multiple decision trees and is trained through the following steps. The historical alarm data is divided into a training set and a validation set in chronological order. Each decision tree is independently trained based on a randomly sampled subset of the data set and a subset of features, and the splitting node is selected through information gain or the Gini coefficient. The classification results of all decision trees are determined by a voting mechanism to obtain the final output.

[0092] The alarm data has high-dimensional and sparse features. The random forest reduces the risk of overfitting through multi-tree integration and supports parallel training at the same time.

[0093] After the multi-dimensional combined feature vector is input into the trained random forest model, each decision tree independently traverses the feature space and determines the belonging path of the feature vector layer by layer according to the node splitting rules, such as "whether the alarm level is emergency" and "whether the text contains 'parsing failed'".

[0094] Each tree outputs a prediction result, such as "parsing failed scenario" or "cache pollution scenario". The prediction results of all decision trees determine the final scenario number through majority voting. For example, among 10 trees, 7 predict "parsing failed", and the alarm recognition model outputs this scenario and triggers the corresponding handling process.

[0095] In some embodiments, it is possible to determine whether a case scenario is recognized through the voting result and a preset threshold. For example, if the voting result does not reach the preset threshold, the alarm recognition model is marked as not recognized.

[0096] S400: In the case of recognizing a case scenario, generate an alarm handling script based on the target information and the corresponding handling command template in the case library.

[0097] In the case where a case scenario is recognized, that is, when the voting result reaches a preset threshold, data is extracted from the case library. The case library includes the basic information of the case, the case scenario, and the disposal command template. In an exemplary scenario, according to the number of the case scenario in the basic information of the case, the corresponding disposal command template is queried from the case library.

[0098] The basic information of the case is used to uniquely identify and describe the attributes of the alarm case. Exemplarily, the basic information includes the fields in the following table:

[0099]

[0100] The case number is a globally unique identifier for retrieving the case; the case name is a descriptive name for easy identification and classification; the alarm level is divided into three categories: emergency, major, and minor, which defines the severity of the alarm; the alarm type is divided according to the business scenario, and the function alarm can be further subdivided; the alarm text content is the original alarm message text that triggers the case; the alarm time is the timestamp when the alarm is first triggered, which is used to associate periodic events.

[0101] In some embodiments, in the case where a case scenario is recognized, the case type of the case scenario is obtained, a disposal command template is matched based on the case type, and an alarm disposal script is generated based on the disposal command template.

[0102] The case types include three major categories: capacity alarm, performance alarm, and function alarm. A capacity alarm means that the system resources reach or exceed a preset threshold. For example, when the disk usage rate of the DNS cache server exceeds 90%, a capacity alarm is triggered. A performance alarm means that the service metrics are lower than the expected level. For example, when the average response time of the recursive DNS server exceeds 500 ms, a performance alarm is triggered. A function alarm means that the core business function is abnormal or a security event occurs. For example, when the authoritative DNS server fails to return the A record resolution result, or an illegal Zone Transfer request is detected, a function alarm is triggered.

[0103] The case types are defined based on historical alarm data analysis and business rules, and are pre-configured in the case library by operation and maintenance personnel or automated tools. Among them, capacity alarms, performance alarms, and function alarms are the main case types, and each main case type can also correspond to multiple sub-case types. For example, the function alarm corresponds to a sub-case type of resolution failure and a sub-case type of hijacking detection.

[0104] After recognizing the case scenario, retrieve the corresponding disposal command template from the case library according to the case type. Specifically, obtain the main case type and sub-case type from the metadata of the case scenario. If there are multiple disposal command templates for the same type, select the disposal command template according to the preset priority (such as disposal success rate, execution time). According to the matching template, call the command group stored in the case library, and then generate an alarm disposal script.

[0105] The case scenario defines the specific problem type and context conditions corresponding to the alarm. Exemplarily, the case scenario includes the fields in the following table:

[0106] Serial number Field name Remarks 1 Case scenario number Identify unique command 2 Case scenario name 3 Alarm handling command group Ordered list of commands 4 Automatic execution flag Whether the command group is automatically executed

[0107] The case scenario number is the unique identifier, which is associated with the case number in the basic case information; the case scenario name describes the specific scenario; the alarm disposal command group is an ordered list of command sequence numbers, pointing to the specific operation steps in the disposal command template library; the automatic execution flag is a boolean value, which defines whether the disposal script generated by this scenario is allowed to be automatically executed.

[0108] In the domain name system, DNS alarm disposal involves key operations. For example, refreshing the root domain name cache, blocking hijacked IPs, etc. Incorrect execution may lead to parsing failures or security vulnerabilities. Therefore, the review mechanism needs to strengthen permission control.

[0109] In some embodiments, read the automatic execution flag of the case scenario to generate review information according to the automatic execution flag. In the case where the review information requires manual review, generate a review request, and the review request represents the execution of a manual review of the alarm disposal script; obtain the review information, and the review information is the information generated after performing a manual review based on the review request; based on the review information, generate an alarm disposal script.

[0110] In the case scenario of the case library, each scenario is associated with an automatic execution flag. For example, high-risk operations are marked as "false" and require manual review; low-risk operations are marked as "true" and allow automatic execution.

[0111] When the alarm recognition model matches the case scenario, read the automatic execution flag of the case scenario. If the automatic execution flag is false, generate review information, and the review information includes alarm details, the pending disposal script, and the review basis.

[0112] The alarm details include key information such as the alarm level, type, associated domain name, server IP, etc. The pending disposal script is the script to be reviewed generated based on the case scenario command template. The review basis is the risk description that triggers the review. The review information is pushed to the manual review interface through the message queue or API interface and recorded in the list of tasks to be processed.

[0113] After manual review is completed, update the status of the review information and generate an alarm handling script.

[0114] The disposal command template is a predefined set of operation instructions, including specific execution logic and dynamic variables. Exemplarily, the disposal command template includes the fields in the following table:

[0115] Serial number Field name Remarks 1 Command number Identify unique command 2 Command name 3 Command template Named content 4 Pre-action Operations to be performed before the command is executed 5 Post-action Operations to be performed after the command is executed

[0116] The command number is a unique identifier associated with the command group in the case scenario; the command name is a descriptive name; the command template is a text template of the executable command, containing placeholder variables; the pre-action is the verification or preparation operation before the command execution; the post-action is the subsequent operation after the command execution.

[0117] Exemplarily, in the case of identifying a case scenario, when the alarm message "The A record resolution of example.com by the DNS server fails" is identified as the "resolution failure" scenario, retrieve the command group associated with the resolution failure scenario from the case library. For example, the command group is CMD-101 "Check the status of the authoritative server" and CMD-102 "Refresh the local cache", and substitute the variables extracted from the information (domain name "example.com", server IP "192.168.1.1") into the disposal command template to generate an executable alarm handling script.

[0118] In the case where no case scenario is identified, in some embodiments, mark the alarm message as an unrecognized status and generate a new case.

[0119] Store the new case in the case library, and record the number of new cases and the training cycle of the alarm recognition model.

[0120] If the number is greater than or equal to the number threshold, and / or, if the training cycle is greater than or equal to the preset cycle threshold, trigger the alarm recognition model to perform iterative training.

[0121] Extract the alarm text, alarm level, alarm type, and alarm time of the new case.

[0122] Based on the alarm text, alarm level, alarm type, and alarm time of the new case, update the parameters of the alarm recognition model to generate an updated alarm recognition model.

[0123] It can be understood that for a new case, a temporary case number, alarm details, manual handling records, etc. can be recorded. The temporary case number is a unique identifier for subsequent tracking. The alarm details include alarm text, level, type, timestamp, and extracted entities. The manual handling record is the disposal command written manually by the operation and maintenance personnel and its execution result.

[0124] For new cases, to improve the efficiency of alarm handling, model update training can be triggered. When the number of new cases reaches a threshold or a preset cycle threshold is reached, a new round of training is performed on the model. For example, if the preset cumulative number of new cases is 50, model training is triggered; if the preset time interval is 7 days and the number threshold is not reached after the timeout, training is still forced to be triggered.

[0125] As Figure 3 shown, during the new round of model training, the alarm text, level, type, and time of new cases are extracted from the case library, merged with historical data to form a new training set, and then feature extraction is performed on the alarm text to generate word vectors and splice them with metadata into multi-dimensional joint feature vectors.

[0126] Based on the original decision tree, subtrees are generated or node splitting rules are adjusted for new data, rather than a full-scale reconstruction. According to the high-frequency features in the new cases, the selection probability of the feature subset in the random forest is adjusted.

[0127] In the domain name system, DNS unrecognized alarms may involve new types of attacks or protocol evolution, and need to be quickly incorporated into the case library to address threats. For other systems, for example, unrecognized alarms in web servers mostly originate from code logic errors or configuration changes, and do not require high-frequency model updates.

[0128] After adding new cases, a new round of training can be performed on the model through the case scenario. In some embodiments, the preset scenario of the new case is obtained;

[0129] According to the preset scenario, the new cases are classified and stored in the training set of the case library;

[0130] If the preset scenario corresponding to the new case is the first scenario set, corpus data is obtained;

[0131] The corpus data is supplemented to the target training set to generate a supplemented training set, and the target training set is the training set corresponding to the first scenario set;

[0132] Based on the supplemented training set, incremental training is performed on the alarm recognition model to generate an alarm recognition model.

[0133] The preset scenario is a classification label for predefined high-risk or special business scenarios. The preset scenario at least includes a first scenario set and a second scenario set. The first scenario set includes scenarios with high security risks and wide business impacts, such as DNS hijacking, cache poisoning, and root server unreachability. Alarms in the first scenario set need to be processed first, and the model is required to have high recognition accuracy. The second scenario set is for routine operation and maintenance problems, such as performance alarms and full log storage, and its handling logic is relatively fixed and the risk is low.

[0134] Match the preset scenario classification rules according to the alarm type, level, and content keywords of the newly added cases. For example, when the alarm text contains "DNSSEC verification failed", it is automatically classified into the first scenario set. If the confidence level of automatic classification is insufficient, manual review can be performed, and the operation and maintenance personnel can specify the scenario set attribution.

[0135] For the cases in the first scenario set, they can be stored in a dedicated training set for high-priority incremental training of the model. For the cases in the second scenario set, they can be stored in a regular training set for batch training at regular intervals.

[0136] The corpus data is extended training data related to the first scenario set, such as historical alarm data, other feature data, etc. The historical alarm data is historical cases of the same type, and the other feature data is associated attack feature data obtained from external sources.

[0137] Merge the newly added cases with the supplemented corpus data into the target training set. For example, the newly added "DNS hijacking" cases are merged with historical hijacking data and external malicious IP intelligence, and then the process of model training is executed. For model training, refer to the above content and will not be elaborated here.

[0138] As Figure 4 shown, in some embodiments, based on the newly added cases, generate a template addition instruction, then generate a disposal command template, match the disposal command template based on the case type, and based on the disposal command template, generate an alarm disposal script.

[0139] In the situation where the case scenario is not recognized, after manually processing the alarm message, extract the processed disposal command sequence, and convert the disposal command sequence into a standardized command template and store it in the case library.

[0140] In some other embodiments, replace the dynamic variables in the executed disposal command template, generate a template addition instruction, then generate a disposal command template, match the disposal command template based on the case type, and based on the disposal command template, generate an alarm disposal script.

[0141] The template addition instruction refers to a standardized command generation instruction generated based on the newly added cases or dynamic variable replacement.

[0142] Among them, the disposal command template includes dynamic variables, and the dynamic variables are preset placeholders in the disposal command template, which are used to be replaced with actual values during script generation. The dynamic variables include host names, system identifiers, and time parameters.

[0143] The hostname variable represents the unique identifier of the target server, which is derived from the host information extracted from the alarm message; the system identifier variable identifies the associated DNS subsystem and is derived from the system information field in the alarm metadata; the time parameter variable represents the time point or time period when the operation is executed and is derived from the alarm trigger timestamp or the manually specified execution window.

[0144] In some embodiments, the corresponding dynamic variables are replaced with host information and system information to generate a command sequence; a pre-check instruction and a post-check instruction are embedded in the command sequence to generate a template addition instruction.

[0145] Information such as the hostname and system identifier is extracted from the alarm text, and then the extracted entities are matched with the placeholders in the template, and the replaced command sequence generates specific executable instructions.

[0146] The pre-check instruction is used to verify the execution environment parameters, that is, a pre-check is added to the command sequence, such as permission verification, and the post-check instruction is used to verify the command execution result and trigger the environment recovery operation, that is, a post-check is added to the command sequence, such as result logging.

[0147] The variables in the DNS handling template need to include the domain name and record type, while other system templates can be based on the port number and interface name.

[0148] The alarm handling command group can be matched by the number of the case scenario, and then an alarm handling script is generated. In some embodiments, the number of the case scenario is obtained.

[0149] According to the number, the corresponding alarm handling command group in the case library is queried. The alarm handling command group contains multiple handling command numbers arranged in the execution order.

[0150] According to the execution order, logical splicing is performed on the handling command template generated based on the template addition instruction to generate an alarm handling script.

[0151] The case scenario number is a string or digital sequence that uniquely identifies the case scenario. The case scenario number can be composed of a prefix plus an incrementing serial number. For example, "DNS-101" represents the "recursive resolution timeout" scenario. The case scenario number corresponds to the scenario record in the case library. When a new scenario is added to the case library, a number is assigned. For example, when manually processing an unrecognized alarm and defining a new scenario "DNSSEC verification failed", the number "DNS-105" is generated and associated with this scenario.

[0152] Each number corresponds to a specific disposal command template in the case library. According to the number, retrieve the corresponding command group list from the scenario table in the case library, and parse the disposal command numbers in the command group. Query the template content item by item from the disposal command template table in the case library according to the command number list, then replace the placeholders in the template with the extracted dynamic variables, and splice the commands in the order of the command group to generate an alarm disposal script.

[0153] Through the domain name system alarm handling method provided in this embodiment, the recognition and automatic disposal of alarm messages can be realized, the stability and reliability of the domain name system can be improved, the alarm recognition model can accurately recognize alarm scenarios, reduce the need for manual intervention, and reduce the response time for handling alarms. The automatically generated alarm disposal script can provide a script according to specific problems, avoid the risk of misoperation that may be brought by a general script, and thus improve the problem of low efficiency in solving alarm disposal.

[0154] Based on the above-mentioned method for handling domain name system alarms, some embodiments of the present application also provide a system for handling domain name system alarms, including:

[0155] A disposal unit, configured to receive the alarm message of the domain name system;

[0156] An information extraction unit, configured to perform target information extraction on the alarm message, where the target information includes an alarm level, an alarm type, system information, host information, and alarm text;

[0157] A recognition unit, configured to input the target information into an alarm recognition model to enable the alarm recognition model to recognize the case scenario of the alarm message;

[0158] The disposal unit is further configured to, in the case of recognizing the case scenario, generate an alarm disposal script based on the target information and the corresponding disposal command template in the case library, where the case library includes the basic information of the case, the case scenario, and the disposal command template.

[0159] For the effect of the above system embodiment during operation, reference can be made to the effect of the above method embodiment, which will not be elaborated here.

[0160] As can be seen from the above technical solutions, the present application provides a method and system for processing domain name system alarms. The method includes: receiving an alarm message of the domain name system; performing target information extraction on the alarm message, where the target information includes an alarm level, an alarm type, system information, host information, and an alarm text; inputting the target information into an alarm recognition model to enable the alarm recognition model to recognize a case scenario of the alarm message; and in the case of recognizing the case scenario, generating an alarm handling script based on the target information and a corresponding handling command template in a case library, where the case library includes basic information of the case, the case scenario, and the handling command template. The method can recognize alarm messages and automatically generate corresponding handling scripts, so as to solve the problem of low alarm handling efficiency caused by the long manual analysis and handling process.

[0161] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other embodiments extended based on the solution of the present application without creative efforts belong to the protection scope of the present application.

Claims

1. A method for processing domain name system alarms, characterized in that, Including: Receiving an alarm message of the Domain Name System; Performing target information extraction on the alarm message, where the target information includes an alarm level, an alarm type, system information, host information, and an alarm text; Inputting the target information into an alarm recognition model to enable the alarm recognition model to recognize a case scenario of the alarm message; When the case scenario is recognized, generating an alarm handling script based on the target information and a corresponding handling command template in a case library, where the case library includes basic information of the case, the case scenario, and the handling command template.

2. The method for processing domain name system alarms according to claim 1, wherein The method further includes: When the case scenario is not recognized, marking the alarm message as an unrecognized state and generating a new case; Storing the new case in the case library and recording the quantity of the new cases and the training period of the alarm recognition model; If the quantity is greater than or equal to a quantity threshold, and / or if the training period is greater than or equal to a preset period threshold, triggering the alarm recognition model to perform iterative training; Extracting the alarm text, alarm level, alarm type, and alarm time of the new case; Updating parameters of the alarm recognition model based on the alarm text, alarm level, alarm type, and alarm time of the new case to generate an updated alarm recognition model.

3. The method for processing domain name system alarms according to claim 2, wherein The generating of the alarm handling script includes: When the case scenario is not recognized, generating a handling command template based on a template addition instruction, where the template addition instruction is an instruction generated based on the new case or by replacing dynamic variables in the executed handling command template; Matching a handling command template based on the type of the case, and generating an alarm handling script based on the handling command template.

4. The method for processing domain name system alarms according to claim 3, wherein The handling command template includes dynamic variables, and the dynamic variables include a host name, a system identifier, and a time parameter; The generating of the handling command template based on the template addition instruction includes: Replacing corresponding dynamic variables with the host information and system information to generate a command sequence; Embedding a pre-check instruction and a post-check instruction in the command sequence to generate a template addition instruction, where the pre-check instruction is used to verify execution environment parameters, and the post-check instruction is used to verify a command execution result and trigger an environment recovery operation.

5. The method for processing domain name system alarms according to claim 4, wherein The generating of the alarm handling script includes: Obtaining a number of the case scenario; Querying a corresponding alarm handling command group in the case library according to the number, where the alarm handling command group includes a plurality of handling command numbers arranged in an execution order; Performing logical splicing on the handling command template generated based on the template addition instruction in accordance with the execution order to generate an alarm handling script.

6. The method for processing domain name system alarms according to claim 1, wherein The generating of the alarm handling script includes: When the case scenario is recognized, obtaining a case type of the case scenario; Matching a handling command template based on the case type, and generating an alarm handling script based on the handling command template.

7. The method for processing domain name system alarms according to claim 2, wherein After generating the new case, it further includes: Obtaining a preset scenario of the new case; According to the preset scenario, classify and store the new case in the training set of the case library, where the preset scenario includes at least a first scenario set and a second scenario set; If the preset scenario corresponding to the new case is the first scenario set, obtain corpus data; Supplement the corpus data to the target training set to generate a supplemented training set, where the target training set is the training set corresponding to the first scenario set; Based on the supplemented training set, perform incremental training on the alarm recognition model to generate an alarm recognition model.

8. The method for processing domain name system alarms according to claim 1, wherein The generation of the alarm handling script includes: Read the automatic execution flag of the case scenario to generate review information according to the automatic execution flag; In the case where the review information requires manual review, generate a review request, where the review request represents a manual review of the alarm handling script; Obtain review information, where the review information is the information generated after performing a manual review based on the review request; Generate an alarm handling script based on the review information.

9. The method for processing domain name system alarms according to claim 1, wherein The method further includes: Perform text feature extraction on the alarm text to generate word vectors through a construction model, where the construction model is one of a binary bag-of-words model, an n-gram model, and a recurrent neural network; Perform feature splicing on the word vectors with the alarm level, alarm type, and alarm time to construct a multi-dimensional joint feature vector; Use the random forest algorithm to perform training on the multi-dimensional joint feature vector to generate the alarm recognition model.

10. A processing system for domain name system alarms, characterized in that, It includes: A handling unit for receiving the alarm message of the domain name system; An information extraction unit for extracting target information from the alarm message, where the target information includes an alarm level, an alarm type, system information, host information, and an alarm text; An identification unit for inputting the target information into the alarm recognition model to enable the alarm recognition model to identify the case scenario of the alarm message; The handling unit is further configured to, in the case where the case scenario is recognized, generate an alarm handling script based on the target information and the corresponding handling command template in the case library, where the case library includes the basic information of the case, the case scenario, and the handling command template.