Domain Name Authorization Risk Assessment Methods, Devices, Equipment and Products

The domain name authorization risk assessment method, which integrates pre-trained tree models and neural network models, and combines hierarchical analysis and large language models, solves the problem of inaccurate assessment in existing technologies, and achieves more efficient and accurate domain name authorization risk assessment.

CN121542676BActive Publication Date: 2026-05-26JINAN UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JINAN UNIVERSITY
Filing Date
2026-01-16
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing domain name authorization risk assessment methods suffer from incomplete indicators, insufficient linear weighted calculation, difficulty in balancing interpretability and data-driven capabilities, and high cost and low efficiency of manual annotation, leading to inaccurate assessments.

Method used

By fusing pre-trained tree models and pre-trained neural network models, risk assessment is conducted using multi-dimensional feature data. The model is then labeled using hierarchical analysis (AHP) models and large language models to construct a domain name authorization risk prediction model for comprehensive risk assessment.

Benefits of technology

It improves the accuracy and reliability of domain name authorization risk assessment, solves the problem of incomplete indicators in existing methods, reduces the cost of manual annotation, and improves assessment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542676B_ABST
    Figure CN121542676B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and product for domain name authorization risk assessment, relating to the field of network security technology. The method includes: receiving multi-dimensional feature data of a domain name to be assessed; and assessing the authorization risk of the multi-dimensional feature data using a pre-built domain name authorization risk prediction model to obtain an assessment result. The domain name authorization risk prediction model is trained using a pre-trained tree model and a pre-trained neural network model. This solves the problem of incomplete indicators and inefficient annotation in existing methods, leading to inaccurate domain name authorization risk assessment, and improves the accuracy of domain name authorization risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to a method, apparatus, device, and product for assessing domain name authorization risks. Background Technology

[0002] The Domain Name System (DNS) is a critical component of the Internet, responsible for resolving domain names into IP addresses to ensure that users can access network services through easy-to-remember domain names. The DNS authorization mechanism is the core of the system, which delegates the responsibility of resolving subdomains to other DNS servers through a hierarchical and delegation mechanism. The resolution process involves a complex authorization dependency structure involving multiple levels, cross-domains, and cross-borders. Improper configuration or security vulnerabilities at any node in the authorization path can easily lead to attacks such as hijacking and tampering of resolution, which can damage the correctness and integrity of DNS and related applications and trigger a chain of security risks. Current domain name authorization dependency security risk assessment mainly adopts rule scoring or linear weighting methods.

[0003] However, existing assessment methods have significant shortcomings. First, the indicator system is incomplete and structurally ambiguous, focusing on a single dimension and lacking a systematic indicator system based on RFC standards, historical security events, and analysis path structures, making it difficult to comprehensively reflect multidimensional risks. Second, the linear weighted calculation method cannot accurately characterize the impact of fatal anomalies and the amplified risks of extreme non-fatal indicators, masking the anomaly coupling effect and leading to assessment bias. Third, it is difficult to balance interpretability, prior knowledge, and data-driven capabilities. Rules rely on human experience, and data-driven approaches are easily affected by noise, neither of which can adapt to complex dynamic analysis environments. Fourth, manual annotation is costly, inefficient, and inconsistent, affecting the quality of model training.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a method, apparatus, device, and product for assessing domain name authorization risk, aiming to solve the technical problem that existing methods suffer from incomplete indicators and inefficient labeling, leading to inaccurate domain name authorization risk assessment.

[0006] To achieve the above objectives, this application proposes a domain name authorization risk assessment method, which includes:

[0007] Receive multidimensional feature data of the domain name to be evaluated;

[0008] The multidimensional feature data is assessed for authorization risk using a pre-built domain name authorization risk prediction model to obtain the assessment results. The domain name authorization risk prediction model is trained using a pre-trained tree model and a pre-trained neural network model.

[0009] In one embodiment, before the step of assessing the authorization risk of the multidimensional feature data using a pre-built domain name authorization risk prediction model to obtain the assessment result, the method further includes:

[0010] Collect Domain Name System (DNS) metrics;

[0011] The DNS metrics are filtered to obtain the hierarchical analysis (AHP) model;

[0012] The weights of the analytic hierarchy process (AHP) model are calculated to obtain the prior risk value.

[0013] Domain name samples are obtained by hierarchical sampling based on the prior risk values, and the domain name samples are annotated by DeFiel using a large language model to obtain an annotated dataset;

[0014] The pre-trained tree model and the pre-trained neural network model are trained and validated using the labeled dataset until their performance converges. Then, the pre-trained tree model and the pre-trained neural network model are fused to obtain the domain name authorization risk prediction model.

[0015] In one embodiment, the step of filtering the DNS indicators to obtain the hierarchical analysis (AHP) model includes:

[0016] A stability analysis was performed on the DNS metrics, and the results were obtained.

[0017] Based on the stability analysis results, low-value DNS indicators are removed to obtain high-value DNS indicators.

[0018] Based on the security impact scope, risk triggering intensity, and historical security events of the high-value DNS indicators, the high-value DNS indicators are classified and merged to obtain the AHP model.

[0019] In one embodiment, the step of calculating the weights of the analytic hierarchy process (AHP) model to obtain the prior risk value includes:

[0020] The AHP judgment matrix is ​​obtained by traversing and comparing the importance of high-value DNS indicators in the AHP model through a scaling system.

[0021] The consistency of the AHP judgment matrix is ​​checked to obtain the check result;

[0022] If the verification result is passed, the index weights of several high-value DNS indicators in the AHP model are calculated using the feature vector method.

[0023] The basic weighted average and basic score of the several high-value DNS indicators are calculated based on the preset standardized mapping rules and the indicator weights.

[0024] The basic weighted average and basic score are enhanced with risk through hard failure rules and soft extreme nonlinear amplification mechanism, and the enhanced basic weighted average and basic score are mapped to the risk scoring interval to obtain the prior risk value.

[0025] In one embodiment, the step of obtaining domain name samples through stratified sampling based on the prior risk value, and then annotating the domain name samples using a large language model to obtain an annotated dataset includes:

[0026] The stratified sampling criteria are determined based on the prior risk value;

[0027] The unlabeled domain names are sampled evenly using the hierarchical sampling annotation to obtain domain name samples;

[0028] After converting the prior risk value into a structured risk warning, based on the structured risk warning, the domain name samples are labeled as high-risk samples using a large language model to obtain high-quality supervised samples.

[0029] The high-quality supervised samples are divided into a training set and a validation set, and the labeled dataset is obtained through the training set and the validation set.

[0030] In one embodiment, the step of training and validating the pre-trained tree model and the pre-trained neural network model using the labeled dataset until the model performance of the pre-trained tree model and the pre-trained neural network model converges, and then fusing the pre-trained tree model and the pre-trained neural network model to obtain the domain name authorization risk prediction model includes:

[0031] The training set of the labeled dataset is input into the pre-trained tree model and the pre-trained neural network model for supervised training to obtain the initial tree model and the initial neural network model.

[0032] Obtain an unlabeled sample set, and use the initial tree model and initial neural network model to predict and estimate the uncertainty of the unlabeled sample set to obtain a consistency measure.

[0033] Based on the consistency metric, the unlabeled sample set is divided into high-confidence samples, medium-confidence samples, and low-confidence samples;

[0034] The training set is corrected by using the high-confidence samples, medium-confidence samples, and low-confidence samples to obtain the corrected training set;

[0035] The initial tree model and the initial neural network model are iteratively trained using the modified training set until their performance converges. Then, the initial tree model and the initial neural network model are fused to obtain the domain name authorization risk prediction model.

[0036] In one embodiment, the step of fusing the initial tree model and the initial neural network model to obtain the domain name authorization risk prediction model includes:

[0037] The model performance metrics of the initial tree model and the initial neural network model are calculated using the validation set of the labeled dataset.

[0038] The fusion weights of the initial tree model and the initial neural network model are determined based on the model performance metrics.

[0039] Based on the fusion weights, the initial tree model and the initial neural network model are fused using a weighted summation formula to obtain the domain name authorization risk prediction model.

[0040] Furthermore, to achieve the above objectives, this application also proposes a domain name authorization risk assessment device, which includes:

[0041] The receiving module is used to receive multidimensional feature data of the domain name to be evaluated;

[0042] The evaluation module is used to evaluate the authorization risk of the multidimensional feature data through a pre-built domain name authorization risk prediction model and obtain the evaluation result. The domain name authorization risk prediction model is trained by a pre-trained tree model and a pre-trained neural network model.

[0043] In addition, to achieve the above objectives, this application also proposes a domain name authorization risk assessment device, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the domain name authorization risk assessment method as described above.

[0044] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the domain name authorization risk assessment method described above.

[0045] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the domain name authorization risk assessment method described above.

[0046] One or more technical solutions proposed in this application have at least the following technical effects:

[0047] This application proposes a domain name authorization risk assessment method, apparatus, device, and product. It receives multi-dimensional feature data of the domain name to be assessed; and performs authorization risk assessment on the multi-dimensional feature data using a pre-built domain name authorization risk prediction model to obtain an assessment result. The domain name authorization risk prediction model is trained using a pre-trained tree model and a pre-trained neural network model. Therefore, by inputting multi-dimensional feature data of the domain name to be assessed, it overcomes the limitations of existing single indicators. Furthermore, the model training employs a pre-trained tree model and a neural network model for training and fusion, solving the problem of inefficient annotation and improving the quality of data assessment. Finally, the trained domain name authorization risk prediction model is used for authorization risk assessment, thereby improving the accuracy of the assessment. This solves the problem of incomplete indicators and inefficient annotation in existing methods, which leads to inaccurate domain name authorization risk assessment, and improves the accuracy of domain name authorization risk assessment. Attached Figure Description

[0048] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0049] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart illustrating the first embodiment of the domain name authorization risk assessment method for this application.

[0051] Figure 2 This is a diagram illustrating the dataset annotations involved in the domain name authorization risk assessment method used in this application.

[0052] Figure 3 This diagram illustrates the AHP analysis method used in the domain name authorization risk assessment of this application, which incorporates expert scoring.

[0053] Figure 4 This is a flowchart illustrating the second embodiment of the domain name authorization risk assessment method for this application.

[0054] Figure 5 A simplified flowchart illustrating the domain name authorization risk assessment method provided in Embodiment 2 of this application;

[0055] Figure 6 This is a schematic diagram of the module structure of the domain name authorization risk assessment device according to an embodiment of this application;

[0056] Figure 7 This is a schematic diagram of the hardware operating environment involved in the domain name authorization risk assessment method in this application embodiment.

[0057] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0058] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0059] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0060] The main solution of this application embodiment is as follows: Collect Domain Name System (DNS) metrics; perform metric filtering on the DNS metrics to obtain an Analytic Hierarchy Process (AHP) model; calculate weights on the AHP model to obtain prior risk values; perform stratified sampling based on the prior risk values ​​to obtain domain name samples; perform DeFi annotation on the domain name samples using a large language model to obtain an annotated dataset; train and validate a pre-trained tree model and a pre-trained neural network model using the annotated dataset until the model performance of the pre-trained tree model and the pre-trained neural network model converges; then fuse the pre-trained tree model and the pre-trained neural network model to obtain a domain name authorization risk prediction model. Perform stability analysis on the DNS metrics to obtain stability analysis results; based on the stability analysis results, remove low-value DNS metrics to obtain high-value DNS metrics; classify and merge the high-value DNS metrics according to their security impact range, risk triggering intensity, and historical security events to obtain an AHP model. The importance of high-value DNS indicators in the AHP model is compared through a scaling system to obtain the AHP judgment matrix. The AHP judgment matrix is ​​then subjected to consistency verification to obtain the verification result. If the verification result is satisfactory, the indicator weights of several high-value DNS indicators in the AHP model are calculated using the eigenvector method. Based on preset standardization mapping rules and the indicator weights, the basic weights and basic scores of the several high-value DNS indicators are calculated. The basic weights and basic scores are then enhanced with risk through hard failure rules and soft extreme nonlinear amplification mechanisms, and the enhanced basic weights and basic scores are mapped to risk scoring intervals to obtain prior risk values. Based on the prior risk values, a stratified sampling standard is determined. Unlabeled domain names are sampled evenly using the stratified sampling labeling to obtain domain name samples. After converting the prior risk values ​​into structured risk warnings, high-risk samples are labeled using a large language model based on the structured risk warnings to obtain high-quality supervised samples. The high-quality supervised samples are divided into training and validation sets, and a labeled dataset is obtained from the training and validation sets.The training set of the labeled dataset is input into the pre-trained tree model and the pre-trained neural network model for supervised training to obtain the initial tree model and the initial neural network model. An unlabeled sample set is obtained, and the initial tree model and the initial neural network model are used to predict and estimate the uncertainty of the unlabeled sample set to obtain a consistency metric. Based on the consistency metric, the unlabeled sample set is divided into high-confidence samples, medium-confidence samples, and low-confidence samples. The training set is then corrected using the high-confidence, medium-confidence, and low-confidence samples to obtain a corrected training set. The initial tree model and the initial neural network model are iteratively trained using the corrected training set until their performance converges. The initial tree model and the initial neural network model are then fused to obtain the domain name authorization risk prediction model. The model performance metrics of the initial tree model and the initial neural network model are calculated using the validation set of the labeled dataset. The fusion weights of the initial tree model and the initial neural network model are determined based on the model performance metrics. Based on the fusion weights, the initial tree model and the initial neural network model are fused using a weighted summation formula to obtain the domain name authorization risk prediction model. This invention solves the problems of incomplete indicators and inefficient annotation in existing methods, which lead to inaccurate domain name authorization risk assessment. It achieves accurate assessment of domain name authorization risk. Based on this invention, addressing the challenges of indicator coverage and annotation efficiency in current domain name authorization risk assessment, and considering the diverse configuration characteristics, risk dimensions, and data distributions across different domain name authorization scenarios, the lack of a systematic multi-dimensional indicator system and efficient annotation mechanism makes it impossible to guarantee the accuracy of risk assessment results, resulting in low accuracy, this invention designs a domain name authorization risk assessment method. The effectiveness of this method is verified during domain name authorization risk assessment, and the accuracy of domain name authorization risk assessment using this method is significantly improved.

[0061] In this embodiment, for ease of description, the domain name authorization risk assessment device will be used as the execution subject in the following description.

[0062] Due to the limitations of existing domain name authorization risk assessment methods, which mainly rely on rule-based scoring or linear weighting, the comprehensiveness and accuracy of risk assessment need to be improved. Firstly, the indicator system is incomplete and structurally ambiguous, lacking a systematic design that makes it difficult to reflect multidimensional risks. Secondly, linear weighting cannot characterize fatal anomalies and extreme indicators that amplify risks, masking coupling effects and leading to bias. Furthermore, it is difficult to balance interpretability, prior knowledge, and data-driven capabilities, making it unsuitable for complex and dynamic resolution environments. Finally, manual annotation is costly, inefficient, and inconsistent, affecting model training quality. Therefore, in scenarios with complex dependency structures in DNS authorization, such as multi-level, cross-domain, and cross-border processes, existing methods are difficult to adapt and are prone to triggering chain risks due to improper configuration or security vulnerabilities, failing to guarantee the reliability of the assessment.

[0063] This application provides a solution that inputs multi-dimensional feature data of the domain name to be evaluated, overcoming the limitations of existing single indicators. Furthermore, it employs a pre-trained tree model and a neural network model for training and fusion during model training, addressing the problem of inefficient annotation and improving the quality of data evaluation. Finally, the trained domain name authorization risk prediction model is used for authorization risk assessment, thereby improving the accuracy of the assessment. This solution addresses the problems of incomplete indicators and inefficient annotation in existing methods, which lead to inaccurate domain name authorization risk assessment, thus improving the accuracy of domain name authorization risk assessment and providing users with a higher quality service.

[0064] Based on this, the embodiments of this application provide a domain name authorization risk assessment method, referring to Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the domain name authorization risk assessment method for this application.

[0065] In this embodiment, the domain name authorization risk assessment method includes steps S01-S02:

[0066] Step S01: Receive multidimensional feature data of the domain name to be evaluated;

[0067] Before the implementation of this embodiment, it should be clear that DNS is a critical part of the Internet, responsible for domain name resolution. Its authorization mechanism has a complex structure with multiple levels and cross-domains. Node vulnerabilities can easily trigger a chain of security risks. Existing assessments mostly use rule scoring or linear weighting methods, which have problems such as incomplete indicators, insufficient characterization of extreme risks, difficulty in balancing interpretability and data-driven approaches, and inefficient labeling.

[0068] Therefore, in order to solve the above problems, this embodiment first receives multi-dimensional feature data of the domain name to be evaluated. The multi-dimensional feature data in this embodiment covers multiple dimensions of information that can reflect the domain name authorization risk, including but not limited to the domain name's own attribute features, domain name resolution record features, domain name registration information features, network behavior features, and associated domain name features.

[0069] Step S02: The multidimensional feature data is assessed for authorization risk using a pre-built domain name authorization risk prediction model to obtain the assessment result. The domain name authorization risk prediction model is trained using a pre-trained tree model and a pre-trained neural network model.

[0070] After receiving multidimensional feature data, authorization risk assessment can be performed using a pre-built domain name authorization risk prediction model to obtain the authorization risk assessment result. In this embodiment, the domain name authorization risk prediction model is trained using a pre-trained tree model and a pre-trained neural network model. Therefore, during the assessment process, if... Figure 2 As shown, the model draws on the multi-agent risk assessment logic from the labeled dataset construction stage. Based on the input multi-dimensional feature data, it analyzes the risk dimensions that LLM and DNS experts are concerned with in collaborative annotation. Specifically, it uses the rule-based feature matching capability of the pre-trained tree model and the complex feature mining capability of the pre-trained neural network to perform dual verification and analysis on the risk points related to domain name authorization dependency. Then, it combines the prediction results of the two to make a fusion judgment and finally outputs the authorization risk assessment result with the same dimension as the consensus analysis result, ensuring the accuracy and reliability of the assessment result.

[0071] Specifically, before step S02 above, which involves assessing the authorization risk of the multidimensional feature data using a pre-built domain name authorization risk prediction model to obtain the assessment result, the method further includes:

[0072] Step S0201: Collect Domain Name System (DNS) metrics;

[0073] Step S0202: Filter the DNS indicators to obtain the Hierarchical Analysis (AHP) model;

[0074] Step S0203: Calculate the weights of the hierarchical analysis (AHP) model to obtain the prior risk value;

[0075] Step S0204: Based on the prior risk value, perform stratified sampling to obtain domain name samples, and use a large language model to perform DeFell annotation on the domain name samples to obtain an annotated dataset;

[0076] Step S0205: Train and validate the pre-trained tree model and the pre-trained neural network model using the labeled dataset until the model performance of the pre-trained tree model and the pre-trained neural network model converges. Then, fuse the pre-trained tree model and the pre-trained neural network model to obtain the domain name authorization risk prediction model.

[0077] First, candidate indicators related to DNS authorization dependency security were collected from relevant research findings and RFC standards. Standardized mapping rules were defined for each indicator, uniformly mapping the original feature data to a risk degree vi∈[0,1], ensuring the comparability and quantifiability of different types of indicators. Candidate indicators were obtained by searching RFCs and recent academic papers and white papers on topics such as DNS authorization dependency, Glue records, resolution dependency chains, DNSSEC, AXFR, resolution availability, and Start of Authority (SOA) records. Specific indicators are shown in Table 1 below.

[0078]

[0079] Table 1

[0080] When collecting DNS metrics, we comprehensively cover various characteristic indicators from five perspectives: resource record configuration risk, DNS attack surface risk, data inconsistency risk, service vulnerability risk, and service stability risk, as shown in Table 1, to ensure the completeness and relevance of the collected data.

[0081] During the process of screening the collected DNS metrics, the candidate metrics were first systematically reviewed. Simultaneously, to achieve a unified assessment of DNS risk metrics, standardized mapping rules were defined, mapping various DNS risk metrics to a unified risk value range of 0-1, where 0 represents the safest and 1 represents the most dangerous. Based on the nature of the metrics, they were divided into the following five types and their corresponding mapping methods:

[0082] (1) Linear indicators, where risk increases or decreases monotonically with the value, such as SOA expire value, etc., with the following mapping formula:

[0083]

[0084] (2) Contrarian indicators: the smaller the value, the higher the risk. For example, the number of NS records. The mapping formula is:

[0085]

[0086] (3) Binary threshold type indicators, Boolean type or state type, directly determined by the existence of risk, such as whether AXFR is prohibited, the mapping formula is:

[0087]

[0088] (4) Proportional indicators, anomaly ratio indicators, where risk is calculated based on the anomaly ratio, such as the proportion of expired Glue records and the proportion of unresponsive NS records. The mapping formula is:

[0089]

[0090] (5) Segmented indicators, where the indicators have distinct safe / medium / high risk ranges and the risk values ​​change non-linearly, such as NS TTL and SOA retry values. The mapping formula is:

[0091]

[0092] After preprocessing the candidate indicators based on the above standardized mapping rules, an AHP hierarchical structure is established based on the preprocessed candidate indicators. By screening and removing indicators with high noise, poor stability, or low practical value, the high-value indicators that are retained are classified and merged according to the scope of security impact, risk trigger intensity, and historical security events, and finally the hierarchical analysis AHP model is obtained.

[0093] When calculating the weights of the Analytic Hierarchy Process (AHP) model, a scoring matrix is ​​constructed based on the opinions of experts after multiple rounds of convergence. After passing the consistency test, the weights of each indicator are output. The risk degree vi of each indicator is calculated in combination with the previously defined standardized mapping rules. The basic risk score is obtained by weighted summation. After the corresponding risk enhancement mechanism is introduced, the prior risk value is finally obtained.

[0094] When obtaining domain name samples through stratified sampling based on prior risk values, the prior risk values ​​are used as the stratification standard. Data is extracted evenly from a large number of unlabeled domain names according to risk intervals to ensure the representativeness and coverage of the samples. During the Delphi annotation process of the domain name samples through the large language model, the prior risk values ​​are used as the structured risk prompt input annotation system to assist the large language model and human experts in focusing on high-risk samples. After multiple rounds of iterations such as multi-party independent scoring, anonymized result aggregation, and consensus analysis, high-quality supervised samples are formed and divided into training set and validation set to obtain the labeled dataset.

[0095] When training and validating pre-trained tree models and pre-trained neural network models using labeled datasets, the training set of the labeled dataset is input into the two pre-trained models respectively. Supervised training is carried out based on the risk data after standardization and mapping of various indicators in Table 1. Simultaneously, the model performance is monitored using the validation set. The training and validation are iterated repeatedly until the performance of both models reaches a convergence state. At this point, the converged pre-trained tree model and pre-trained neural network model are fused to finally obtain the domain name authorization risk prediction model.

[0096] More specifically, step S0202 above, the step of filtering the DNS indicators to obtain the hierarchical analysis (AHP) model, includes:

[0097] Step S02021: Perform stability analysis on the DNS index to obtain the stability analysis results;

[0098] Step S02022: Based on the stability analysis results, low-value DNS indicators are removed to obtain high-value DNS indicators;

[0099] Step S02023: Based on the security impact scope, risk triggering intensity, and historical security events of the high-value DNS indicators, classify and merge the high-value DNS indicators to obtain the AHP model.

[0100] When conducting stability analysis on DNS metrics, the Delphi method was used to conduct multiple rounds of anonymous expert discussions. A large model was simultaneously introduced as an "expert" to participate in the analysis process. Combined with the reference opinions of the large model based on historical data and experience, the stability of each DNS metric was jointly assessed, identifying metrics with high noise and poor stability. For metrics with significant deviations in opinions during the discussions, anonymous feedback reports were generated and iterative discussions were conducted to bring expert opinions closer together. Finally, a stability analysis result containing information such as the stability level and degree of deviation of each metric was formed.

[0101] When eliminating low-value DNS indicators based on stability analysis results, and combining the discussion conclusions after multiple rounds of convergence of the Delphi method, in addition to eliminating indicators that are clearly instable in the stability analysis results, further screening and removal of indicators with low actual security value are carried out. These are those indicators that contribute little to the assessment of DNS authorization dependency risk and have a low correlation with security risk. After the above screening process, the indicators that are finally retained are the high-value DNS indicators. At the same time, the participation of the large model also improves the scientificity and efficiency of indicator value judgment.

[0102] Based on the security impact scope, risk triggering intensity, and historical security events of high-value DNS indicators, the classification and merging of high-value DNS indicators begins with identifying the inherent relationships between indicators according to the three core factors mentioned above. On this basis, a three-layer AHP model—target layer, criterion layer, and indicator layer—is constructed. The target layer is set as the assessment of DNS authorization dependency risk elements; the criterion layer is divided into five risk sources: resource record configuration risk, service exposure risk, data consistency risk, service vulnerability risk, and service stability risk; and the indicator layer corresponds to the previously screened high-value DNS indicators, specifically 21 feature values. After classifying and merging to form a preliminary hierarchical structure, the elements of the criterion layer and indicator layer are then analyzed in pairs. The importance principle of comparison is used for scoring, and a 1-9 scale is used for scoring. Subjective judgment is made by domain experts, while large models are also introduced as "experts" to participate in the scoring, providing reference scores based on historical data and experience. This balances expert experience with model data support to form a judgment matrix for AHP weight calculation, improving scoring efficiency and reproducibility. Considering that the scoring of the judgment matrix may have subjective bias, a consistency check is required to ensure the reasonable allocation of subsequent weights. This prevents contradictions in the matrix scores from affecting the weight calculation. If the consistency ratio CR > 0.1, experts are organized to re-score and iteratively construct the judgment matrix until the consistency check is passed, finally obtaining an AHP model with a reasonable structure and clear hierarchy.

[0103] Further, step S0203 above, the step of calculating the weights of the analytic hierarchy process (AHP) model to obtain the prior risk value, includes:

[0104] Step S02031: The importance of high-value DNS indicators in the AHP model is compared by scaling system to obtain the AHP judgment matrix;

[0105] Step S02032: Perform a consistency check on the AHP judgment matrix to obtain the check result;

[0106] Step S02033: If the verification result is passed, the indicator weights of several high-value DNS indicators in the AHP model are calculated using the feature vector method.

[0107] Step S02034: Calculate the basic weighted average and basic score of the several high-value DNS indicators according to the preset standardized mapping rules and the indicator weights.

[0108] Step S02035: The basic weighted average and basic score are enhanced with risk through hard failure rules and soft extreme nonlinear amplification mechanism, and the enhanced basic weighted average and basic score are mapped to the risk scoring interval to obtain the prior risk value.

[0109] Firstly, when comparing the importance of high-value DNS indicators in the AHP model through a scaling system, the 1-9 scaling method is used as the core scaling system. For each high-value DNS indicator in the AHP model's criterion layer (resource record configuration risk, service exposure risk, data consistency risk, service vulnerability risk, and service stability risk) and indicator layer, pairwise importance comparisons are carried out one by one. During the comparison process, in addition to the subjective judgment of domain experts based on experience, a large model is simultaneously introduced as an "extended expert" to participate in the scoring. The large model outputs reference scores based on historical security event data and a DNS authorization dependency risk case library, taking into account both expert experience and data-driven support. Finally, an AHP judgment matrix containing the importance ratios between each indicator is formed, improving the scientificity and reproducibility of the comparison results.

[0110] When performing consistency verification on the AHP judgment matrix, the consistency ratio (CR) is used to measure the degree of consistency of the matrix. The core judgment criterion is that if CR > 0.1, the judgment matrix has a scoring contradiction, and experts need to be organized to re-conduct the importance comparison and correct the judgment matrix based on the reference information fed back by the large model. If CR ≤ 0.1, the consistency of the judgment matrix meets the requirements, and finally, a verification result containing information such as consistency ratio and whether it passes is generated.

[0111] If the verification result is satisfactory, the weights of several high-value DNS indicators in the AHP model are calculated using the eigenvector method. Specifically, this is done by solving for the eigenvector corresponding to the largest eigenvalue of the judgment matrix and normalizing the eigenvector to obtain the weight set of each high-value DNS indicator. This set of weights intuitively reflects the contribution of each indicator to the assessment of DNS authorization dependency risk.

[0112] When calculating the basic weighted average and basic score of several high-value DNS indicators based on the preset standardized mapping rules and indicator weights, the basic score is first calculated using AHP weights. The specific formula is as follows:

[0113]

[0114] Then, the preset standardized mapping rules are invoked to normalize the raw data of each high-value DNS indicator to the [0, 1] interval, thereby obtaining the risk score corresponding to each indicator. (0 represents the safest, 1 represents the most dangerous), where the linear index uses the formula:

[0115]

[0116] The mapping, inverse indicator uses the following formula:

[0117]

[0118] For mapping, binary threshold, proportional, and segmented indicators are mapped using corresponding preset formulas. Based on this, the base score is calculated using a weighted sum formula, specifically:

[0119] B is the base score. Assigning weights to each indicator. The basic weighted calculation is performed simultaneously during the calculation of the risk degree after normalization of each indicator to obtain the basic weighted result.

[0120] When enhancing the risk of the basic weighting and basic score through hard failure rules and soft extreme nonlinear amplification mechanisms, soft extreme nonlinear amplification is performed first. A set of indicators with soft extreme anomaly effects is selected (such as Glue expiration rate >70%, NS failure rate >70%, excessive cross-country dependence, excessively long CNAME chain, etc.), and the anomaly factor E is calculated. The specific formula is as follows:

[0121]

[0122] In the formula For soft extreme index weights, To amplify the index (optional value 2-3), use the following formula:

[0123] The score after soft amplification is calculated, where α>0 is the amplification factor (optional value 0.2–0.5).

[0124] Then, hard failure rule validation is performed, and a pre-defined hard failure set is executed. This includes critical but low-probability errors such as NS server count = 0, all NS servers being invalid, all Glue servers being expired, DNS resolution loops (CNAME / SOA), SOA anomalies causing the domain to become completely unresolvable, open AXFR, and exposed private IP addresses. If any high-value DNS metric triggers a hard failure set, then the following formula will be used:

[0125] Where, in the formula This indicates that the corresponding indicator has triggered a hard failure, directly pushing the score into the high-risk range. After risk enhancement, the following mapping formula is used:

[0126] The risk-enhanced score is converted to a risk score range of 1–10, and the prior risk value S_AHP is finally obtained. This value can accurately reflect the high-impact, low-probability failure characteristics of DNS.

[0127] Furthermore, step S0204 above, which involves obtaining domain name samples through stratified sampling based on the prior risk value, and then labeling the domain name samples using a large language model to obtain a labeled dataset, includes:

[0128] Step S02041: Determine the stratified sampling criteria based on the prior risk value;

[0129] Step S02042: Perform balanced sampling on the unlabeled domain names through the hierarchical sampling annotation to obtain domain name samples;

[0130] Step S02043: After converting the prior risk value into a structured risk warning, based on the structured risk warning, the domain name samples are labeled as high-risk samples using a large language model to obtain high-quality supervised samples;

[0131] Step S02044: Divide the high-quality supervised samples into a training set and a validation set, and obtain the labeled dataset through the training set and the validation set.

[0132] When determining the stratified sampling standard based on the prior risk value, the previously obtained prior risk value S_AHP (i.e., the score value in the range of 1–10) is directly used as the core stratified sampling standard, providing a clear quantitative basis for subsequent equalization sampling. At the same time, this prior risk value will also be used as an important feature in subsequent model training.

[0133] When performing balanced sampling on unlabeled domains using stratified sampling criteria, the samples are evenly extracted from a massive number of unlabeled domains according to the risk range corresponding to the prior risk value Score. This ensures that domain samples of different risk levels are fully covered, guarantees the uniformity and representativeness of the sampling data to a certain extent, avoids the category imbalance problem caused by sampling bias, and finally obtains domain samples covering multiple risk levels.

[0134] After converting the prior risk value into a structured risk warning, when labeling high-risk samples of domain names using a large language model based on this structured risk warning, the prior risk value (Score) is first converted into a structured risk warning containing information such as risk level and core risk correlation indicators. This risk warning is then input into the labeling system to help the labeling entity focus on high-risk samples, improve labeling consistency, and reduce subjective noise. Simultaneously, a large language model (LLM) is introduced as an extended expert, collaborating with human experts to conduct multiple rounds of labeling using an improved Delphi method that integrates the knowledge of both. The specific labeling process is as follows: Figure 3As shown, in the first round of evaluation, based on expert and large model knowledge, each annotation subject independently evaluates all domain name samples and outputs an integer risk score of 1-10. During this stage, any form of communication or reference to external opinions is prohibited. Following this, the consensus analysis and feedback stage begins, summarizing the evaluation opinions of experts and LLMs, calculating the maximum absolute difference (MAD) for each domain name sample, and creating anonymous reports for disputed domain names (i.e., domain names with MAD > 2), clearly presenting the score differences among the subjects and the core points of contention. Then, iterative evaluation is conducted. Evaluators must fully consider the content of the anonymous feedback reports and reassess the domain name authorization dependency security risks. The evaluation terminates when any of the following conditions are met:

[0135] a) The MAD of all domain name samples is ≤ 2;

[0136] b) Reaching the maximum number of iteration rounds (e.g., the default 3 rounds);

[0137] c) The reduction rate of disputed domain names is less than 10% for two consecutive rounds.

[0138] Through the above multi-round iterative annotation process, high-quality supervised samples are finally obtained.

[0139] High-quality supervised samples are divided into training and validation sets. When obtaining the labeled dataset from the training and validation sets, the high-quality supervised samples, which have been confirmed by multiple rounds of Delphi annotation, are randomly divided into training and validation sets according to a preset division ratio (such as 7:3 or 8:2) to ensure the consistency of the distribution of the two sets of data. The two sets of data together constitute the labeled dataset required for subsequent model training and validation. At the same time, the prior risk value S_AHP, which was previously used as a structured hint, will also be incorporated into the labeled dataset as an important feature.

[0140] This embodiment, through the above-described scheme, specifically receives multi-dimensional feature data of the domain name to be evaluated; it then uses a pre-built domain name authorization risk prediction model to assess the authorization risk of the multi-dimensional feature data, obtaining the assessment result. The domain name authorization risk prediction model is trained using a pre-trained tree model and a pre-trained neural network model. Therefore, by inputting multi-dimensional feature data of the domain name to be evaluated, it overcomes the limitations of existing single indicators. Furthermore, the model training employs a pre-trained tree model and a neural network model for training and fusion, solving the problem of inefficient annotation and improving the quality of data evaluation. Finally, the trained domain name authorization risk prediction model is used for authorization risk assessment, thereby improving the accuracy of the assessment. This solves the problem of incomplete indicators and inefficient annotation in existing methods, which leads to inaccurate domain name authorization risk assessment, and improves the accuracy of domain name authorization risk assessment.

[0141] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 In step S0205, the pre-trained tree model and the pre-trained neural network model are trained and validated using the labeled dataset until their performance converges. Then, the pre-trained tree model and the pre-trained neural network model are fused to obtain the domain name authorization risk prediction model. In this step, the domain name authorization risk assessment method further includes steps S02051-S02055:

[0142] Step S02051: Input the training set of the labeled dataset into the pre-trained tree model and the pre-trained neural network model for supervised training to obtain the initial tree model and the initial neural network model.

[0143] Step S02052: Obtain an unlabeled sample set, and use the initial tree model and the initial neural network model to predict and estimate the uncertainty of the unlabeled sample set to obtain a consistency measure;

[0144] Step S02053: Based on the consistency metric, divide the unlabeled sample set into high-confidence samples, medium-confidence samples, and low-confidence samples;

[0145] Step S02054: The training set is corrected using the high-confidence samples, medium-confidence samples, and low-confidence samples to obtain the corrected training set;

[0146] Step S02055: Iteratively train the initial tree model and the initial neural network model using the corrected training set until the performance of the initial tree model and the initial neural network model converges. Then, fuse the initial tree model and the initial neural network model to obtain the domain name authorization risk prediction model.

[0147] First, when inputting the training set of the labeled dataset into the pre-trained tree model and the pre-trained neural network model for supervised training, a heterogeneous dual-teacher architecture is adopted for training. Among them, for the pre-trained tree model, models such as LightGBM or XGBoost, which have good adaptability to structured features (such as DNS features and protocol statistical features), are selected. Their training is stable and they have good interpretability (able to output feature importance). For the pre-trained neural network model, MLP or TabNet is selected, which can capture the non-linear relationships of high-dimensional and complex interaction features (such as domain name security-related features, ASN, geographical distribution, TTL, etc.). The training set L of the labeled dataset is input into the two types of pre-trained models respectively. Through supervised learning, the models learn the risk mapping rules in the labeled data, and finally obtain the initial tree model (Teacher A) and the initial neural network model (Teacher B). The two models are complementary in paradigm, taking into account stability, interpretability, and complex feature modeling capabilities.

[0148] After obtaining the unlabeled sample set, the initial tree model and the initial neural network model are used to predict and estimate the uncertainty of the unlabeled sample set U, and a consistency metric is obtained. Specifically, first, Teacher A and Teacher B respectively predict the unlabeled sample set U to obtain their respective prediction outputs and , and then calculate the average value of the predictions of the two models. The formula is:

[0149]

[0150] Next, calculate the degree of dispersion (variance) of the predictions of the two models. The formula is:

[0151]

[0152] Introduce an exponential mapping to map the variance to the confidence level in the [0, 1] interval as the consistency metric. The formula is:

[0153] In the formula is the scaling factor, which can adjust the sensitivity of the variance to the confidence level. The smaller the Var value, the larger the Conf value, indicating that the predictions of the two models are more consistent and the reliability of the pseudo-label is stronger.

[0154] When dividing the unlabeled sample set into high-confidence samples, medium-confidence samples, and low-confidence samples based on the consistency metric, a clear confidence threshold is set. Samples with Conf > 0.8 are classified as high-confidence samples, and such samples are used for self-training. Samples with 0.5 < Conf ≤ 0.8 are classified as medium-confidence samples and are used for cross-pseudo-label training. Samples with Conf ≤ 0.5 are classified as low-confidence samples and are used for active learning and manual review before being added to the next round of iterative training.

[0155] Specifically, when correcting the training set using high-confidence, medium-confidence, and low-confidence samples, a Cross-Teacher collaborative pseudo-label mechanism is employed. For high-confidence samples, pseudo-labels are directly generated and added to the training set for self-training. For medium-confidence samples, a cross-pseudo-label strategy is implemented, using teacher A's prediction as teacher B's pseudo-label and teacher B's prediction as teacher A's pseudo-label. During retraining, softlabels or confidence weighting are used for medium-confidence samples to reduce the impact of pseudo-label noise on training, enabling the two models to learn complementary information from the reliable pseudo-labels provided by each other. For low-confidence samples, they are first sent to the active learning pool for manual verification of labels before being added to the training set. Through the above differential processing, the training set is expanded and corrected, resulting in the corrected training set.

[0156] When iteratively training the initial tree model and the initial neural network model using the corrected training set, repeat the above process of sample prediction, confidence stratification, training set correction, and model training for 3-5 rounds. During each iteration, the model performance is evaluated using the validation set of the labeled dataset. Evaluation metrics include MAE, RMSE, QWK, etc., until the validation set performance tends to stabilize and the model is determined to have reached a performance convergence state. At this point, the converged initial tree model and the initial neural network model are fused to integrate the advantages of the two types of models, ultimately obtaining a domain name authorization risk prediction model with better generalization performance.

[0157] Specifically, step S02055 above, which involves fusing the initial tree model and the initial neural network model to obtain the domain name authorization risk prediction model, includes:

[0158] Step S020551: Calculate the model performance metrics of the initial tree model and the initial neural network model using the validation set of the labeled dataset;

[0159] Step S020552: Determine the fusion weights of the initial tree model and the initial neural network model based on the model performance indicators;

[0160] Step S020553: Based on the fusion weights, the initial tree model and the initial neural network model are fused using a weighted summation formula to obtain the domain name authorization risk prediction model.

[0161] First, the model performance metrics of the initial tree model and the initial neural network model are calculated using the validation set of the labeled dataset. The mean absolute error (MAE), root mean square error (RMSE), and quadratic weighted Kappa coefficient (QWK) are selected as the core performance evaluation metrics. The validation set of the labeled dataset is input into the converged initial tree model and the initial neural network model, and the specific values ​​of the two models on the above metrics are calculated respectively, forming a complete set of model performance metrics, which provides a quantitative basis for determining the subsequent fusion weights.

[0162] The fusion weights of the initial tree model and the initial neural network model are determined based on the model performance metrics. The performance of each model is the core criterion, and a rule is set that higher performance corresponds to higher fusion weights. Finally, a weighted fusion method can be used to fuse the tree model and the neural network predictions to improve robustness. Weights can be determined based on the validation set AUC or the consistency metric Conf. The better the model's performance on the validation set, the more reliable its predictions are, and it can be given higher weights in the fusion. Specifically, multiple performance metrics can be selected to adapt to different scenarios. For example, the normalized result of the quadratic weighted Kappa coefficient (QWK) can be used as a reference for weight allocation, or the validation set AUC value can be used to calculate the weights. The corresponding weight calculation formula is as follows:

[0163]

[0164] in, The fusion weights for the initial tree model, The fusion weights for the initial neural network model. This represents the AUC value of the initial tree model on the validation set. The AUC value of the initial neural network model on the validation set is used to determine the weights of the two models by calculating the proportion of their respective performance metrics. This ensures that the weight allocation matches the actual predictive ability of the models, while also ensuring that the sum of the fusion weights of the initial tree model and the initial neural network model is 1, providing a reasonable premise for subsequent weighted fusion.

[0165] Based on the aforementioned fusion weights, the initial tree model and the initial neural network model are fused using a weighted summation formula, which is as follows:

[0166]

[0167] Where Y is the risk prediction value after fusion. The fusion weights for the initial tree model, These are the predicted values ​​from the initial tree model. The fusion weights for the initial neural network model. (The predicted value of the initial neural network model) is used to calculate the risk prediction result after fusion by substituting the prediction results of the two models for the same domain name sample into the formula. This weighted integration method fully integrates the stability and interpretability advantages of the initial tree model with the complex feature modeling capability of the initial neural network model, and finally obtains the domain name authorization risk prediction model.

[0168] This embodiment, through the above scheme, specifically involves inputting the training set of the labeled dataset into the pre-trained tree model and the pre-trained neural network model for supervised training to obtain the initial tree model and the initial neural network model; obtaining an unlabeled sample set, and using the initial tree model and the initial neural network model to predict and estimate the uncertainty of the unlabeled sample set to obtain a consistency metric; dividing the unlabeled sample set into high-confidence samples, medium-confidence samples, and low-confidence samples based on the consistency metric; correcting the training set using the high-confidence samples, medium-confidence samples, and low-confidence samples to obtain a corrected training set; iteratively training the initial tree model and the initial neural network model using the corrected training set until the performance of the initial tree model and the initial neural network model converges, and then fusing the initial tree model and the initial neural network model to obtain the domain name authorization risk prediction model. Therefore, by inputting multidimensional feature data of the domain name to be evaluated, the limitations of existing single indicators are overcome. Furthermore, the pre-trained tree model and neural network model are used for training and fusion in the model training process to solve the problem of inefficient annotation and improve the quality of data evaluation. Finally, the trained domain name authorization risk prediction model is used to conduct authorization risk assessment, thereby improving the accuracy of the assessment. This solves the problem that existing methods suffer from incomplete indicators and inefficient annotation, which leads to inaccurate domain name authorization risk assessment and improves the accuracy of domain name authorization risk assessment.

[0169] For example, to help understand the implementation process of the domain name authorization risk assessment method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 5 , Figure 5 A simplified flowchart illustrating a domain name authorization risk assessment method is provided, specifically:

[0170] First, we construct multi-dimensional initial indicators, collect candidate indicators based on RFC standards and related literature, cover five risk perspectives and define standardized mapping rules to achieve unified quantitative indicators.

[0171] Subsequently, high-value indicators were screened using the Delphi method, and a three-layer AHP model was constructed. Experts and the large model collaborated to compare the importance of the indicators and form an AHP judgment matrix.

[0172] After the consistency test of the AHP judgment matrix, the weights of each indicator and the basic risk scores are calculated.

[0173] The score is enhanced by hard failure rules and nonlinear amplification mechanisms, and the prior risk value is mapped to it.

[0174] Based on the hierarchical sampling of domain name samples using prior risk values, a labeled dataset is formed through Delphi annotation using LLM and human experts.

[0175] A heterogeneous dual-teacher architecture is used to conduct semi-supervised iterative training. The training set is adjusted by stratifying the sample confidence level to optimize model performance.

[0176] After the model performance converges, the fusion weights are determined based on the validation set performance, and the weighted fusion is used to obtain the final domain name authorization risk prediction model.

[0177] In practical applications, inputting the multidimensional feature data of the domain name to be evaluated into the model will output the authorization risk assessment results.

[0178] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the domain name authorization risk assessment method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0179] This application also provides a domain name authorization risk assessment device, please refer to... Figure 6 The domain name authorization risk assessment device includes:

[0180] The receiving module 10 is used to receive multidimensional feature data of the domain name to be evaluated;

[0181] The evaluation module 20 is used to evaluate the authorization risk of the multidimensional feature data through a pre-built domain name authorization risk prediction model and obtain the evaluation result. The domain name authorization risk prediction model is trained by a pre-trained tree model and a pre-trained neural network model.

[0182] The domain name authorization risk assessment device provided in this application, employing the domain name authorization risk assessment method described in the above embodiments, can solve the technical problem of incomplete indicators and inefficient labeling in existing methods, leading to inaccurate domain name authorization risk assessment. Compared with the prior art, the beneficial effects of the domain name authorization risk assessment device provided in this application are the same as those of the domain name authorization risk assessment method provided in the above embodiments, and other technical features in the domain name authorization risk assessment device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0183] This application provides a domain name authorization risk assessment device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the domain name authorization risk assessment method in Embodiment 1 above.

[0184] The following is for reference. Figure 7 The diagram illustrates a structural schematic of a domain name authorization risk assessment device suitable for implementing embodiments of this application. The domain name authorization risk assessment device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The domain name authorization risk assessment device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0185] like Figure 7As shown, the domain name authorization risk assessment device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the xxx device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the domain name authorization risk assessment device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows domain name authorization risk assessment devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0186] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0187] The domain name authorization risk assessment device provided in this application, employing the domain name authorization risk assessment method described in the above embodiments, can solve the technical problem of incomplete indicators and inefficient labeling in existing methods, leading to inaccurate domain name authorization risk assessment. Compared with the prior art, the beneficial effects of the domain name authorization risk assessment device provided in this application are the same as those of the domain name authorization risk assessment method provided in the above embodiments, and other technical features in this domain name authorization risk assessment device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0188] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0189] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0190] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the domain name authorization risk assessment method in the above embodiments.

[0191] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0192] The aforementioned computer-readable storage medium may be included in the domain name authorization risk assessment device; or it may exist independently and not be assembled into the domain name authorization risk assessment device.

[0193] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by the domain name authorization risk assessment device, the domain name authorization risk assessment device: receives multi-dimensional feature data of the domain name to be assessed; performs authorization risk assessment on the multi-dimensional feature data through a pre-built domain name authorization risk prediction model, and obtains an assessment result, wherein the domain name authorization risk prediction model is trained through a pre-trained tree model and a pre-trained neural network model.

[0194] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0195] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0196] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0197] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described domain name authorization risk assessment method. This solves the technical problem of incomplete indicators and inefficient labeling in existing methods, leading to inaccurate domain name authorization risk assessment. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the domain name authorization risk assessment method provided in the above embodiments, and will not be repeated here.

[0198] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the domain name authorization risk assessment method described above.

[0199] The computer program product provided in this application can solve the technical problem that existing methods suffer from incomplete indicators and inefficient labeling, leading to inaccurate domain name authorization risk assessment. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the domain name authorization risk assessment method provided in the above embodiments, and will not be repeated here.

[0200] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for assessing domain name authorization risk, characterized in that, The domain name authorization risk assessment method includes: Receive multidimensional feature data of the domain name to be evaluated; The multidimensional feature data is assessed for authorization risk using a pre-built domain name authorization risk prediction model to obtain the assessment result. The domain name authorization risk prediction model is trained using a pre-trained tree model and a pre-trained neural network model. Prior to the step of assessing the authorization risk of the multidimensional feature data using a pre-built domain name authorization risk prediction model to obtain the assessment result, the method further includes: Collect Domain Name System (DNS) metrics; The DNS metrics are filtered to obtain the hierarchical analysis (AHP) model; The step of filtering the DNS indicators to obtain the hierarchical analysis (AHP) model includes: A stability analysis was performed on the DNS metrics, and the results were obtained. Based on the stability analysis results, low-value DNS indicators are removed to obtain high-value DNS indicators. Based on the security impact scope, risk triggering intensity, and historical security events of the high-value DNS indicators, the high-value DNS indicators are classified and merged to obtain the AHP model. The weights of the analytic hierarchy process (AHP) model are calculated to obtain the prior risk value. The step of calculating the weights of the hierarchical analysis (AHP) model to obtain the prior risk value includes: The AHP judgment matrix is ​​obtained by traversing and comparing the importance of high-value DNS indicators in the AHP model through a scaling system. The consistency of the AHP judgment matrix is ​​checked to obtain the check result; If the verification result is passed, the index weights of several high-value DNS indicators in the AHP model are calculated using the feature vector method. The basic weighted average and basic score of the several high-value DNS indicators are calculated based on the preset standardized mapping rules and the indicator weights. The basic weighted average and basic score are enhanced with risk through hard failure rules and soft extreme nonlinear amplification mechanism, and the enhanced basic weighted average and basic score are mapped to the risk scoring interval to obtain the prior risk value. Domain name samples are obtained by hierarchical sampling based on the prior risk values, and Delphi annotation is performed on the domain name samples using a large language model to obtain an annotated dataset; The pre-trained tree model and the pre-trained neural network model are trained and validated using the labeled dataset until their performance converges. Then, the pre-trained tree model and the pre-trained neural network model are fused to obtain the domain name authorization risk prediction model.

2. The domain name authorization risk assessment method as described in claim 1, characterized in that, The steps of obtaining domain name samples through stratified sampling based on the prior risk value, and then annotating the domain name samples using a large language model to obtain an annotated dataset include: The stratified sampling criteria are determined based on the prior risk value; The unlabeled domain names are sampled evenly using the aforementioned stratified sampling criteria to obtain a domain name sample; After converting the prior risk value into a structured risk warning, based on the structured risk warning, the domain name samples are labeled as high-risk samples using a large language model to obtain high-quality supervised samples. The high-quality supervised samples are divided into a training set and a validation set, and the labeled dataset is obtained through the training set and the validation set.

3. The domain name authorization risk assessment method as described in claim 1, characterized in that, The step of training and validating the pre-trained tree model and the pre-trained neural network model using the labeled dataset until the model performance of the pre-trained tree model and the pre-trained neural network model converges, and then fusing the pre-trained tree model and the pre-trained neural network model to obtain the domain name authorization risk prediction model includes: The training set of the labeled dataset is input into the pre-trained tree model and the pre-trained neural network model for supervised training to obtain the initial tree model and the initial neural network model. Obtain an unlabeled sample set, and use the initial tree model and initial neural network model to predict and estimate the uncertainty of the unlabeled sample set to obtain a consistency measure. Based on the consistency metric, the unlabeled sample set is divided into high-confidence samples, medium-confidence samples, and low-confidence samples; The training set is corrected by using the high-confidence samples, medium-confidence samples, and low-confidence samples to obtain the corrected training set; The initial tree model and the initial neural network model are iteratively trained using the modified training set until their performance converges. Then, the initial tree model and the initial neural network model are fused to obtain the domain name authorization risk prediction model.

4. The domain name authorization risk assessment method as described in claim 3, characterized in that, The step of fusing the initial tree model and the initial neural network model to obtain the domain name authorization risk prediction model includes: The model performance metrics of the initial tree model and the initial neural network model are calculated using the validation set of the labeled dataset. The fusion weights of the initial tree model and the initial neural network model are determined based on the model performance metrics. Based on the fusion weights, the initial tree model and the initial neural network model are fused using a weighted summation formula to obtain the domain name authorization risk prediction model.

5. A domain name authorization risk assessment device, characterized in that, The domain name authorization risk assessment device includes: The receiving module is used to receive multidimensional feature data of the domain name to be evaluated; An evaluation module is used to evaluate the authorization risk of the multidimensional feature data using a pre-built domain name authorization risk prediction model to obtain an evaluation result. The domain name authorization risk prediction model is trained using a pre-trained tree model and a pre-trained neural network model. The evaluation module is also used to collect Domain Name System (DNS) metrics. The DNS metrics are filtered to obtain the hierarchical analysis (AHP) model; The evaluation module is also used to perform stability analysis on the DNS index and obtain stability analysis results. Based on the stability analysis results, low-value DNS indicators are removed to obtain high-value DNS indicators. Based on the security impact scope, risk triggering intensity, and historical security events of the high-value DNS indicators, the high-value DNS indicators are classified and merged to obtain the AHP model. The weights of the analytic hierarchy process (AHP) model are calculated to obtain the prior risk value. The AHP judgment matrix is ​​obtained by traversing and comparing the importance of high-value DNS indicators in the AHP model through a scaling system. The consistency of the AHP judgment matrix is ​​checked to obtain the check result; If the verification result is passed, the index weights of several high-value DNS indicators in the AHP model are calculated using the feature vector method. The basic weighted average and basic score of the several high-value DNS indicators are calculated based on the preset standardized mapping rules and the indicator weights. The basic weighted average and basic score are enhanced with risk through hard failure rules and soft extreme nonlinear amplification mechanism, and the enhanced basic weighted average and basic score are mapped to the risk scoring interval to obtain the prior risk value. Domain name samples are obtained by hierarchical sampling based on the prior risk values, and Delphi annotation is performed on the domain name samples using a large language model to obtain an annotated dataset; The pre-trained tree model and the pre-trained neural network model are trained and validated using the labeled dataset until their performance converges. Then, the pre-trained tree model and the pre-trained neural network model are fused to obtain the domain name authorization risk prediction model.

6. A domain name authorization risk assessment device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the domain name authorization risk assessment method as described in any one of claims 1 to 4.

7. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the domain name authorization risk assessment method as described in any one of claims 1 to 4.