Domain name protection policy determination method and device and related equipment

CN122764686APending Publication Date: 2026-09-15CHINA INTERNET NETWORK INFORMATION CENTER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611110017.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-15

Smart Images

  • Figure CN122764686A_ABST
    Figure CN122764686A_ABST
Patent Text Reader

Abstract

A protection policy determination method and device related to domain names and related equipment, relating to the technical field of network security. The method comprises: obtaining a to-be-detected domain name and associated data formed in a data processing process related to the to-be-detected domain name; determining statistical features according to the to-be-detected domain name and the associated data, and determining semantic features according to the to-be-detected domain name, the semantic features being used to represent the context features of the to-be-detected domain name based on a domain name query sequence in a historical time period, the domain name query sequence comprising a plurality of domain names arranged according to query time; and determining a protection policy for the to-be-detected domain name according to the statistical features and the semantic features. Thus, the statistical features and the semantic features can provide feature basis for the determination of the protection policy from different dimensions, wherein the semantic features can collectively represent the context information formed based on the domain name query sequence, enhance the ability to distinguish the context differences between different to-be-detected domain names, and thus improve the accuracy of the protection policy determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to a method, apparatus and related equipment for determining protection strategies for domain names. Background Technology

[0002] Domain names, as crucial resource identifiers in network access, are widely used in scenarios such as web page access, service scheduling, service connection, and network communication. With the continuous evolution of network attack methods, malicious domain names may be used for phishing attacks, malware downloads, command and control communications, and data leaks. To ensure network access security, it is typically necessary to determine appropriate protection strategies for each domain name.

[0003] In the current process of determining domain name protection strategies, risk assessment is typically based on domain-related data, network access records, existing security data, or pre-defined handling rules. Based on the risk assessment results, protective measures such as alerts, blocking, redirection, or manual review are then implemented. However, in practical applications, determining and implementing protection strategies in this way may result in insufficient or excessive protection. For example, protection strategies determined for some malicious domains may be insufficient to effectively restrict related access, leading to poor domain name protection. Summary of the Invention

[0004] This application provides a method for determining domain name protection strategies to improve the accuracy of strategy determination. Furthermore, this application also provides a corresponding device, computing equipment, computer-readable storage medium, and computer program product for determining domain name protection strategies.

[0005] Firstly, this application provides a method for determining a protection strategy for a domain name, which can be executed by a device or computing device with data processing capabilities. For ease of explanation, the following description uses a protection strategy determination device executing the method as an example. Specifically, the protection strategy determination device acquires the domain name to be detected and its associated data, which is formed during data processing related to the domain name to be detected. Then, the protection strategy determination device determines the statistical characteristics and semantic characteristics of the domain name to be detected, wherein the statistical characteristics are determined based on the domain name to be detected and the associated data; the semantic characteristics are used to characterize the contextual characteristics of the domain name to be detected based on a domain name query sequence within a historical time period, the domain name query sequence including multiple domain names arranged according to query time, and the semantic characteristics are determined based on the domain name to be detected. Next, the protection strategy determination device determines a protection strategy for the domain name to be detected based on the statistical characteristics and semantic characteristics.

[0006] Thus, when determining a protection strategy, the protection strategy determination device can simultaneously utilize statistical features determined by the domain name to be detected and its associated data, as well as semantic features determined based on the domain name to be detected. This allows the determination of the protection strategy to no longer rely on a single type of feature, but rather to characterize the domain name to be detected from different dimensions through statistical and semantic features, thereby providing a more comprehensive feature basis for accurately determining the protection strategy. Furthermore, the domain name query sequence includes multiple domain names arranged according to query time, and semantic features can characterize the contextual features of the domain name to be detected based on this domain name query sequence. Compared to scattered domain name query records, semantic features can centrally reflect the contextual information formed by domain name query sequences within a historical time period, enhancing the ability of the protection strategy determination process to distinguish contextual differences between different domain names to be detected, thereby improving the accuracy of the protection strategy determination.

[0007] In one possible implementation, the protection strategy determination device determines scenario information of the domain name to be detected based on statistical features. This scenario information indicates the probability that the domain name belongs to each of multiple scenarios. Then, the device calculates a first contribution level of the statistical features and a second contribution level of the semantic features based on the statistical features, semantic features, and scenario information. The first contribution level characterizes the contribution of the statistical features in determining the protection strategy, and the second contribution level characterizes the contribution of the semantic features. Next, the device determines a protection strategy for the domain name to be detected based on the statistical features, semantic features, the first contribution level, and the second contribution level. In this way, the contribution levels of the statistical features and semantic features in determining the protection strategy can be dynamically adjusted according to the probability that the domain name belongs to each of the multiple scenarios, avoiding the two types of features always participating in the protection strategy determination with a fixed contribution level, thereby improving the scenario adaptability of the protection strategy determination process.

[0008] In one possible implementation, the protection strategy determination device projects statistical features and semantic features into a feature space, respectively, to obtain a first projection result corresponding to the statistical features and a second projection result corresponding to the semantic features. The first and second projection results have the same dimension. Then, based on the first and second projection results and scene information, the protection strategy determination device calculates a first weight for the statistical features and a second weight for the semantic features. The first weight characterizes a first degree of contribution, and the second weight characterizes a second degree of contribution. Next, the protection strategy determination device fuses the first and second projection results according to the first and second weights to obtain fused features, and determines a protection strategy for the domain name to be detected based on the fused features. Thus, through projection processing, statistical features and semantic features can be converted into projection results of the same dimension, enabling the first projection result and the second projection result to be processed together. Furthermore, the first weight and the second weight respectively characterize the contribution degree of the two types of features, allowing the first projection result and the second projection result to participate in the fusion processing according to their respective weights. Based on the obtained fusion features, a protection strategy for the domain name to be detected is determined. This realizes the process of the two types of features participating in the determination of the protection strategy according to their respective contribution degree as a calculable and executable processing method, improving the feasibility of the protection strategy determination process.

[0009] In one possible implementation, the protection strategy determination device processes statistical features through a scene evaluation model to obtain scene information. The scene evaluation model is trained based on multiple sample statistical features and the scene annotation results corresponding to each of the multiple sample statistical features. The scene annotation results indicate the scene to which the sample statistical features belong. Thus, by training the scene evaluation model using multiple sample statistical features with scene annotation results, the scene evaluation model can learn the feature differences of sample statistical features under different scenes. When processing the statistical features of the domain name to be detected, it determines the probability that the domain name belongs to each of the multiple scenes, thereby clarifying the method of obtaining scene information and enabling automatic and reliable determination of the scene information of the domain name to be detected.

[0010] In one possible implementation, the protection strategy determination device determines policy matching information for the domain name to be detected based on statistical and semantic features. The policy matching information includes at least one of the risk level and security type of the domain name to be detected. Then, the protection strategy determination device determines a protection strategy for the domain name to be detected based on the policy matching information. Thus, the risk level, security type, or a combination thereof of the domain name to be detected can be determined first based on statistical and semantic features, and the determined result can be used as policy matching information. Then, a corresponding protection strategy can be determined based on the policy matching information, enabling the protection strategy to be determined according to the risk level, security type, or a combination thereof of the domain name to be detected, thereby improving the targeting of the protection strategy determination.

[0011] In one possible implementation, the policy matching information includes the risk level and security type of the domain name to be detected. The protection policy determination device queries the policy mapping relationship based on the risk level and security type to obtain the protection policy for the domain name to be detected. The policy mapping relationship is the correspondence between risk level, security type, and protection policy. Thus, the policy mapping relationship can be queried based on different combinations of risk levels and security types to obtain corresponding protection policies, ensuring that different combinations of risk levels and security types can be matched with appropriate protection policies, thereby improving the granularity of protection policy determination.

[0012] Furthermore, before querying the policy mapping relationship based on risk level and security type, if the domain name to be detected meets the strong rule triggering conditions, the protection policy determination device determines the risk level of the domain name to be detected as high-risk. Then, based on the high-risk level and the security type mapping relationship of the domain name to be detected, it obtains the protection policy for the domain name to be detected. In this way, domain names to be detected that meet the strong rule triggering conditions can be matched with corresponding protection policies according to their high-risk level and security type, making the protection policy determined for such domain names adaptable to their high-risk situation, thereby improving the overall reliability of the protection policy determination.

[0013] Furthermore, after obtaining the protection strategy for the domain to be tested, if the domain to be tested meets the whitelist downgrade conditions, the protection strategy determination device adjusts the protection actions in the protection strategy to alarm actions or manual review actions. In this way, after determining the protection strategy based on risk level and security type, the protection actions can be further adjusted according to the whitelist downgrade conditions, allowing the protection actions for the domain to be tested to be adapted to the whitelist downgrade conditions, thereby improving the flexibility of protection strategy determination.

[0014] In one possible implementation, the protection strategy determination device determines multiple statistical sub-features of the domain name to be detected based on the domain name and associated data. These multiple statistical sub-features include at least two of the following: domain name character features, domain name resolution features, domain name registration information features, digital certificate features, threat intelligence features, and encrypted domain name resolution traffic features. Then, the protection strategy determination device generates statistical features of the domain name to be detected based on these multiple statistical sub-features. In this way, the statistical features can comprehensively characterize the statistical information of the domain name to be detected across at least two data dimensions, thereby enriching the information contained in the statistical features and providing a more comprehensive statistical feature foundation for subsequent determination of protection strategies by combining semantic features.

[0015] In one possible implementation, the protection strategy determination device processes the domain name to be detected using a domain name semantic model to obtain semantic features. The domain name semantic model is trained based on multiple training samples. Each training sample includes a center domain and associated domains, which are domains in the domain query sequence that satisfy the context selection criteria. Thus, by constructing training samples using center domains and associated domains that satisfy the context selection criteria in the domain query sequence, the domain name semantic model can learn the contextual relationships between domain names. Furthermore, the trained domain name semantic model can be used to determine the semantic features of the domain name to be detected, thereby clarifying the method of semantic feature determination and enabling the semantic features to centrally reflect the contextual information formed by domain query sequences within a historical time period. This provides a more effective semantic feature basis for subsequent determination of protection strategies by combining statistical features.

[0016] Secondly, this application provides a device for determining a protection strategy for a domain name. The device includes an acquisition module, a feature determination module, and a strategy determination module. The acquisition module acquires the domain name to be detected and its associated data, which is formed during data processing related to the domain name. The feature determination module determines the statistical and semantic features of the domain name to be detected. The statistical features are determined based on the domain name and associated data, while the semantic features characterize the contextual features of the domain name based on a domain name query sequence within a historical time period. The domain name query sequence includes multiple domain names arranged according to query time, and the semantic features are determined based on the domain name. The strategy determination module determines a protection strategy for the domain name to be detected based on the statistical and semantic features.

[0017] In one possible implementation, the strategy determination module is used to determine the scenario information of the domain name to be detected based on statistical features, wherein the scenario information indicates the probability that the domain name to be detected belongs to each of multiple scenarios; calculate a first contribution degree of the statistical features and a second contribution degree of the semantic features based on the statistical features, semantic features, and scenario information, wherein the first contribution degree characterizes the contribution degree of the statistical features in determining the protection strategy, and the second contribution degree characterizes the contribution degree of the semantic features in determining the protection strategy; and determine the protection strategy for the domain name to be detected based on the statistical features, semantic features, first contribution degree, and second contribution degree.

[0018] In one possible implementation, the strategy determination module is used to project statistical features and semantic features onto a feature space respectively, to obtain a first projection result corresponding to the statistical features and a second projection result corresponding to the semantic features, wherein the dimension of the first projection result and the dimension of the second projection result are the same; based on the first projection result, the second projection result and scene information, a first weight of the statistical features and a second weight of the semantic features are calculated, wherein the first weight is used to characterize a first degree of contribution and the second weight is used to characterize a second degree of contribution; based on the first weight and the second weight, the first projection result and the second projection result are fused to obtain a fused feature; and based on the fused feature, a protection strategy for the domain name to be detected is determined.

[0019] In one possible implementation, the strategy determination module is used to process statistical features through a scene evaluation model to obtain scene information. The scene evaluation model is trained based on multiple sample statistical features and the scene annotation results corresponding to each sample statistical feature. The scene annotation results are used to indicate the scene to which the sample statistical features belong.

[0020] In one possible implementation, the policy determination module is used to determine policy matching information of the domain name to be detected based on statistical features and semantic features, the policy matching information including at least one of the risk level and security type of the domain name to be detected; and to determine a protection policy for the domain name to be detected based on the policy matching information.

[0021] In one possible implementation, the policy matching information includes the risk level and security type of the domain name to be detected. The policy determination module is used to query the policy mapping relationship based on the risk level and security type to obtain the protection policy for the domain name to be detected. The policy mapping relationship is the correspondence between the risk level, security type and protection policy.

[0022] Furthermore, the strategy determination module is also used to determine the risk level of the domain to be detected as high-risk if the domain to be detected meets the strong rule triggering conditions before querying the strategy mapping relationship based on the risk level and security type; and to query the strategy mapping relationship based on the high-risk level and the security type of the domain to be detected to obtain the protection strategy for the domain to be detected.

[0023] Furthermore, the policy determination module is also used to adjust the protection actions in the protection policy to alarm actions or manual review actions after obtaining the protection policy for the domain name to be detected, provided that the domain name to be detected meets the whitelist downgrade conditions.

[0024] In one possible implementation, the feature determination module is used to determine multiple statistical sub-features of the domain name to be detected based on the domain name to be detected and associated data. The multiple statistical sub-features include at least two of the following: domain name character features, domain name resolution features, domain name registration information features, digital certificate features, threat intelligence features, and encrypted domain name resolution traffic features; and to generate statistical features of the domain name to be detected based on the multiple statistical sub-features.

[0025] In one possible implementation, the feature determination module is used to process the domain name to be detected through a domain name semantic model to obtain semantic features. The domain name semantic model is trained based on multiple training samples. Each training sample includes a central domain name and associated domain names, which are domain names in the domain name query sequence that meet the context selection conditions.

[0026] In one possible implementation, the protection strategy determination device may further include a data preprocessing module. The data preprocessing module performs data preprocessing on at least one of the domain name to be detected and associated data. For example, the data preprocessing module may perform format conversion, duplicate data removal, abnormal data processing, missing data processing, normalization processing, or encoding processing, depending on the specific characteristics of the data to be processed. The feature determination module determines the statistical or semantic features of the domain name to be detected based on the processing results output by the data preprocessing module.

[0027] In one possible implementation, the protection strategy determination device may further include a data storage module. The data storage module may be used to store the domain name to be detected, associated data, statistical features, semantic features, protection strategies, or other data related to the determination of the protection strategy.

[0028] It should be noted that the above modules are merely examples. In other possible implementations, the protection strategy determination device may also include other modules, omit some optional modules, divide the functions implemented by one module into multiple modules, or integrate the functions implemented by multiple modules into the same module. This application does not limit this.

[0029] The domain name protection strategy determination device provided in the second aspect corresponds to the domain name protection strategy determination method provided in the first aspect. Therefore, the technical effects of each embodiment in the second aspect can be found in the corresponding embodiments in the first aspect, and will not be repeated here.

[0030] Thirdly, this application provides a computing device. The computing device includes at least one processor and at least one memory, the at least one memory for storing instructions, and the at least one processor for executing the instructions stored in the at least one memory, such that the computing device performs the domain name protection strategy determination method provided in the first aspect or any possible implementation thereof.

[0031] In one possible implementation, communication is possible between at least one processor and at least one memory. The at least one memory may be integrated into or independent of the at least one processor. The computing device may also include a bus through which the at least one processor can be connected to the at least one memory.

[0032] Fourthly, this application provides a computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed on a computing device, cause the computing device to perform the domain name protection strategy determination method provided in the first aspect or any possible implementation thereof.

[0033] Fifthly, this application provides a computer program product. The computer program product includes instructions that, when executed on a computing device, cause the computing device to perform the domain name protection strategy determination method provided in the first aspect or any possible implementation thereof.

[0034] Based on the implementation methods provided in the above aspects, this application can be further combined to obtain more implementation methods. Attached Figure Description

[0035] Figure 1 This application provides a schematic diagram of the structure of a domain name protection system;

[0036] Figure 2 A schematic diagram illustrating the processing relationship for determining protection strategies based on statistical and semantic features, provided for this application;

[0037] Figure 3 A flowchart illustrating a method for determining a domain name protection strategy provided in this application;

[0038] Figure 4 A schematic diagram illustrating a statistical feature generation process provided in this application;

[0039] Figure 5 A schematic diagram illustrating the domain name semantic model training and semantic feature determination process provided in this application;

[0040] Figure 6 A schematic diagram illustrating a process for adjusting and fusing feature contribution levels based on scene information, as provided in this application;

[0041] Figure 7 This application provides a schematic diagram illustrating the process of determining and adjusting protection strategies based on risk level and security type.

[0042] Figure 8 A schematic diagram of a domain name protection strategy determination device provided in this application;

[0043] Figure 9 A schematic diagram of the structure of a computing device provided in this application. Detailed Implementation

[0044] The technical solutions of the embodiments of this application are described below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of this application, and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application.

[0045] It should be noted that the terms "first" and "second" in this application specification and claims are used only to distinguish objects with the same or similar meanings, and are not used to limit the execution order or quantitative relationship between objects. For example, the first degree of contribution and the second degree of contribution are used to distinguish the degree of contribution of statistical features and semantic features in the process of determining the protection strategy; the first projection result and the second projection result are used to distinguish the projection results corresponding to statistical features and semantic features, respectively.

[0046] To facilitate understanding of the embodiments of this application, the relevant concepts involved in this application will be explained below.

[0047] The domain name to be detected can be a domain name for which a protection strategy needs to be determined. The domain name to be detected can be represented by a domain name string, a domain name identifier corresponding to the domain name string, or other data forms that can indicate the corresponding domain name. This application embodiment does not limit this.

[0048] Correlated data can be data generated during data processing related to the domain name to be detected. For example, related data may include domain name resolution data generated during domain name resolution, registration information data generated during domain name registration, certificate data generated during digital certificate application or use, intelligence data generated during threat intelligence processing, and traffic data generated during encrypted domain name resolution communication. Correlated data can reflect the network activity and security-related information of the domain name to be detected from different dimensions.

[0049] Statistical features can be characteristics determined based on the domain name to be detected and related data. Statistical features can be used to characterize the statistical information of the domain name under different data dimensions. For example, statistical features may include domain character features to characterize the composition of the domain name string, domain name resolution features to characterize domain name resolution behavior, domain name registration information features to characterize domain name registration information, digital certificate features to characterize digital certificate-related information, threat intelligence features to characterize threat intelligence hit rates, and encrypted domain name resolution traffic features to characterize encrypted domain name resolution traffic behavior, etc.

[0050] A domain query sequence can include multiple domains arranged according to query time. The domain query sequence can be constructed based on multiple domain query records formed within a historical time period. For example, domain query records can originate from domain query logs, which can record query time and the queried domain. The domain query sequence can be formed by arranging multiple domains according to query time based on the time information in multiple domain query records. The length of the historical time period, the selection range of domain query records, and the grouping method can be determined according to the specific application scenario. For example, the historical time period can be several minutes, several hours, or several days in the past; this embodiment does not limit this.

[0051] Semantic features, determined based on the domain name to be detected, characterize the contextual features of the domain name's query sequence within a historical time period. Contextual features reflect the contextual information contained within the domain name query sequence. This contextual information can include the consecutive occurrences, proximity occurrences, co-occurrences, and order of occurrence of multiple domain names arranged by query time. For example, during access to the same network service, the main site domain, content distribution domain, resource domain, and interface domain may be queried within the same or similar time periods. Arranging the corresponding query records by query time, the order of occurrence, proximity occurrences, and co-occurrences of multiple domain names in the domain name query sequence can reflect the corresponding contextual information. Semantic features can represent this contextual information in a form suitable for computer processing.

[0052] Scenario information can be determined based on the statistical characteristics of the domain name to be detected, indicating the probability that the domain name belongs to each of multiple scenarios. For example, the multiple scenarios may include newly registered domain name scenarios, encrypted domain name resolution scenarios, low-activity domain name scenarios, high-activity and stable domain name scenarios, and high-threat zone scenarios. Scenario information may include multiple probability values ​​corresponding to each scenario, with each probability value indicating the probability that the domain name to be detected belongs to the corresponding scenario. Therefore, scenario information can reflect the probability that the domain name to be detected belongs to multiple scenarios, rather than limiting the domain name to being detected to only one scenario.

[0053] Fusion features are features obtained by fusing statistical and semantic features. Statistical and semantic features can participate in the fusion process according to their respective contributions to determining the protection strategy. Fusion features can be used to determine the protection strategy for the domain name to be detected.

[0054] Risk levels can be used to characterize the degree of risk of a domain name to be detected. For example, risk levels may include low risk, medium risk, and high risk. The number of risk levels and the criteria for classifying each risk level can be determined according to actual application needs, and this application embodiment does not limit this.

[0055] Security types can be used to indicate whether a domain name under test is judged to be normal, suspicious, or related to a corresponding attack activity. For example, security types may include phishing / spoofing, command and control, algorithm-generated domain names, malicious downloads or propagation, mining communication, DNS tunneling or data infiltration, suspicious, and normal. The suspicious type can indicate that the domain name under test is abnormal, but cannot yet be determined as a specific attack type; the normal type can indicate that the domain name under test is judged to be a normal domain, and this application embodiment does not limit this.

[0056] A protection strategy can be a handling plan determined for the domain name to be detected. A protection strategy can include a single protection action or a combination of multiple protection actions. For example, protection actions can include allowing access, logging, alerting, manual review, blocking, redirecting, diverting traffic, or adding to a blacklist, etc., and this application embodiment does not limit this.

[0057] In the embodiments of this application, the data or intermediate processing results involved in the processing of statistical features, semantic features, fusion features, and other features can be represented in the form of vectors, tensors, or other data formats suitable for computer processing. The above data formats are merely exemplary representations of the relevant features or intermediate processing results and do not constitute a limitation on the embodiments of this application.

[0058] To facilitate understanding of the practical application environment of the embodiments of this application, the following is combined with Figure 1The domain name protection system provided in the embodiments of this application will be described.

[0059] like Figure 1 As shown, the domain name protection system may include a protection policy determination device 100, a domain name resolution server 200, an associated data source 300, and a domain name query device 400. The domain name query device 400 can send a domain name query request to the domain name resolution server 200, and the domain name resolution server 200 can provide the queried domain name in the domain name query request as the domain name to be detected to the protection policy determination device 100. The protection policy determination device 100 can send a data query request to the associated data source 300 based on the domain name to be detected, and the associated data source 300 can provide the protection policy determination device 100 with associated data related to the domain name to be detected. The protection policy determination device 100 can process the domain name to be detected and the associated data to obtain a protection policy for the domain name to be detected. Figure 1 In the embodiment shown, the protection policy determination device 100 can also provide the protection policy to the domain name resolution server 200.

[0060] The protection strategy determination device 100 can be used to acquire the domain name to be detected and its associated data, and determine a protection strategy for the domain name to be detected based on the domain name to be detected and its associated data. The protection strategy determination device 100 can be implemented by a server, computing device, security management platform, or other device with data processing capabilities, or by a functional module deployed in a corresponding device. This application embodiment does not limit the specific implementation form of the protection strategy determination device 100.

[0061] Domain name resolution server 200 can be used to receive domain name query requests and obtain the queried domain name from the query request. Figure 1 In the illustrated embodiment, the domain name resolution server 200 can provide the queried domain name as the domain name to be detected to the protection policy determination device 100. The domain name resolution server 200 can be a recursive domain name resolution server or other devices capable of providing domain name resolution services.

[0062] In one possible implementation, the domain name resolution server 200 may also receive protection policies provided by the protection policy determination device 100 and process domain name query requests accordingly based on the protection policies. For example, the domain name resolution server 200 may perform actions such as allowing the request, recording the domain name query request or processing result, returning a response indicating that the domain name does not exist, returning a response indicating that there is no data, redirecting the request, or returning a preset address, based on the protection policies. The domain name resolution server 200 executing the protection policies is only one possible implementation; the protection policies may also be executed by the protection policy determination device 100 or other devices with policy execution capabilities.

[0063] For example, the protection policy determination device 100 can provide the protection policy to the domain name resolution server 200 or other devices with policy execution capabilities by means of responding to policy area file synchronization, dynamic injection through recursive domain name system management interface or loading of domain name system proxy plugin.

[0064] The associated data source 300 can be used to store or provide associated data of the domain name to be detected. The associated data source 300 can be implemented by a database, data platform, server, data acquisition device, or other device capable of storing or providing associated data. The protection strategy determination device 100 can use the domain name to be detected, the domain name identifier corresponding to the domain name to be detected, or other information that can indicate the domain name to be detected as query, filtering, or matching conditions to obtain associated data related to the domain name to be detected from the associated data source 300.

[0065] It should be noted that the associated data is data already formed during the data processing related to the domain name to be detected. The protection policy determination device 100 can query, filter, match, or aggregate the already formed data based on a data query request to obtain the associated data of the domain name to be detected. The above acquisition process does not mean that the associated data needs to be temporarily generated in response to this data query request, nor does it mean that the domain name to be detected and the associated data need to be pre-combined into a complete set of data before being provided to the protection policy determination device 100.

[0066] like Figure 1 As shown, the associated data source 300 may include a domain name resolution log data source 301, a domain name registration information data source 302, a certificate information data source 303, a threat intelligence data source 304, and an encrypted domain name resolution traffic data source 305, etc. These data sources can be implemented by log storage systems, databases, data query platforms, or data collection devices, etc. Specifically, the domain name resolution log data source 301 can provide data generated during domain name queries or resolution; the domain name registration information data source 302 can provide data generated during domain name registration and management; the certificate information data source 303 can provide data generated during the application, issuance, verification, or use of digital certificates; the threat intelligence data source 304 can provide data generated during threat intelligence collection, analysis, or updating; and the encrypted domain name resolution traffic data source 305 can provide traffic data generated during encrypted domain name resolution communication. The above data sources are merely examples; the type and number of data sources included in the associated data source 300 can be determined according to actual application needs, and this application embodiment does not limit this.

[0067] The domain name query device 400 can be used to send domain name query requests to the domain name resolution server 200. The domain name query device 400 can be a terminal device, server, network gateway, or other device capable of initiating domain name query requests. The domain name query request may include the domain name to be queried, the query type, and other information used for domain name resolution.

[0068] It should be noted that, Figure 1 The system composition and data interaction relationships shown are only one possible implementation. In other possible implementations, the domain name to be detected may also be obtained from domain name query logs, network access records, security alarms, task queues, or a set of domain names to be processed. The protection policy determination device 100 may also simply output or store the protection policy, or provide the protection policy to other devices with policy execution capabilities. The embodiments of this application do not limit the specific source of the domain name to be detected or the specific entity executing the protection policy.

[0069] To further explain Figure 1 The process of determining the protection strategy by the protection strategy determination device 100 is described below in conjunction with... Figure 2 Please provide an explanation. Figure 2 This is a schematic diagram illustrating the processing relationship for determining a protection strategy based on statistical and semantic features, as provided in an embodiment of this application.

[0070] like Figure 2 As shown, in one possible implementation, the protection strategy determination device 100 can acquire the domain name to be detected and its associated data, determine statistical features based on the domain name to be detected and the associated data, and determine semantic features based on the domain name to be detected. The statistical features can be used to characterize the statistical information of the domain name to be detected under different data dimensions, and the semantic features can be used to characterize the contextual features of the domain name to be detected based on the domain name query sequence within a historical time period. The protection strategy determination device 100 can further determine a protection strategy for the domain name to be detected based on the statistical features and semantic features.

[0071] It should be noted that, Figure 2 This is used to illustrate the processing relationship between the domain name to be detected, associated data, statistical features, semantic features, and protection strategies. Figure 2 The connections in the table represent the basis for determining the corresponding processing results. They do not limit the specific implementation method used to obtain the processing results based on the corresponding basis, nor do they limit the specific methods by which statistical and semantic features participate in the determination of protection strategies. Statistical and semantic features can be used directly as the basis for determining protection strategies, or they can be used to determine protection strategies after further processing.

[0072] The connection between the domain name to be detected and its semantic features indicates that the semantic features are determined based on the domain name to be detected, but does not mean that the semantic features are generated solely from the string content of the domain name itself. In one possible implementation, the protection strategy determination device 100 can input the domain name to be detected into a pre-trained domain name semantic model and obtain the semantic features of the domain name to be detected output by the domain name semantic model. In other possible implementations, the protection strategy determination device 100 can also obtain the semantic features of the domain name to be detected based on a pre-established mapping relationship between domain names and semantic features, contextual statistical results based on historical domain name query sequences, or other processing criteria.

[0073] In practical applications, domain protection strategies are typically determined through methods such as blacklist or reputation database matching, domain string feature analysis, and domain resolution behavior analysis. These methods primarily rely on information directly matchable or statistically relevant to the domain being tested and its associated data, making it difficult to fully utilize the contextual information presented by multiple domains arranged according to query time. Therefore, for domains with different query contexts, the distinguishing ability of these methods is limited, easily leading to a mismatch between the determined protection strategy and the actual risk of the domain being tested, resulting in either insufficient or excessive protection.

[0074] To address the aforementioned issues, in the embodiments of this application, Figure 1 The protection strategy determination device 100 in the domain name protection system shown can determine the protection strategy according to... Figure 2 The processing relationship shown is used to obtain the domain name to be detected and its associated data. Statistical features are determined based on the domain name to be detected and its associated data, and semantic features are also determined based on the domain name to be detected. The statistical features can be used to characterize the statistical information of the domain name to be detected under different data dimensions, and the semantic features can be used to characterize the contextual features of the domain name to be detected based on the domain name query sequence within a historical time period. The protection strategy determination device 100 can further determine a protection strategy for the domain name to be detected based on the statistical features and semantic features.

[0075] Therefore, the protection strategy determination device 100 can simultaneously utilize statistical features determined based on the domain name to be detected and its associated data, as well as semantic features determined based on the domain name to be detected. This means that the determination of the protection strategy no longer relies solely on directly matchable or statistically relevant information from the domain name itself or its associated data, but also utilizes the contextual information presented by historical domain name query sequences. This provides a more comprehensive feature basis for accurately determining the protection strategy. Furthermore, semantic features can centrally reflect the contextual information formed by domain name query sequences within a historical time period, enhancing the ability of the protection strategy determination process to distinguish contextual differences between different domain names to be detected, thereby improving the accuracy of the protection strategy determination.

[0076] based on Figure 1 The application environment shown and Figure 2 The data processing relationships shown below, combined with... Figure 3 The method for determining domain name protection strategies provided in the embodiments of this application will be described. Figure 3 This is a flowchart illustrating a method for determining a domain name protection strategy, as provided in an embodiment of this application.

[0077] like Figure 3 As shown, the method may include steps S301 to S303.

[0078] S301: The protection strategy determination device 100 acquires the domain name to be detected and the associated data of the domain name to be detected.

[0079] In one possible implementation, the protection strategy determination device 100 can first obtain the domain name to be detected, and then use the domain name to be detected as a query condition to query, match or aggregate data related to the domain name to be detected from the pre-formed data to obtain the associated data of the domain name to be detected.

[0080] The domain to be detected can be the domain indicated by a domain query request, historical domain query logs, security alerts, pending tasks, or a set of domains to be processed. For example, the domain to be detected can be the domain indicated by the current domain query request, or it can be a domain in the historical domain query log that meets the preset detection conditions.

[0081] The associated data is generated during the data processing related to the domain name to be detected, and can be generated and stored before the protection policy determination device 100 obtains the domain name to be detected. For example, the associated data may be one or more of the following: query records and resolution results of the domain name to be detected, registration information of the domain name to be detected or its corresponding registrable domain name, digital certificate information covering the domain name to be detected, threat intelligence related to the domain name to be detected or its corresponding response Internet Protocol address, and encrypted domain name resolution traffic data related to the domain name to be detected.

[0082] It should be noted that the associated data for different domain names to be tested may have different data types and levels of completeness. This application does not limit each domain name to have all types of associated data. For example, some domain names to be tested may have relatively complete registration information, but lack sufficient historical DNS records; queries for some domain names to be tested may mainly use encrypted DNS resolution methods, thus lacking corresponding plaintext DNS resolution data.

[0083] In one possible implementation, the protection strategy determination device 100 may further preprocess at least one of the domain name to be detected and associated data. Exemplarily, preprocessing may include format conversion, deduplication, anomaly handling, missing data handling, normalization, standardization, or encoding.

[0084] As one implementation of missing data processing, when some related data is missing, the protection strategy determination device 100 can set a missing data marker, a validity mask, or a preset padding value for the missing data, or it can perform subsequent processing using the currently obtained related data. Specifically, the missing data marker can be used to indicate whether the corresponding data is missing, the validity mask can be used to control whether the corresponding data participates in subsequent processing, and the preset padding value can be used to fill in the missing positions in the corresponding data.

[0085] S302: The protection strategy determination device 100 determines the statistical and semantic features of the domain name to be detected.

[0086] The protection strategy determination device 100 can determine statistical features based on the domain name to be detected and associated data, and determine semantic features based on the domain name to be detected. The statistical features and semantic features can be represented using vectors, tensors, or other data formats suitable for computer processing. The protection strategy determination device 100 can determine the statistical features and semantic features sequentially or in parallel; this embodiment does not limit this approach.

[0087] The following is combined with Figure 4 This section describes one possible way to determine statistical characteristics.

[0088] Figure 4 This is a schematic diagram of a statistical feature generation process provided in an embodiment of this application, used to illustrate a possible way in which the protection strategy determination device 100 determines statistical features.

[0089] like Figure 4 As shown, in one possible implementation, the protection strategy determination device 100 can determine multiple statistical sub-features based on the domain name to be detected and associated data, and generate statistical features of the domain name to be detected based on the multiple statistical sub-features.

[0090] It should be noted that, Figure 4 The device 100 for illustrating protection strategy determination can determine at least two of the statistical sub-features shown in the figure, and generate statistical features of the domain name to be detected based on the determined statistical sub-features, but does not mean that the statistical features are generated based on all the statistical sub-features shown in the figure. Figure 4The connection relationship in the text is used to represent a processing relationship between the domain name to be detected, associated data, statistical sub-features and statistical features, without limiting the specific processing method adopted by the protection strategy determination device 100 when generating statistical features based on multiple statistical sub-features.

[0091] In one possible implementation, when the domain name to be detected or associated data is preprocessed in step S301, the protection strategy determination device 100 can determine a variety of statistical sub-features based on the preprocessed domain name to be detected and associated data.

[0092] For example, the various statistical sub-features may include at least two of the following: domain name character features, domain name resolution features, domain name registration information features, digital certificate features, threat intelligence features, and encrypted domain name resolution traffic features. Except... Figure 4 In addition to the statistical sub-features shown, the protection strategy determination device 100 can also determine other statistical sub-features based on the domain name to be detected and associated data. These other statistical sub-features can be used to characterize the statistical information of the domain name to be detected in other data dimensions.

[0093] Domain name character features can be determined based on the domain name string of the domain name to be detected, and are used to characterize the character composition or string structure of the domain name to be detected. For example, domain name character features may include the total length of the domain name, the length of registrable domain names, the length of subdomains, the number of subdomain levels, the proportion of numeric characters, the proportion of hyphens, the number of character types, character entropy, the rarity of character fragments, or homonymous character markers, etc.

[0094] Character entropy can be used to reflect the dispersion of character distribution in the domain name to be detected. The protection strategy determination device 100 can calculate character entropy based on the occurrence frequency of each character in the domain name to be detected. Character fragment rarity can be determined based on the occurrence frequency of consecutive character combinations contained in the domain name to be detected. In one possible implementation, the protection strategy determination device 100 can pre-calculate the occurrence frequency of different character combinations in multiple domain names. When determining the character fragment rarity of the domain name to be detected, the protection strategy determination device 100 can extract character combinations consisting of two or three consecutive characters from the domain name string of the domain name to be detected, and determine the character fragment rarity based on the occurrence frequency of the corresponding character combinations. Homographs can be used to indicate whether there are characters with similar shapes but different character codes in the domain name to be detected.

[0095] Domain name resolution features can be determined based on the query records and resolution results of the domain name to be detected, and are used to characterize the behavior of the domain name to be detected during the domain name query or resolution process. For example, domain name resolution features may include the number of queries within a set time window, the number of duplicate queries from different clients, the query burst coefficient, the query time interval statistics, the query type distribution, the proportion of non-existent domain name responses, the proportion of responses with no data, the lifetime statistics, the proportion of records with low lifetimes, the number of duplicate response Internet Protocol addresses, the number of response Internet Protocol address switches, the depth of the canonical name chain, or the concentration of autonomous systems to which the response Internet Protocol addresses belong, etc.

[0096] The query burst coefficient can be used to reflect the degree of change in the query volume of the domain name to be detected within the current time window relative to the historical query volume benchmark. For example, the protection strategy determination device 100 can determine the query burst coefficient as the ratio of the query volume within the current time window to the median of the query volume within the same time period over the past several days. In other possible implementations, the protection strategy determination device 100 can also determine the query volume benchmark based on the mean, quantiles, or other statistical results of historical query volumes.

[0097] Domain registration information features can be determined based on the registration information of the domain to be tested or the registration information of the corresponding registrable domain, and are used to characterize the registration and management status of the domain to be tested. For example, domain registration information features may include domain age, remaining time until expiration date, abnormal registration status markers, privacy protection enabled markers, number of name servers, number of name server changes, registration information completeness, or newly registered domain markers, etc.

[0098] The domain age can be determined based on the time interval between the domain registration time and the current time, the remaining time until the expiration date can be determined based on the time interval between the domain expiration time and the current time, and the newly registered domain tag can be used to indicate whether the domain age is less than the preset duration.

[0099] Digital certificate characteristics can be determined based on digital certificate information matching the domain name to be detected, and are used to characterize the certificate usage of the domain name to be detected. In one possible implementation, when the certificate subject name of the digital certificate is consistent with the domain name to be detected, the alternative subject name of the digital certificate includes the domain name to be detected, or the wildcard range of the digital certificate can match the domain name to be detected, the protection policy determination device 100 can determine the digital certificate as a digital certificate matching the domain name to be detected. For example, digital certificate characteristics may include certificate validity period, remaining certificate validity time, number of certificate fingerprint reuses, number of domain names in the alternative subject name, abnormal flags for the alternative subject name, certificate chain verification results, certificate issuing authority credibility, or the interval between the certificate's first discovery time and the domain name registration time, etc.

[0100] Threat intelligence features can be determined based on threat intelligence related to the domain name to be detected or the corresponding Internet Protocol address (IPA address), and are used to characterize the matching between the domain name to be detected and known risk information. For example, threat intelligence features may include blacklist hit rate, whitelist hit flag, intelligence source credibility statistics, number of associated malicious intrusion indicators, number of associated malicious infrastructures, or the latest intelligence update time, etc.

[0101] Encrypted domain name resolution traffic characteristics can be determined based on encrypted domain name resolution traffic data related to the domain name to be detected, and are used to characterize the traffic behavior formed when the domain name to be detected is queried using encrypted domain name resolution. For example, encrypted domain name resolution traffic characteristics may include uplink message length statistics, downlink message length statistics, message direction sequence, message time interval statistics, connection duration, number of connections per unit time, or data transmission volume, etc. When the query of the domain name to be detected mainly uses encrypted domain name resolution, resulting in some plaintext domain name resolution data being unavailable or incomplete, the protection strategy determination device 100 can use the aforementioned encrypted domain name resolution traffic data to determine the corresponding statistical sub-features, so that the traffic data formed during the encrypted domain name resolution process can be used to determine the statistical characteristics of the domain name to be detected.

[0102] It should be noted that, Figure 4 The various statistical sub-features shown can be determined based on different types of data, and these sub-features can differ in their value range, dimension, or representation. In one possible implementation, the protection strategy determination device 100 can process each statistical sub-feature separately according to its value range, dimension, or representation, so that the processed statistical sub-features can be used together to generate statistical features.

[0103] For example, for numerical statistical sub-features with different value ranges, the protection strategy determination device 100 can perform min-max normalization or standardization processing to adjust the numerical scale of the corresponding statistical sub-features; for statistical sub-features including category labels, the protection strategy determination device 100 can perform one-hot encoding, embedded encoding, or label value processing. The protection strategy determination device 100 can select the appropriate processing method according to the value range, dimension, or representation of each statistical sub-feature, without needing to perform all of the above processing on each statistical sub-feature.

[0104] The protection strategy determination device 100 can arrange or combine the processed statistical sub-features according to a preset method, or it can use a model to process the processed statistical sub-features to obtain the statistical features of the domain name to be detected. This embodiment of the application is not limited in this respect. In one possible implementation, when the processed statistical sub-features can all be represented in a multidimensional numerical form, the protection strategy determination device 100 can concatenate the corresponding processing results in a preset order to obtain statistical features represented by feature vectors.

[0105] It should be noted that the normalization, standardization, encoding, and concatenation processes described above are merely specific implementations for generating statistical features based on multiple statistical sub-features, and do not limit the scope of these processes. Figure 4 The specific processing method in the processing relationship is shown. In other possible implementations, the protection strategy determination device 100 may also generate statistical features of the domain name to be detected based on multiple statistical sub-features through feature selection models, feature transformation models, neural networks or other feature processing methods.

[0106] pass Figure 4 The processing method shown allows the protection strategy determination device 100 to determine at least two statistical sub-features based on the domain name to be detected and associated data, and to generate statistical features based on the determined statistical sub-features, so that the statistical features can comprehensively characterize the statistical information of the domain name to be detected under different data dimensions.

[0107] The following is combined with Figure 5 This section describes one possible way to determine semantic features.

[0108] Figure 5 This is a schematic diagram illustrating a domain name semantic model training and semantic feature determination process provided in an embodiment of this application. Figure 5 As shown, the training process of the domain name semantic model and the process of determining semantic features using the trained domain name semantic model can be performed separately. The domain name semantic model can be trained before determining the semantic features of the domain name to be detected.

[0109] The training process of the domain name semantic model can be performed by the protection strategy determination device 100, the model training device, or other devices with model training capabilities. This application embodiment does not limit this. For ease of explanation, the following description uses the model training device performing the domain name semantic model training process as an example.

[0110] like Figure 5As shown, in one possible implementation, during the model training phase, the model training device can first acquire multiple domain name query records formed within a historical time period. Each domain name query record can include the query time and the queried domain name. These multiple domain name query records can be obtained from domain name resolution servers, recursive domain name system nodes, or domain name query logs recorded by other devices. The historical time period can be set according to the training data scale, model update cycle, or deployment requirements; for example, the historical time period can be the past several days, several months, or other preset time ranges.

[0111] Next, the model training device can extract the query time and the queried domain name from multiple domain name query records, and arrange the multiple queried domain names according to the query time to construct one or more domain name query sequences. Each domain name query sequence can include multiple domain names arranged according to the query time. The consecutive occurrence, adjacent occurrence, co-occurrence, and order of the multiple domain names in the domain name query sequence can reflect the contextual information of multiple domain names in the historical query process.

[0112] For example, the model training device can group multiple domain name query records based on group identifiers included in the multiple domain name query records, and then arrange the multiple queried domain names in each group according to the query time to form a domain name query sequence corresponding to each group. The group identifiers may include client identifiers, terminal identifiers, user identifiers, enterprise identifiers, tenant identifiers, network segment identifiers, session identifiers, or recursive domain name system node identifiers, etc.

[0113] Then, the model training device can select domain names from the domain query sequence to construct training samples according to preset context selection conditions. Specifically, the model training device can use one domain name from the domain query sequence as the center domain name, and one or more domain names that satisfy the context selection conditions with the center domain name as associated domain names. The center domain name and each associated domain name are then used to form training samples, which can include one center domain name and one associated domain name. The model training device can use multiple domain names from the domain query sequence as center domain names and construct multiple training samples in the above manner.

[0114] It should be noted that associated domains can be domains that meet the context selection criteria with the central domain in the domain query sequence, which is different from the associated data used as a data concept mentioned above. Context selection criteria can be determined based on the sequence position interval between the central domain and associated domains, query time interval, time window, query segment, session affiliation, frequency of co-occurrence, relevance, or other conditions that reflect contextual information. The model training device can use one or more of the above criteria to select associated domains.

[0115] For example, the model training device can divide the domain name query sequence into multiple query segments according to a preset time window, and select a domain name from one query segment as the central domain name. The model training device can determine one or more domain names located within a preset selection radius before and after the central domain name as associated domain names, and form training samples based on the central domain name and each associated domain name. The above-mentioned query segments and preset selection radius are only one specific way to determine the central domain name and associated domain names, and the embodiments of this application do not limit the specific form of the context selection conditions.

[0116] like Figure 5 As shown, the model training device can use multiple training samples to train the domain name semantic model, enabling the domain name semantic model to learn the contextual information reflected in the historical domain name query sequence of the central domain name and related domain names.

[0117] After obtaining the trained domain name semantic model, such as Figure 5 As shown, in the semantic feature determination stage, the protection strategy determination device 100 can use the trained domain name semantic model to process the domain name to be detected and obtain the semantic features of the domain name to be detected. Since the domain name semantic model is trained based on the above-mentioned multiple training samples, the semantic features can reflect the contextual information learned by the domain name semantic model from the above-mentioned multiple training samples, and thus be used to characterize the contextual features of the domain name to be detected based on the domain name query sequence within the historical time period.

[0118] In one possible implementation, the model training device can train a domain name semantic model using a skip-word model architecture. The model training device can construct a domain name vocabulary set based on different domain names appearing in multiple domain name query sequences. The domain name semantic model can include an input embedding matrix and an output embedding matrix. The number of rows in the input and output embedding matrices can be the same as the number of domain names included in the domain name vocabulary set, and the number of columns in the input and output embedding matrices can be a preset embedding dimension.

[0119] For including the central domain name and associated domains The training samples and model training equipment can utilize the central domain name. Predict related domains And adjust the model parameters of the domain name semantic model to improve performance on a given central domain name. Associated domains under the condition The probability of occurrence. For example, this conditional probability can be expressed as:

[0120] ;

[0121] in, Indicates that in a given central domain name Associated domains under the condition The probability of occurrence Represents the input embedding matrix. This indicates the output embedding matrix. This indicates the relationship between the input embedding matrix and the center domain name. The corresponding vector, This indicates the relationship between the output embedding matrix and the associated domain name. The corresponding vector, Indicates transpose. Represents a set of domain name terms. Represents a set of domain name terms One of the domain names, This indicates exponentiation.

[0122] Model training equipment can improve performance on a given central domain. Associated domains under the condition The probability of occurrence is used as the training objective. The training loss is determined based on multiple training samples, and the input embedding matrix and output embedding matrix are adjusted according to the training loss.

[0123] During training, the model training device can use full softmax to calculate the normalized probability of all domain names in the domain name vocabulary set; the model training device can also use negative sampling, hierarchical softmax, or other methods to reduce the amount of training computation.

[0124] Furthermore, when the domain name semantic model adopts the aforementioned skip character model, and the domain name to be detected has been included in the domain name vocabulary set, the protection strategy determination device 100 can read the row vector corresponding to the domain name to be detected from the input embedding matrix and determine the semantic features of the domain name to be detected based on the row vector. In other possible embodiments, the protection strategy determination device 100 can also determine the semantic features of the domain name to be detected based on the row vector corresponding to the domain name to be detected in the output embedding matrix, or based on a combination of the row vectors corresponding to the domain name to be detected in the input embedding matrix and the output embedding matrix.

[0125] It should be noted that historical domain query sequences are primarily used to train the domain semantic model. During the semantic feature determination phase, the complete historical domain query sequence, as well as other domains preceding and following the domain to be detected in the current query, are not necessarily required data for determining semantic features. The protection strategy determination device 100 can determine semantic features based on the domain to be detected, and the contextual features represented by these semantic features can originate from the model parameters learned by the domain semantic model from the historical domain query sequences during the training phase.

[0126] For domain names to be detected that are not included in the domain name vocabulary set, the protection strategy determination device 100 can determine semantic features in a way that is adapted to the domain name semantic model. For example, the domain name semantic model can process the domain name to be detected using character-level, word-level, or subdomain-level encoding methods, or it can determine the semantic features of the domain name to be detected after completing incremental training or model updates. This application embodiment does not limit this.

[0127] In other possible implementations, the domain name semantic model can also employ a transformer model. For example, when computational resources are sufficient, the model training device can use a transformer-based masked language model (MLM) pre-training method, or use a simple framework for contrastive learning of visual representations (SimCLR) for contrastive learning to train the domain name semantic model. This application does not limit the specific type of domain name semantic model.

[0128] Through the above processing, the model training device can utilize multiple training samples, including the central domain name and related domain names, to enable the domain name semantic model to learn the contextual information reflected by the central domain name and related domain names in historical domain name query sequences. When the protection strategy determination device 100 processes the domain name to be detected using the trained domain name semantic model, it can obtain semantic features that reflect the contextual features of the domain name to be detected based on historical domain name query sequences, thereby improving the accuracy of the semantic features in representing the contextual features of the domain name to be detected.

[0129] S303: The protection strategy determination device 100 determines the protection strategy for the domain name to be detected based on statistical and semantic features.

[0130] This application does not limit the specific method of determining the protection strategy based on statistical features and semantic features. For example, the protection strategy determination device 100 can directly determine the protection strategy based on statistical features and semantic features, or it can first determine the contribution of statistical features and semantic features in the process of determining the protection strategy, and then determine the protection strategy based on the corresponding contribution.

[0131] In one possible implementation, the protection strategy determination device 100 can determine a first contribution level of statistical features and a second contribution level of semantic features, and determine a protection strategy for the domain name to be detected based on the statistical features, semantic features, the first contribution level, and the second contribution level. The first contribution level characterizes the contribution of statistical features in determining the protection strategy, and the second contribution level characterizes the contribution of semantic features. Therefore, the protection strategy determination device 100 can utilize both types of features—statistical features and semantic features—based on their respective roles in the protection strategy determination process, avoiding the determination of a protection strategy without distinguishing the differences in the roles of the two types of features, thereby improving the comprehensive utilization of statistical and contextual information.

[0132] The first and second contribution levels can be represented by weight parameters, attention coefficients, scaling coefficients, or other parameters that characterize the participation level of the corresponding features. The protection strategy determination device 100 can use the first and second contribution levels to adjust the influence of statistical and semantic features on subsequent processing results. When a unified feature representation is required, the protection strategy determination device 100 can fuse statistical and semantic features to obtain fused features; when a unified feature representation is not required, the protection strategy determination device 100 can directly determine the protection strategy using statistical features, semantic features, the first contribution level, and the second contribution level.

[0133] Furthermore, in the process of determining the protection strategy based on statistical and semantic features, the protection strategy determination device 100 can first determine the strategy matching information of the domain name to be detected, and then determine the protection strategy for the domain name to be detected based on the strategy matching information. The strategy matching information may include one or more pieces of information used to match the protection strategy. For example, the strategy matching information may include at least one of the risk level and security type of the domain name to be detected. The strategy matching information can be determined based on statistical and semantic features, or based on the result obtained after processing the two types of features by combining the first contribution degree and the second contribution degree.

[0134] In one possible implementation, the protection strategy determination device 100 can determine the scenario information of the domain name to be detected based on statistical characteristics, determine a first contribution level and a second contribution level based on the statistical characteristics, semantic characteristics, and scenario information, and determine a protection strategy for the domain name to be detected by combining the corresponding contribution levels. The scenario information may include multiple probability values ​​corresponding to multiple scenarios, representing the probability distribution of the domain name to be detected belonging to multiple scenarios, without limiting the use to only the largest probability value. Multiple probability values ​​can jointly participate in determining the first and second contribution levels, allowing the statistical characteristics and semantic characteristics to participate in the determination of the protection strategy with contribution levels adapted to the scenario probability distribution, thereby improving the adaptability of the protection strategy determination process to different domain name detection scenarios.

[0135] It should be noted that the scene information is used to participate in the determination of the first and second contribution levels, in order to adjust the contribution levels of statistical features and semantic features in the process of determining the protection strategy, rather than determining the contribution level of scene information alone. Furthermore, determining scene information and combining it with the scene information to determine the first and second contribution levels is only one possible implementation of step S303, and is not a necessary step in determining the protection strategy based on statistical features and semantic features.

[0136] The following combination Figure 6 This paper describes a computable implementation for determining and fusing the first and second contribution levels using scenario information. The first weight and the second weight are specific representations of the first and second contribution levels, respectively, and the fused features can serve as an intermediate representation for determining the protection strategy for the domain name to be detected.

[0137] See Figure 6 , Figure 6 This is a schematic diagram illustrating scene information, contribution level, and fusion processing procedure provided in an embodiment of this application. For example... Figure 6 As shown, the protection strategy determination device 100 can determine scene information based on statistical features, the scene information including multiple probability values ​​corresponding to multiple scenes respectively; project the statistical features and semantic features respectively to obtain a first projection result and a second projection result; determine a first weight and a second weight based on the first projection result, the second projection result and the scene information; and then perform a fusion process on the first projection result and the second projection result based on the first weight and the second weight to obtain a fused feature.

[0138] Figure 6Taking the scenario evaluation model as an example of processing statistical features and outputting scenario information, one method of determining scenario information is shown. The scenario evaluation model is only one optional implementation for determining scenario information. In other possible implementations, the protection strategy determination device 100 can match the statistical features with preset scenario rules corresponding to multiple scenarios to obtain multiple matching scores corresponding to multiple scenarios, and determine multiple probability values ​​corresponding to multiple scenarios based on the multiple matching scores.

[0139] In one possible implementation, the multiple scenarios may include one or more of the following: newly registered domain name scenario, encrypted domain name resolution scenario, low-activity domain name scenario, high-activity and stable domain name scenario, and high-threat area scenario.

[0140] The newly registered domain scenario indicates that the domain to be detected was registered relatively recently. For example, a domain age of less than 7 days can be used as an exemplary criterion for determining the newly registered domain scenario. Since there may be relatively little long-term domain name resolution behavior data generated by newly registered domains, the probability corresponding to the newly registered domain scenario can be used to increase the contribution of domain character features, digital certificate features, or threat intelligence features, while decreasing the contribution of long-term domain name resolution behavior features.

[0141] An encrypted domain name resolution scenario indicates that the domain name under test is primarily queried through encrypted domain name resolution methods, resulting in the loss or incompleteness of some plaintext domain name resolution behavior data, but the corresponding encrypted domain name resolution traffic data can still be obtained. The probability corresponding to the encrypted domain name resolution scenario can be used to reduce the contribution of plaintext domain name resolution features and increase the contribution of encrypted domain name resolution traffic features.

[0142] A low-activity domain scenario indicates that the domain being monitored receives fewer queries per unit of time. For example, fewer than 10 queries per hour can be used as an exemplary criterion for a low-activity domain scenario. Since low-activity domains are unlikely to generate sufficient domain name resolution behavior information in a short period of time, the probability corresponding to a low-activity domain scenario can be used to increase the contribution of domain character features, digital certificate features, or threat intelligence features.

[0143] A highly active and stable domain scenario indicates that the domain to be detected has a high number of queries, and the current query behavior fluctuates little compared to the historical query baseline. For example, the protection strategy determination device 100 can determine the probability that the domain to be detected has the attribute of a highly active and stable domain scenario based on the number of queries and the query burst coefficient. The probability corresponding to the highly active and stable domain scenario can be used to improve the contribution of domain name resolution behavior baseline and behavior change information.

[0144] High-threat area scenarios can indicate a high proportion of historically malicious domains in the top-level domain to which the domain to be detected belongs, or a high proportion of historically malicious Internet Protocol (IP) addresses in the autonomous system to which the response IP address of the domain to be detected belongs. The probabilities corresponding to high-threat area scenarios can be used to enhance the contribution of threat intelligence features and relevant regional statistics.

[0145] The aforementioned scenarios can overlap, and the domain name to be detected can simultaneously possess multiple scenario attributes. The number, type, name, judgment criteria, and corresponding contribution adjustment methods of multiple scenarios can be set according to the actual deployment environment, and this application embodiment does not limit this.

[0146] Furthermore, scene information can also be used to adjust the contribution levels of various statistical sub-features in statistical features. The protection strategy determination device 100 can generate initial statistical features based on various statistical sub-features, determine scene information based on the initial statistical features, and adjust the various statistical sub-features using the scene information to obtain adjusted statistical features. For example, the protection strategy determination device 100 can determine a weight matrix based on the scene information and perform weighted processing on various statistical sub-features based on the weight matrix. The protection strategy determination device 100 can further determine a first contribution level and a second contribution level based on the adjusted statistical features, semantic features, and scene information. It should be noted that the above processing is only one optional implementation method, belonging to one application method of scene information within statistical features, and does not indicate that each statistical sub-feature is at the same processing level as the statistical features and semantic features, nor does it affect... Figure 6 The scenario information shown is used to determine the processing method for the overall contribution of statistical features and semantic features.

[0147] like Figure 6 As shown, in one possible implementation, the protection strategy determination device 100 can utilize a scene evaluation model to process statistical features and obtain scene information. The scene evaluation model may include one or more fully connected layers and a normalization processing layer for outputting multiple scene probabilities. For example, the scene information can be represented as:

[0148] ;

[0149] in, Indicates statistical characteristics, This represents the weight matrix of the scenario evaluation model. This represents the bias vector. Represents the normalized exponential function, Represents scene information. Scene information A scene probability distribution can be generated, consisting of multiple probability values, and represented in vector form. Different dimensions in the scene probability distribution correspond to different scenes, and the value of each dimension represents the probability that the domain name to be detected belongs to the corresponding scene. In the case of a function, the sum of multiple probability values ​​can be 1.

[0150] The scenario evaluation model can be trained based on the statistical features of multiple samples and the scenario annotation results corresponding to each sample's statistical features. The scenario annotation results can be generated according to preset scenario rules. For example, the scenario annotation results can be generated based on domain age, the degree of missing plaintext domain name resolution behavior data, query volume per unit time, query burst coefficient, the proportion of historical malicious domain names in the top-level domain, or the proportion of historical malicious Internet Protocol addresses in the autonomous system.

[0151] In one possible implementation, when the scenario evaluation model uses single-label supervised training, if the same sample domain name satisfies multiple scenario labeling rules, the scenario label corresponding to that sample domain name can be determined according to a preset priority. For example, the scenario labels can be determined in the order of encrypted domain name resolution scenario, high-threat area scenario, newly registered domain name scenario, low-activity domain name scenario, and high-activity stable domain name scenario. During online processing, the protection strategy determination device 100 can retain the complete scenario probability vector output by the scenario evaluation model, without limiting itself to using only the probability value corresponding to the scenario with the highest probability. In other possible implementations, the scenario labeling result can also include multiple scenario labels, and the scenario evaluation model can output the probability of the domain name to be detected belonging to each scenario respectively. The probability corresponding to each scenario can be determined separately, without limiting the sum of multiple probability values ​​to 1. This application embodiment does not limit the training method of the scenario evaluation model or the numerical relationship between multiple scenario probabilities.

[0152] exist Figure 6 In the illustrated embodiment, the protection strategy determination device 100 can project statistical features and semantic features onto the same feature space to obtain a first projection result and a second projection result. The dimensions of the first projection result and the second projection result can be the same, so that they can be used together to calculate the first weight and the second weight, and can be fused based on the first weight and the second weight. For example, the first projection result and the second projection result can be represented as follows:

[0153] ;

[0154] ;

[0155] in, Indicates statistical characteristics, Represents semantic features, and These represent the projection matrices used for projecting statistical features and semantic features, respectively. and These represent the corresponding bias vectors. Represents the hyperbolic tangent function. This represents the first projection result. This represents the second projection result.

[0156] It should be noted that, Figure 6 This only illustrates the participation of scene information in the calculation of the first and second weights, and does not limit the specific processing method of the scene information before participating in the weight calculation. In one possible implementation, the protection strategy determination device 100 can encode the scene information before calculating the first and second weights to obtain a scene encoding result. For example, the scene encoding result can be expressed as:

[0157] ;

[0158] in, This represents scene information composed of multiple scene probabilities. Represents the scene encoding weight matrix. This represents the bias vector. Represents the sigmoid activation function. This represents the scene encoding result. The above encoding process can convert scene information into a data representation suitable for use in conjunction with the first projection result and the second projection result in the calculation of the first and second weights.

[0159] like Figure 6 As shown, the protection strategy determination device 100 can calculate a first weight and a second weight based on the first projection result, the second projection result, and scene information. In one possible implementation, the protection strategy determination device 100 can first encode the scene information to obtain a scene encoding result, and then generate a gating vector based on the first projection result, the second projection result, and the scene encoding result. For example, the gating vector can be represented as:

[0160] ;

[0161] in, This represents a vector concatenation operation. Represents the gate weight matrix. This represents the bias vector. This represents the first projection result. This represents the second projection result. This represents the scene encoding result obtained by encoding the scene information. This represents the gating vector. It's important to note that the scene encoding result is used to incorporate scene information into the gating vector calculation; it is not considered a third feature alongside the first and second projection results in the fusion process.

[0162] Furthermore, gating vector It can be used as the first weight. This can be used as a second weight. The first weight is used to represent the first contribution of the statistical features, and the second weight is used to represent the second contribution of the semantic features. The protection strategy determination device 100 can fuse the first projection result and the second projection result according to the first weight and the second weight to obtain the fused feature:

[0163] ;

[0164] in, This represents element-wise multiplication. This indicates the fusion feature.

[0165] When using the first and second weights in the aforementioned vector form, different feature dimensions in the first and second projection results can have different weights, thereby achieving adjustment of the contribution level per feature dimension. Since the calculation of the first and second weights utilizes the complete scene probability distribution, changes in each probability value in the scene probability distribution can be passed to the first and second weights through scene information encoding and gating calculation, thus affecting the contribution level of statistical features and semantic features in one or more feature dimensions.

[0166] In other possible implementations, the first weight and the second weight can also be represented by scalar weights or feature group-level weights. The protection strategy determination device 100 can also fuse the first projection result and the second projection result by weighted summation, splicing followed by feature processing, or other methods. This application embodiment does not limit the specific representation of the first contribution degree and the second contribution degree, nor the specific method of fusion processing.

[0167] pass Figure 6 The processing described above involves the protection strategy determination device 100 first converting statistical features and semantic features into a first projection result and a second projection result of the same dimension. Then, based on the complete probability distribution of the domain name to be detected belonging to multiple scenarios, it calculates a first weight and a second weight, and fuses the first projection result and the second projection result according to the first weight and the second weight. Thus, the above processing can translate the requirement to adjust the contribution level of statistical features and semantic features based on scene information into computable and executable processing steps such as feature projection, scene information processing, weight calculation, and feature fusion.

[0168] Figure 6The fused features obtained through the processing shown can be used as the result of processing statistical features and semantic features by combining the first and second contribution levels, and can be used to determine the policy matching information for the domain name to be detected.

[0169] See Figure 7 , Figure 7 This is a schematic diagram illustrating a protection strategy matching and adjustment process provided in an embodiment of this application. Figure 7 This diagram illustrates the processing relationship between risk level, security type, and protection strategy, as well as the impact of strong rule triggering and whitelist downgrading on the corresponding processing procedures. Specifically, the protection strategy determination device 100 can query the policy mapping relationship based on the risk level and security type to obtain the protection strategy for the domain name to be detected; if the domain name to be detected meets the strong rule triggering conditions, the protection strategy determination device 100 can adjust the risk level to a high-risk level; if the domain name to be detected meets the whitelist downgrading conditions, the protection strategy determination device 100 can adjust the protection strategy to obtain the adjusted protection strategy.

[0170] The protection strategy determination device 100 can determine the risk level and security type of the domain name to be detected based on statistical features and semantic features, respectively, and can also determine the risk level and security type of the domain name to be detected based on statistical features and semantic features, respectively. Figure 6 The fusion features shown determine the risk level and security type of the domain name to be detected, but this embodiment does not limit this. For ease of explanation, the following description uses the determination of risk level and security type based on fusion features as an example.

[0171] In one possible implementation, the protection strategy determination device 100 can utilize a scoring network to process fused features to obtain a risk score for the domain name to be detected. For example, the risk score can be expressed as:

[0172] ;

[0173] in, This represents the weight matrix of the scoring network. This represents the bias vector. Indicates fusion characteristics, Represents the sigmoid activation function. This represents the risk score, which can range from 0 to 1.

[0174] The protection strategy determination device 100 can convert a risk score into a risk level based on one or more risk thresholds. For example, the protection strategy determination device 100 can set a first threshold. Second threshold When the risk score is less than the first threshold, the protection strategy determination device 100 can determine the domain name to be detected as low-risk; when the risk score is greater than or equal to the first threshold and less than the second threshold, the protection strategy determination device 100 can determine the domain name to be detected as medium-risk; when the risk score is greater than or equal to the second threshold, the protection strategy determination device 100 can determine the domain name to be detected as high-risk.

[0175] For example, the first threshold can be 0.40, and the second threshold can be 0.80. The above risk thresholds are merely examples; the number of risk levels, the number of risk thresholds, and the specific values ​​of each risk threshold can be set according to the deployment environment, risk tolerance, or historical assessment results. In other possible implementations, the protection strategy determination device 100 can also directly determine the risk level using a multi-classification model, or determine the risk level according to preset rules, without first determining a continuous risk score.

[0176] On the other hand, the protection strategy determination device 100 can also determine the security type of the domain name to be detected based on the fusion characteristics. For example, the security type may include phishing imitation, command and control, algorithm-generated domain name, malicious download or propagation, mining communication, domain name system tunneling or data infiltration, suspicious, and normal types.

[0177] Among them, phishing and impersonation domains can indicate that the domain being tested is used to impersonate legitimate brands, business portals, or trusted websites, and redirect users to phishing pages; command and control domains can indicate that the domain being tested is used as remote control infrastructure for Trojans, backdoor programs, or other malicious programs; algorithm-generated domains can indicate that the domain being tested is generated by a domain generation algorithm; malicious downloading or propagation domains can indicate that the domain being tested points to the distribution location of malicious software or other malicious payloads; mining communication domains can indicate that the domain being tested is related to cryptocurrency mining pools or mining communication infrastructure; domain name system tunneling or data leakage domains can indicate that the domain being tested is used for covert communication, tunneling, or data leakage through the domain name system; suspicious domains can indicate that the domain being tested has abnormal characteristics but cannot yet be classified into a specific malicious type; and normal domains can indicate that the domain being tested is determined to be a normal domain.

[0178] In one possible implementation, the protection strategy determination device 100 can utilize a security type classification model to process fused features and output multiple probability values ​​corresponding to multiple candidate security types. For example, the security type probability vector can be represented as:

[0179] ;

[0180] in, The weight matrix represents the security type classification model. This represents the bias vector. Indicates fusion characteristics, Represents the sigmoid activation function. This represents a probability vector indicating the security type.

[0181] For example, the security type classification model can be trained using a multi-label classification method, allowing the same sample domain name to have multiple security type labels during the training phase. During online processing, the protection strategy determination device 100 can determine the candidate security type with the highest probability as the security type of the domain name to be detected based on the multiple probability values ​​output by the security type classification model.

[0182] The above-mentioned security type determination method is only one possible implementation. The protection strategy determination device 100 can also set corresponding judgment thresholds for one or more candidate security types, and determine the security type of the domain name to be detected based on the comparison results of multiple probability values ​​and corresponding judgment thresholds. When multiple candidate security types meet the corresponding judgment conditions, the protection strategy determination device 100 can determine a security type according to the preset selection rules, or it can retain multiple security types.

[0183] exist Figure 7 In the illustrated implementation, the policy mapping relationship is used to represent the correspondence between different combinations of risk levels and security types and protection policies. The protection policy determination device 100 can use both the risk level and security type of the domain name to be detected as matching conditions, and determine the protection policy corresponding to the combination of the risk level and security type according to the policy mapping relationship.

[0184] For example, the combination of medium-risk level and phishing / spoofing type can correspond to a protection strategy that redirects the domain name to be detected to an alert page for manual review; the combination of high-risk level and phishing / spoofing type can correspond to a protection strategy that redirects the domain name to be detected to an alert page and blocks the corresponding original Internet Protocol address.

[0185] The combination of high-risk level and command control can correspond to a protection strategy that returns a response indicating a non-existent domain name and adds the corresponding Internet Protocol address to the outbound blacklist; the combination of high-risk level and algorithm-generated domain name can correspond to a protection strategy that returns a response indicating a non-existent domain name; the combination of high-risk level and mining communication can correspond to a protection strategy that returns a response indicating a non-existent domain name or returns a locally aggregated Internet Protocol address.

[0186] A combination of high-risk levels and DNS tunneling or data infiltration methods can correspond to a protection strategy that resolves the domain name to be detected to a preset honeypot address and records the corresponding communication data. The preset honeypot address can be a server address controlled by the security administrator, used to redirect access to the domain name to be detected to a controlled environment, thereby isolating or recording the corresponding communication behavior.

[0187] The combination of low-risk level and normal category can correspond to the protection strategy of release or release and record assessment results; the combination of medium-risk level and suspicious category can correspond to the protection strategy of observation, alarm or manual review.

[0188] The above mapping is for illustrative purposes only and does not limit the combination of each risk level and security type to correspond only to the above protection strategies, nor does it require the corresponding protection strategy to include all of the above protection actions. In actual deployment, the policy mapping relationship can be set or adjusted according to the network architecture, the importance of the business, the level of risk tolerance, or the capabilities of the protection equipment.

[0189] Furthermore, in Figure 7 In the illustrated embodiment, before querying the policy mapping relationship based on risk level and security type, the protection policy determination device 100 can obtain strong rule judgment data for the domain name to be detected and match or compare the strong rule judgment data with preset strong rules. If the domain name to be detected meets the strong rule triggering conditions, the protection policy determination device 100 can set the risk level of the domain name to be detected to a high-risk level, and then obtain the protection policy based on the high-risk level and the security type of the domain name to be detected. Therefore, the above process allows domain names to be detected with preset high-risk evidence to be matched with protection policies according to high-risk levels, avoiding insufficient protection strength due to a low determined risk level.

[0190] For example, strong rule triggering conditions may include at least one of the following conditions: the domain name to be detected is blacklisted in multiple independent threat intelligence sources, and the number of threat intelligence sources hit reaches a preset number; the domain name to be detected matches the domain name in the ban list issued by an authoritative body; the digital certificate information of the domain name to be detected matches the digital certificate information of a known malicious infrastructure; the response Internet Protocol address of the domain name to be detected matches the Internet Protocol address corresponding to a known command and control infrastructure.

[0191] like Figure 7 As shown, after obtaining the protection strategy, the protection strategy determination device 100 can also acquire the whitelist judgment data of the domain name to be detected, and match or compare the whitelist judgment data with the preset whitelist downgrade rules. If the domain name to be detected meets the whitelist downgrade conditions, the protection strategy determination device 100 can adjust the automatic blocking action in the protection strategy to an alarm action or an action to trigger the manual review process, thus obtaining the adjusted protection strategy. If the domain name to be detected does not meet the whitelist downgrade conditions, the protection strategy determination device 100 can retain the protection strategy obtained according to the strategy mapping relationship.

[0192] For example, whitelist downgrade conditions may include at least one of the following: the domain to be tested matches a domain in the enterprise's core business domain list; the domain to be tested matches a domain in the trusted domain list published by an authoritative entity; the domain to be tested has been registered for a preset duration and no malicious records have been detected within the set historical time period.

[0193] It should be noted that both strong rule triggering and whitelist downgrade processing are optional. The protection strategy determination device 100 can execute one of these processes, according to... Figure 7 The order in which the two processes are executed or the above processes are not executed is shown, and the embodiments of this application do not limit this.

[0194] The above describes the process of determining and adjusting the protection strategy during the online processing phase. The training process of the relevant models will be described below.

[0195] For example, in Figure 6 In the implementation method using the gating mechanism shown, the network that calculates the first and second weights based on scene information and fuses statistical and semantic features according to the first and second weights can be a gated fusion network. The gated fusion network, the scoring network for outputting risk scores, and the safety type classification model can be trained using an end-to-end joint training method, and the model parameters of the scene evaluation model can also be adjusted synchronously during the training process.

[0196] Joint training samples can include sample statistical features, sample semantic features, risk labels, safety type annotation results, and scene annotation results. The following explanation uses the example of risk labels as binary classification labels, safety type annotation results including one or more safety type labels, and scene annotation results as single labels.

[0197] During training, the scene evaluation model can output multiple probability values ​​corresponding to multiple scenes based on the sample statistical features, thus obtaining the scene information of the sample domain name. The gating fusion network can project the sample statistical features and sample semantic features onto the same feature space, obtaining the first projection result and the second projection result of the sample; it calculates the first weight and the second weight based on the first projection result, the second projection result, and the scene information, and then fuses the first projection result and the second projection result of the sample based on the first weight and the second weight to obtain the sample fusion feature. The scoring network and the security type classification model can output the risk score of the sample domain name and multiple probability values ​​corresponding to multiple candidate security types based on the sample fusion feature.

[0198] Furthermore, the risk scoring loss can be determined based on the difference between the risk score and the risk label; the safety type classification loss can be determined based on the difference between the probability values ​​of multiple candidate safety types and the safety type labeling results; and the scene classification auxiliary loss can be determined based on the difference between scene information and scene labeling results. These three losses can together form a joint loss. For example, the joint loss function can be expressed as:

[0199] ;

[0200] in, Indicates risk score loss. Indicates the type of loss classified by safety. This represents the scene classification auxiliary loss. The weight coefficients represent the auxiliary loss for scene classification.

[0201] Risk scoring loss can be achieved using binary cross-entropy loss:

[0202] ;

[0203] in, Indicates a risk label. This represents the risk score output by the scoring network.

[0204] The classification loss for security type can be achieved using multi-label binary classification cross-entropy loss:

[0205] ;

[0206] in, Indicates the first Each security type's label value, The output of the security type classification model represents the first... The probability of each security type.

[0207] Scene classification auxiliary loss can be multi-class cross-entropy loss:

[0208] ;

[0209] in, Indicates the first corresponding to the sample domain name Each scene annotation value, The output of the scene evaluation model represents the first... The probability of each scenario.

[0210] During model training, the model parameters of the gating fusion network, scoring network, security type classification model, and scenario assessment model can be adjusted through backpropagation based on the joint loss. Specifically, the risk scoring loss and security type classification loss can jointly backpropagate the model parameters of the gating fusion network, ensuring that the calculation of the first and second weights is constrained by both losses. The scenario classification auxiliary loss constrains the scenario assessment model to maintain its ability to distinguish between different scenarios, allowing scenario information to participate in the calculation of the first and second weights. Thus, the scenario classification auxiliary loss enables scenario information to more accurately represent the probability that a sample domain name belongs to multiple scenarios, while the risk scoring loss and security type classification loss jointly constrain the gating fusion network to calculate the first and second weights in conjunction with scenario information. This ensures that the fusion results of statistical and semantic features meet the processing needs of both risk scoring and security type classification, thereby improving the accuracy of risk level and security type determination, and consequently, the accuracy of protection strategy determination.

[0211] The above joint training method and the method of adjusting model parameters based on joint loss are only examples. During model training, the weights of each loss can be set according to actual needs, and the model parameters participating in joint training can be adjusted using an adaptive moment estimation optimizer or other optimization algorithms.

[0212] It should be noted that after the domain name semantic model is trained, its model parameters can remain unchanged during the joint training process described above; or, the domain name semantic model can be trained together with other models participating in the joint training to synchronously adjust the model parameters of the domain name semantic model and other models.

[0213] The above model training process is used to obtain the model parameters of each model in advance, and is not part of the process that needs to be performed in step S303 each time a protection strategy is determined.

[0214] The above combination Figures 1 to 7 The method for determining domain name protection strategies provided in this application embodiment will be introduced. Next, the structure of the domain name protection strategy determination device and computing device provided in this application embodiment will be described with reference to the accompanying drawings.

[0215] Based on the domain name protection strategy determination method provided above, this embodiment also provides a domain name protection strategy determination device. This device includes multiple functional modules that interact to implement the method executed by the protection strategy determination device 100. The following is in conjunction with... Figure 8 This paper presents an implementation example of the multiple functional modules included in the protection strategy determination device.

[0216] See Figure 8The diagram illustrates the structure of a device for determining domain name protection strategies. Figure 8 As shown, the protection strategy determination device 800 includes:

[0217] The acquisition module 801 is used to acquire the domain name to be detected and its associated data, which are generated during the data processing related to the domain name to be detected.

[0218] The feature determination module 802 is used to determine the statistical features and semantic features of the domain name to be detected. The statistical features are determined based on the domain name to be detected and related data. The semantic features are used to characterize the contextual features of the domain name to be detected based on the domain name query sequence within a historical time period. The domain name query sequence includes multiple domain names arranged according to query time. The semantic features are determined based on the domain name to be detected.

[0219] The strategy determination module 803 is used to determine the protection strategy for the domain name to be detected based on statistical and semantic features.

[0220] In one possible implementation, the strategy determination module 803 is used to determine the scenario information of the domain name to be detected based on statistical features, wherein the scenario information is used to indicate the probability that the domain name to be detected belongs to each of multiple scenarios; calculate a first contribution degree of the statistical features and a second contribution degree of the semantic features based on the statistical features, semantic features and scenario information, wherein the first contribution degree is used to characterize the contribution degree of the statistical features in the process of determining the protection strategy and the second contribution degree is used to characterize the contribution degree of the semantic features in the process of determining the protection strategy; and determine the protection strategy for the domain name to be detected based on the statistical features, semantic features, first contribution degree and second contribution degree.

[0221] In one possible implementation, the strategy determination module 803 is used to project statistical features and semantic features onto a feature space respectively to obtain a first projection result corresponding to the statistical features and a second projection result corresponding to the semantic features, wherein the dimension of the first projection result and the dimension of the second projection result are the same; based on the first projection result, the second projection result and scene information, a first weight of the statistical features and a second weight of the semantic features are calculated, wherein the first weight is used to characterize a first degree of contribution and the second weight is used to characterize a second degree of contribution; based on the first weight and the second weight, the first projection result and the second projection result are fused to obtain a fused feature; based on the fused feature, a protection strategy for the domain name to be detected is determined.

[0222] In one possible implementation, the strategy determination module 803 is used to process statistical features through a scene evaluation model to obtain scene information. The scene evaluation model is trained based on multiple sample statistical features and the scene labeling results corresponding to each sample statistical feature. The scene labeling results are used to indicate the scene to which the sample statistical features belong.

[0223] In one possible implementation, the policy determination module 803 is used to determine the policy matching information of the domain name to be detected based on statistical features and semantic features. The policy matching information includes at least one of the risk level and security type of the domain name to be detected. Based on the policy matching information, a protection policy for the domain name to be detected is determined.

[0224] In one possible implementation, the policy matching information includes the risk level and security type of the domain name to be detected;

[0225] The policy determination module 803 is used to query the policy mapping relationship based on the risk level and security type to obtain the protection policy for the domain name to be detected. The policy mapping relationship is the correspondence between the risk level, security type and protection policy.

[0226] In one possible implementation, the policy determination module 803 is further configured to determine the risk level of the domain name to be detected as high risk level when the domain name to be detected meets the strong rule triggering conditions, and to obtain the protection policy for the domain name to be detected by querying the policy mapping relationship based on the high risk level and the security type of the domain name to be detected.

[0227] In one possible implementation, the policy determination module 803 is further configured to, after obtaining the protection policy for the domain name to be detected, adjust the protection action in the protection policy to an alarm action or a manual review action if the domain name to be detected meets the whitelist downgrade conditions.

[0228] In one possible implementation, the feature determination module 802 is used to determine multiple statistical sub-features of the domain name to be detected based on the domain name to be detected and associated data. The multiple statistical sub-features include at least two of the following: domain name character features, domain name resolution features, domain name registration information features, digital certificate features, threat intelligence features, and encrypted domain name resolution traffic features; and generate statistical features of the domain name to be detected based on the multiple statistical sub-features.

[0229] In one possible implementation, the feature determination module 802 is used to process the domain name to be detected through the domain name semantic model to obtain semantic features. The domain name semantic model is trained based on multiple training samples. Each training sample includes a central domain name and associated domain names. The central domain name and associated domain names are domain names in the domain name query sequence that meet the context selection conditions.

[0230] because Figure 8 The protection strategy determination device 800 shown corresponds to the above. Figure 3 The protection strategy determination device 100 in the illustrated embodiment, therefore Figure 8 For the specific implementation of the protection strategy determination device 800 and its technical effects, please refer to the above. Figure 3 The relevant details in the illustrated embodiments are described in detail here, and will not be repeated here.

[0231] Figure 9 This is a schematic diagram of the structure of a computing device provided in this application. Figure 9 As shown, the computing device 900 includes a processor 901, a memory 902, a communication interface 903, and a bus 904. The processor 901, memory 902, and communication interface 903 communicate via the bus 904. The bus 904 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The symbol is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. Communication interface 903 is used for communication with external devices.

[0232] It should be understood that in the embodiments of this application, processor 901 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete device assemblies, etc. General-purpose processors may be microprocessors or any conventional processors, etc.

[0233] Memory 902 may include read-only memory and random access memory, and provides instructions and data to processor 901. Memory 902 may also include non-volatile random access memory.

[0234] The memory 902 can be volatile or non-volatile, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which serves as an external cache. For example, random access memory may include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0235] The memory 902 stores executable code, and the processor 901 executes the executable code to perform the aforementioned actions. Figure 3 The method performed by the protection strategy determination device 100 in the illustrated embodiment.

[0236] It should be understood that the computing device 900 in this application embodiment can correspond to the protection strategy determination device 100 in this application embodiment, and can correspond to the execution of the embodiments in this application embodiment. Figure 3 The protection strategy determination device 100 in the illustrated method executes the method. The above and other operations and / or functions implemented by the computing device 900 are respectively used to implement... Figure 3 The corresponding methods and processes are omitted here for the sake of brevity.

[0237] This application can be further combined based on the above-described embodiments. As long as there is no conflict between different embodiments, the corresponding technical features can be combined to obtain other possible embodiments.

[0238] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium accessible to a computing device, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital versatile disc (DVD); or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium stores instructions that, when executed on a computing device, cause the computing device to perform the aforementioned method for determining the domain name protection policy.

[0239] This application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application are generated.

[0240] Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. Wired means may include, for example, coaxial cable, fiber optic cable, or digital subscriber line (DSL), while wireless means may include, for example, infrared, wireless, or microwave.

[0241] The computer program product can be a software installation package. When it is necessary to execute the aforementioned method for determining the domain name protection policy, the computer program product can be downloaded and run on a computing device, causing the computing device to execute the aforementioned method for determining the domain name protection policy.

[0242] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product may include one or more computer instructions. When the computer instructions are loaded or executed on a computing device, the processes or functions described in the embodiments of this application can be generated, in whole or in part.

[0243] Computing devices can be general-purpose computers, special-purpose computers, computer networks, or other programmable devices. Computer instructions can be stored in computer-readable storage media or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. Wired means can include, for example, coaxial cable, fiber optic cable, or digital subscriber line, while wireless means can include, for example, infrared, wireless, or microwave.

[0244] Computer-readable storage media can be any available medium that can be accessed by a computing device, or a data storage device such as a data center that includes one or more available media. Available media can be magnetic media, such as floppy disks, hard disks, or magnetic tapes; optical media, such as digital versatile optical discs; or semiconductor media, such as solid-state drives.

[0245] The terminology used in the above embodiments is for describing the corresponding embodiments and is not intended to limit this application. As used in the specification and claims of this application, the singular forms "a," "an," "the," "the," "the," and "this" may also include one or more instances, unless the context clearly indicates otherwise.

[0246] It should also be understood that in the embodiments of this application, "one or more" can refer to one, two, or more; the character " / " generally indicates an "or" relationship between related objects. In the embodiments of this application, "simultaneously" can refer to the corresponding processes occurring within the same time period, including the situation where the corresponding processes occur at the same moment.

[0247] In this specification, descriptions of "one embodiment," "some embodiments," "one possible implementation," etc., indicate that a specific feature, structure, or characteristic described in connection with the corresponding embodiment or implementation is included in one or more embodiments of this application. Therefore, expressions such as "in one embodiment," "in some embodiments," "in one possible implementation," and "in other possible implementations" appearing in different locations in this specification do not necessarily refer to the same embodiment, but may refer to one or more, but not all, embodiments, unless otherwise expressly stated.

[0248] The terms “including,” “comprising,” “having,” and variations thereof all mean including but not limited to, unless otherwise expressly stated.

[0249] The above description is merely a specific embodiment of this application, and the scope of protection of this application is not limited thereto. Any person skilled in the art can conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and all such modifications or substitutions should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for determining a domain name protection strategy, characterized in that, The method includes: Obtain the domain name to be detected and its associated data, which are generated during data processing related to the domain name to be detected; The statistical features and semantic features of the domain name to be detected are determined, wherein the statistical features are determined based on the domain name to be detected and the associated data, and the semantic features are used to characterize the contextual features of the domain name to be detected based on the domain name query sequence within a historical time period, wherein the domain name query sequence includes multiple domain names arranged according to query time, and the semantic features are determined based on the domain name to be detected. Based on the statistical features and the semantic features, a protection strategy is determined for the domain name to be detected.

2. The method according to claim 1, characterized in that, The step of determining a protection strategy for the domain name to be detected based on the statistical features and the semantic features includes: Based on the statistical characteristics, the scenario information of the domain name to be detected is determined, and the scenario information is used to indicate the probability that the domain name to be detected belongs to each of the multiple scenarios. Based on the statistical features, the semantic features, and the scenario information, a first contribution degree of the statistical features and a second contribution degree of the semantic features are calculated. The first contribution degree is used to characterize the contribution degree of the statistical features in determining the protection strategy, and the second contribution degree is used to characterize the contribution degree of the semantic features in determining the protection strategy. Based on the statistical features, the semantic features, the first contribution level, and the second contribution level, a protection strategy is determined for the domain name to be detected.

3. The method according to claim 2, characterized in that, The step of calculating the first contribution degree of the statistical features and the second contribution degree of the semantic features based on the statistical features, the semantic features, and the scene information includes: The statistical features and the semantic features are projected onto the feature space respectively to obtain a first projection result corresponding to the statistical features and a second projection result corresponding to the semantic features, wherein the dimensions of the first projection result and the second projection result are the same; Based on the first projection result, the second projection result, and the scene information, calculate the first weight of the statistical feature and the second weight of the semantic feature. The first weight is used to characterize the first contribution level, and the second weight is used to characterize the second contribution level. The step of determining a protection strategy for the domain name to be detected based on the statistical features, the semantic features, the first contribution level, and the second contribution level includes: Based on the first weight and the second weight, the first projection result and the second projection result are fused to obtain the fused feature; Based on the fusion characteristics, a protection strategy is determined for the domain name to be detected.

4. The method according to claim 2, characterized in that, The step of determining the scenario information of the domain name to be detected based on the statistical characteristics includes: The statistical features are processed by a scene evaluation model to obtain the scene information. The scene evaluation model is trained based on multiple sample statistical features and the scene annotation results corresponding to each sample statistical feature. The scene annotation results are used to indicate the scene to which the sample statistical features belong.

5. The method according to claim 1, characterized in that, The step of determining a protection strategy for the domain name to be detected based on the statistical features and the semantic features includes: Based on the statistical features and the semantic features, the policy matching information of the domain name to be detected is determined, wherein the policy matching information includes at least one of the risk level and security type of the domain name to be detected; Based on the policy matching information, a protection policy is determined for the domain name to be detected.

6. The method according to claim 5, characterized in that, The policy matching information includes the risk level and security type of the domain name to be detected; The step of determining the protection strategy for the domain name to be detected based on the strategy matching information includes: Based on the risk level and the security type, the policy mapping relationship is queried to obtain the protection policy for the domain name to be detected, wherein the policy mapping relationship is the correspondence between risk level, security type and protection policy.

7. The method according to claim 1, characterized in that, The determination of the statistical characteristics of the domain name to be detected includes: Based on the domain name to be detected and the associated data, multiple statistical sub-features of the domain name to be detected are determined. The multiple statistical sub-features include at least two of the following: domain name character features, domain name resolution features, domain name registration information features, digital certificate features, threat intelligence features, and encrypted domain name resolution traffic features. Based on the various statistical sub-features, statistical features of the domain name to be detected are generated.

8. The method according to any one of claims 1 to 7, characterized in that, Determining the semantic features of the domain name to be detected includes: The domain name to be detected is processed by a domain name semantic model to obtain the semantic features. The domain name semantic model is trained based on multiple training samples. Each training sample includes a central domain name and associated domain names. The central domain name and the associated domain names are domain names in the domain name query sequence that meet the context selection conditions.

9. A device for determining protection strategies for domain names, characterized in that, The device includes: The acquisition module is used to acquire the domain name to be detected and the associated data of the domain name to be detected, which is formed during the data processing related to the domain name to be detected; A feature determination module is used to determine the statistical features and semantic features of the domain name to be detected. The statistical features are determined based on the domain name to be detected and the associated data. The semantic features are used to characterize the contextual features of the domain name to be detected based on the domain name query sequence within a historical time period. The domain name query sequence includes multiple domain names arranged according to query time. The semantic features are determined based on the domain name to be detected. The strategy determination module is used to determine a protection strategy for the domain name to be detected based on the statistical features and the semantic features.

10. A computing device, characterized in that, The computing device includes at least one processor and at least one memory, the at least one memory being used to store instructions, and the at least one processor being used to execute the instructions stored in the at least one memory to cause the computing device to perform the method of any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computing device, cause the computing device to perform the method of any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes instructions that, when executed on a computing device, cause the computing device to perform the method of any one of claims 1 to 8.