Domain name detection method and device, electronic equipment and storage medium
Through the method of word segmentation processing and vocabulary feature matching, the general subdomain names under the main domain name are determined, which solves the problem of insufficient comprehensive subdomain name detection in the existing technology, and achieves more comprehensive detection results.
Patent Information
- Application Number
- CN202311530154.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is prone to omissions when detecting subdomains under the main domain name, resulting in insufficient comprehensive detection results.
By determining the subdomain names under the main domain name based on preset tools, performing word segmentation processing, matching the target vocabulary features in the vocabulary feature library, and determining the general subdomain names, thereby comprehensively detecting the subdomain names under the main domain name.
It realizes comprehensive detection of subdomains under the main domain name, improving the comprehensiveness and accuracy of the detection results.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of network monitoring, and particularly relates to a method for detecting domain names, a device for detecting domain names, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of the Internet, various websites and online applications have emerged like mushrooms after rain, providing users with high-quality network services. In order to accelerate content transmission and reduce network latency, many websites have adopted content delivery network services. By deploying servers globally, users can access the required content more quickly. However, while the content delivery network service improves network performance, it also generates a large number of sub-domains under different main domain names, which is likely to trigger crises in terms of network security and performance. To resolve this crisis, it is crucial to deeply understand the relevant information of each sub-domain under the main domain name.
[0003] In the related art, a common domain name dictionary is used to perform exhaustive blasting on the main domain name. However, when this domain name detection method processes a main domain name containing a large number of sub-domains, it is easy to miss some, resulting in an incomplete detection result of the main domain name. Summary of the Invention
[0004] This application provides a method for detecting domain names, a device for detecting domain names, an electronic device, and a computer-readable storage medium, which helps to comprehensively determine the sub-domains under the main domain name.
[0005] In a first aspect, this application provides a method for detecting domain names, including: Determining at least one sub-domain under the main domain name based on at least one preset tool; Performing word segmentation processing on each sub-domain to obtain a corresponding set of word segmentation results, and each set of word segmentation results includes at least one basic vocabulary; Determining the target vocabulary feature of each basic vocabulary in each set of word segmentation results from a pre-constructed vocabulary feature library; the vocabulary feature library records multiple preset vocabularies and the vocabulary features corresponding to each preset vocabulary; the target vocabulary feature is the vocabulary feature corresponding to the preset vocabulary that matches the basic vocabulary among the preset vocabularies; For each set of word segmentation results, determining at least one general sub-domain based on the target vocabulary features corresponding to the basic vocabularies in the word segmentation results to obtain the detection result of the main domain name.
[0006] In a second aspect, this application provides a device for detecting domain names, including: A first determination module, configured to determine at least one sub-domain under the main domain name based on at least one preset tool; A word segmentation module, configured to perform word segmentation on each sub-domain name to obtain a corresponding set of word segmentation results, and each set of word segmentation results includes at least one basic vocabulary; A second determination module, configured to determine the target vocabulary feature of each basic vocabulary in each set of word segmentation results from a pre-constructed vocabulary feature library; the vocabulary feature library records multiple preset vocabularies and the corresponding vocabulary features of each preset vocabulary, and the target vocabulary feature is the vocabulary feature corresponding to the preset vocabulary that matches the basic vocabulary among the preset vocabularies; A third determination module, configured to, for each set of word segmentation results, determine at least one general sub-domain name based on the target vocabulary features corresponding to the basic vocabularies in the word segmentation results, and obtain the detection result of the main domain name.
[0007] In a third aspect, the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method in the first aspect are implemented.
[0008] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method in the first aspect are implemented.
[0009] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by one or more processors, the steps of the method in the first aspect are implemented.
[0010] The beneficial effect of the present application compared with the prior art is that: for the main domain name to be detected, at least one preset tool can be used to discover as many of its sub-domain names as possible. Among them, the number of sub-domain names is at least one. Different sub-domain names have different corresponding features. Based on the features of each sub-domain name, general sub-domain names can be determined. The general sub-domain names can be used as the basis for the derivation of different sub-domain names, that is, more sub-domain names can be derived according to the general sub-domain names, so as to cover the sub-domain names under the main domain name as much as possible, which helps to obtain a more comprehensive detection result.
[0011] Thus, after obtaining each sub-domain name, in order to improve the comprehensiveness of the detection results, the characteristics of each sub-domain name can be determined, and then the general sub-domain name corresponding to each sub-domain name can be determined. Specifically, for each sub-domain name, word segmentation processing can be performed on it to obtain a corresponding set of word segmentation results; among them, each set of word segmentation results can include at least one basic vocabulary. In order to determine the characteristics of each sub-domain name, for each set of word segmentation results, the target vocabulary characteristics can be matched for each basic vocabulary in the set of word segmentation results from the pre-constructed vocabulary feature library. The vocabulary feature library records multiple preset vocabularies, and each preset vocabulary corresponds to a vocabulary feature; the target vocabulary feature is the vocabulary feature corresponding to the preset vocabulary that matches the basic vocabulary among the preset vocabularies. For a sub-domain name, it can be considered that the target vocabulary characteristics matched by the corresponding word segmentation results are the characteristics of the sub-domain name; based on this, for each word segmentation result, at least one general sub-domain name can be determined according to the corresponding target vocabulary characteristics, and the general sub-domain names corresponding to each word segmentation result can be determined as the detection results of the main domain name.
[0012] It can be understood that the detection method of the present application does not enumerate all sub-domain names under the main domain name, but provides general sub-domain names under the main domain name, creating conditions for extensive sub-domain name detection, so that users can flexibly and extensively detect all sub-domain names under the main domain name based on these general sub-domain names. Therefore, it can be considered that the detection results of the present application are more comprehensive than those in the related art.
[0013] It can be understood that the beneficial effects of the above second aspect to the fifth aspect can refer to the relevant descriptions in the above first aspect and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 is a schematic flowchart of the method for detecting a domain name provided by an embodiment of the present application; Figure 2 is a schematic flowchart of the method for constructing a general vocabulary library provided by an embodiment of the present application; Figure 3 is a schematic flowchart of the method for constructing a professional vocabulary library provided by an embodiment of the present application; Figure 4 is a schematic flowchart of the method for determining a general sub-domain name provided by an embodiment of the present application; Figure 5It is a schematic structural diagram of a domain name detection device provided by an embodiment of the present application; Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Embodiment
[0016] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0017] In the related art, although there are currently many sub-domain name detection tools, these tools can only detect some sub-domains under the main domain name, that is, they cannot comprehensively detect all sub-domains under the main domain name.
[0018] In order to comprehensively detect all sub-domains under the main domain name, the present application proposes a domain name detection method, which helps to comprehensively determine the sub-domains under the main domain name. The control method proposed by the present application will be described below through specific embodiments.
[0019] The domain name detection method provided by the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.
[0020] Figure 1 It shows a schematic flowchart of the domain name detection method provided by the present application. The domain name detection method includes: Step 110: The electronic device determines at least one sub-domain under the main domain name based on at least one preset tool.
[0021] In the process of providing network services for users, in addition to providing high-quality content for users, the supplier should also ensure the performance and security of the network services. In order to ensure the performance and security of the network services, it is crucial to deeply understand all sub-domains under the main domain name of the network services.
[0022] For the main domain names to be deeply understood, that is, the main domain names in this embodiment, the number can be one, or two or more. When the number of main domain names is two or more, the electronic device can choose a parallel processing or sequential processing method to detect each main domain name. The specific processing method can be determined based on actual needs and is not limited in this application. For the convenience of description, subsequent descriptions will be based on the detection of one main domain name as an example.
[0023] In order to improve the comprehensiveness of the detection results of the main domain name, the electronic device can determine the common sub-domains under the main domain name, thereby laying a foundation for flexibly and widely detecting each sub-domain under the main domain name. In order to determine as many common sub-domains under the main domain name as possible, before determining the common sub-domains, the electronic device can determine at least one sub-domain under the main domain name based on at least one preset tool, so as to facilitate subsequent processing of the sub-domain to obtain the corresponding common sub-domain.
[0024] Optionally, the preset tool refers to each tool for determining sub-domains, such as the record information of the Domain Name System (DNS) and / or different sub-domain name dictionaries. Based on this, the foregoing step 110 specifically includes: The electronic device matches at least one sub-domain for the main domain name based on the record information of the domain name system and / or the domain name dictionary to determine at least one sub-domain under the main domain name.
[0025] Specifically, when the preset tool is the record information of DNS, the electronic device can parse the record information of DNS to obtain parsing information, and the parsing information contains the relevant information of the recorded main domain name and its sub-domains. Based on this, the electronic device can use the main domain name as a query condition to match sub-domains for the main domain name from the parsing information.
[0026] When the preset tool is a domain name dictionary, the electronic device can brute-force crack each sub-domain under the main domain name by means of enumeration.
[0027] It can be understood that different preset tools have their own advantages and disadvantages. Therefore, in order to determine more sub-domains, the electronic device can simultaneously use different preset tools to determine the sub-domains under the main domain name.
[0028] Step 120: The electronic device performs word segmentation processing on each sub-domain name to obtain a corresponding set of word segmentation results.
[0029] In order to facilitate the determination of the characteristics of each sub-domain name, after obtaining each sub-domain name, the electronic device can perform word segmentation processing on each sub-domain name to obtain a set of word segmentation results; wherein, each set of word segmentation results includes at least one basic vocabulary.
[0030] For example only, assume that the main domain name is "helloworld.com". Through step 110, its sub-domains "monigongji.helloworld.com" and "shentou.helloworld.com" are determined. After word segmentation of "monigongji.helloworld.com", words such as "moni", "gongji", "helloworld", and "com" can be obtained. Obviously, among the obtained words, "helloworld" and "com" are the components that are the same in the main domain name and the sub-domain name, and it is difficult to reflect the characteristics of the sub-domain name. Therefore, this part is not used as the result of word segmentation, and only "moni" and "gongji" are used as basic words; that is, the electronic device can determine these two basic words as the result of word segmentation of the sub-domain name "monigongji.helloworld.com". By analogy, in the result of word segmentation of the sub-domain name "shentou.helloworld.com", its basic word is "shentou".
[0031] Step 130: The electronic device determines the target word features of each basic word in each set of word segmentation results from the pre-constructed word feature library.
[0032] In order to accurately determine the characteristics of each sub-domain name, the electronic device can pre-construct a word feature library, which records multiple preset words, and each preset word corresponds to a word feature. Among them, the preset words can be historical basic words or basic words analyzed and predicted. Based on the word feature library, for each set of word segmentation results, the electronic device can determine the target word features for each corresponding basic word.
[0033] Optionally, the step of determining the target word feature may include: if a set of word segmentation results only includes one basic word, then the electronic device can directly use this basic word as a query condition to determine the matching preset word from the word feature library, and determine the word feature corresponding to the matching preset word as the target word feature of this basic word. If a set of word segmentation results includes more than two basic words, for each basic word, its target word feature can be determined by referring to the foregoing determination method; in addition, in order to improve the matching efficiency, the electronic device can also perform parallel queries, that is, use the basic words included in this set of word segmentation results as sub-conditions in the query condition, and combine the sub-conditions through specific operators to determine the target word features corresponding to each basic word at one time.
[0034] Step 140: For each set of word segmentation results, the electronic device determines at least one general sub-domain name based on the target word features corresponding to the basic words in the word segmentation results, and obtains the detection result of the main domain name.
[0035] It can be considered that for each sub - domain name, the target word features corresponding to each basic word in its word - segmentation result are the features of the sub - domain name. Thus, for each set of word - segmentation results, the electronic device can determine at least one general sub - domain name based on the matched target word features. Given that the general sub - domain name can be used as the derivation basis for many sub - domain names, the electronic device takes all the determined general sub - domain names as the detection result of the main domain name to improve the comprehensiveness of the main domain name detection result.
[0036] In this embodiment, the electronic device first determines at least one sub - domain name under the main domain name through at least one preset tool, and then processes each sub - domain name to obtain the features of each sub - domain name. For each sub - domain name, based on its features, at least one general sub - domain name can be determined. Given that the general sub - domain name can be used as the derivation basis for many sub - domain names, the general sub - domain names corresponding to each sub - domain name as the detection result of the main domain name can improve the comprehensiveness of the main domain name detection result.
[0037] In some embodiments, to accurately determine the features of each sub - domain name, the preset vocabulary can be divided into general vocabulary and professional vocabulary. Among them, general vocabulary belongs to the commonly used vocabulary in different fields, usually composed of combinations of letters and numbers, and its word features are the vocabulary itself. Professional vocabulary, on the other hand, is the specific keyword vocabulary in a specified field. For example, in the field of penetration testing, for keywords such as "testing", "red team", "blue team", and "simulated attack", because these professional vocabulary have special meanings in the penetration field, they can be used as professional vocabulary in the penetration testing field. The word features of professional vocabulary may be vocabulary abbreviations or vocabulary pinyin, different from general vocabulary. To accurately determine the target word features of each basic word, based on the division of the preset vocabulary, the word feature library can be divided into a general word feature library and a professional word feature library.
[0038] In some embodiments, in order to be able to construct a general word feature library with rich data, refer to Figure 2 which shows a schematic flowchart of the general word feature library construction method, specifically including: Step 210: The electronic device obtains initial general vocabulary from the pre - acquired data set.
[0039] The data set includes the data set formed by manual collection and / or different open - source vocabulary libraries. Only as an example, the open - source vocabulary library can include vocabulary libraries such as WordNet and / or Google's Word2Vec. Due to the low efficiency of the manual collection method and the low coverage of the data set formed by the collection for general vocabulary, it is preferred to obtain the initial general vocabulary through the open - source vocabulary library.
[0040] Step 220: The electronic device pre - processes the initial general vocabulary to obtain general vocabulary.
[0041] For initial general vocabulary, there may be problems such as format and duplication. To avoid excessive redundant data or defective data in the general vocabulary feature library, the electronic device can preprocess the initial general vocabulary to obtain the general vocabulary.
[0042] Only as an example, for the initial general vocabulary, the electronic device can first unify its format and then filter out the duplicate initial general vocabulary to obtain high-quality general vocabulary.
[0043] Step 230: The electronic device constructs a general vocabulary feature library based on the general vocabulary.
[0044] As can be seen from the foregoing, the lexical feature of the general vocabulary is the general vocabulary itself. Based on this, after obtaining the general vocabulary, the electronic device can directly construct a general vocabulary feature library based on the general vocabulary. It can be understood that because the lexical feature of the general vocabulary is the general vocabulary itself, the corresponding relationship between the general vocabulary and its lexical feature may not be recorded in the general vocabulary feature library.
[0045] In the embodiments of the present application, the electronic device obtains the initial general vocabulary from the open-source vocabulary library and / or the dataset collected manually, which can enrich the general vocabulary feature library as much as possible; by preprocessing the initial general vocabulary, general vocabulary with higher data quality can be obtained; the general vocabulary feature library constructed based on this general vocabulary not only has higher data quality but also higher data richness, which helps to accurately determine the target lexical feature of the basic vocabulary subsequently.
[0046] In some embodiments, the electronic device can download and install WordNet and use the nltk library of Python to read and parse the relevant files of WordNet. Then, download the lexical data of WordNet through nltk.download, obtain the initial general vocabulary set from the downloaded lexical data through noun_synsets = wordnet.all_synsets(wordnet.NOUN), preprocess the initial general vocabulary set to obtain the general vocabulary set, and save the general vocabulary set to the local database, thus completing the construction of the general vocabulary feature library.
[0047] In some embodiments, in order to construct a more comprehensive and accurate professional vocabulary feature library, refer to Figure 3 , which shows a schematic flowchart of the method for constructing a professional vocabulary feature library, specifically including: Step 310: The electronic device obtains multiple sub-domain name samples under a preset domain.
[0048] In different fields, the corresponding professional vocabulary is different, that is, the domain characteristics of professional vocabulary are relatively strong. Therefore, by paying attention to which fields, the electronic device can specifically obtain the sub-domains under the corresponding fields as samples, so as to subsequently obtain the professional vocabulary under the preset fields from these samples. That is to say, the preset fields can be at least one of the fields of concern, such as the computer field, the penetration testing field, and the medical field; for each preset field, the electronic device can obtain the sub-domain samples under the preset field.
[0049] Step 320: The electronic device extracts the corresponding professional vocabulary from each sub-domain sample.
[0050] The sub-domain samples corresponding to different preset fields contain different professional vocabularies. In order to obtain the professional vocabulary under each preset field, the electronic device can extract the corresponding professional vocabulary from each sub-domain sample; among them, the professional vocabulary in each sub-domain sample is at least one. For the extraction result corresponding to each sub-domain sample, it is recorded as a group of professional vocabulary. Correspondingly, in step 310, for how many sub-domain samples are obtained, in this step, the electronic device can extract the same number of groups of professional vocabulary.
[0051] Optionally, for each sub-domain sample, the electronic device can perform the word segmentation operation as described in the foregoing step 120 to obtain the word segmentation result of each sub-domain sample. The word segmentation result may contain non-professional vocabulary. In order to accurately determine the professional vocabulary from the word segmentation result, the electronic device can screen the word segmentation result according to preset conditions to obtain a group of professional vocabulary corresponding to the sub-domain sample. Among them, the preset conditions can be set based on different fields and are not limited here.
[0052] Optionally, in addition to extracting professional vocabulary from sub-domain samples, the electronic device can also crawl the professional vocabulary of each preset field. That is to say, the professional vocabulary is not limited to the vocabulary that has already appeared in the sub-domain, but also includes some professional terms or industry keywords that may appear in the sub-domain in the future.
[0053] Step 330: The electronic device inputs each professional vocabulary into a pre-trained feature extraction network model to obtain the vocabulary feature corresponding to each professional vocabulary.
[0054] After obtaining the professional vocabulary, in order to accurately represent the features of each professional vocabulary, the electronic device can input each professional vocabulary into a pre-trained feature extraction network model. The feature extraction network model can process each received professional vocabulary to obtain the output of the hidden layer of the feature extraction network model, that is, obtain the vocabulary feature corresponding to each professional vocabulary. It can be understood that the model can capture key information from the professional vocabulary, which helps to determine the vocabulary feature of the professional vocabulary.
[0055] For example, assume that the professional term is "test". When it is input into the feature extraction network model, the lexical features of "test" that can be obtained include "ceshi" and "cs".
[0056] Step 340: The electronic device constructs a professional vocabulary feature library based on each professional vocabulary and the corresponding lexical features of each professional vocabulary.
[0057] After determining each professional vocabulary, the electronic device can record each professional vocabulary in the database. For each professional vocabulary, the corresponding lexical features can also be further recorded. After the recording is completed, a professional vocabulary feature library can be obtained.
[0058] In the embodiments of the present application, the electronic device first determines the professional vocabulary in each field, then determines the lexical features of each professional vocabulary based on a pre-trained feature extraction network, and finally can construct a relatively comprehensive and accurate professional vocabulary feature library based on each professional vocabulary and the corresponding lexical features of each professional vocabulary.
[0059] In some embodiments, it can be understood that whether it is a general vocabulary or a professional vocabulary, new words will continue to be generated over time. To improve the comprehensiveness of these two vocabulary feature libraries, the electronic device can update these two vocabulary feature libraries regularly.
[0060] In some embodiments, in order to accurately determine the lexical features of each professional vocabulary, a Convolutional Neural Networks (CNN) model can be used as the feature extraction network model. By processing the text sequence through the CNN model, the local features of the text sequence can be accurately extracted, which helps to accurately determine the lexical features of each professional vocabulary. Among them, the convolutional neural network model includes a convolutional layer and a fully connected layer.
[0061] Specifically, the convolutional layer is mainly used to extract the local features of professional vocabulary. It can be understood that for the convenience of the electronic device to process, each professional vocabulary can be encoded to obtain a corresponding vector. That is to say, when the convolutional layer extracts the local features of each professional vocabulary, it processes the vector of the professional vocabulary to output the local features. In other words, the vectors of each professional vocabulary are the input sequences of the convolutional layer.
[0062] For different convolutional layers, they correspond to different convolutional kernels. The adjustable parameters of the convolutional kernel are the size and number of channels of the convolutional kernel, which are mainly used to extract local features from the input sequence. For different input sequences, their lengths are different. In order to accurately extract the most representative local features, the convolutional layer can adjust the stride according to the length of the input sequence, that is, the sliding amplitude of the convolutional kernel on the input sequence. Among them, different strides can extract features of different scales.
[0063] Based on the above definitions, for the convolutional layer, its output can be expressed by the following formula: T[i] = activation(∑(X[i * S + k] * W[k]) + b conv ) Where T is the output of the convolutional layer; i is the index of T, representing the position of the output feature in the input sequence; X is the input sequence; S is the stride; W is the convolutional kernel, k is the index of the convolutional kernel, and k matches i. For different k, the corresponding convolutional kernels are different, and the local features that can be extracted are different. That is to say, different local features of professional terms can be extracted by setting the convolutional kernel; activation is the activation function, and b conv is the bias term of the convolutional layer.
[0064] The output formula corresponding to the convolutional layer is the convolution process of the convolutional layer, which is used to extract the local features in professional terms. Only as an example, assume that the professional terms include "test", "red team", "blue team", and "simulation attack". Each professional term can be encoded to obtain a corresponding vector, and the vectors are combined to obtain the input sequence X. For different vectors, different convolutional kernels can be set. Taking the vector corresponding to "test" as an example, assume that the convolutional kernel W[k] and the stride S are set to perform a convolution operation on the input sequence X. Each time a convolution operation is performed, a vector in the input sequence X can be processed, that is, the vector is multiplied by the convolutional kernel to obtain the convolution result of the vector. After processing each vector in the input sequence X, 5 convolution results can be obtained. Summing these 5 convolution results can obtain the convolution sum. Adding the convolution sum to the bias term b conv and then performing a non-linear transformation based on the activation function activation can capture the specific form and structural information of the professional term "test", that is, determine the local features of "test". For other words, and so on.
[0065] The fully connected layer is used to predict the vocabulary features of professional terms based on local features. It can be understood that the output T[i] of the convolutional layer is the input of the fully connected layer, and the output of the fully connected layer can be expressed by the following formula: Y = activation(T[i] * w + b fc ) Where Y is the output of the fully connected layer, w is the weight of the fully connected layer, and b fc is the bias term of the fully connected layer.
[0066] In the output formula corresponding to the fully connected layer, the fully connected layer multiplies the output T[i] of the convolutional layer by the weight of the fully connected layer wMultiply them and then add the bias term. To meet the prediction requirements of lexical features, the output of the fully connected layer is passed through an activation function to represent similarity, thereby accurately predicting the lexical features corresponding to each professional term.
[0067] In some embodiments, for the CNN model, during its training process, assuming the learning rate is η and the gradient of the loss function with respect to the model parameter θ is ∇G, the formula for learning rate update can be expressed as: θ = θ - η * ∇G Among them, the learning rate update can ensure that the model parameters are optimized in the direction of reducing the loss.
[0068] In some embodiments, the regularization term for the weight decay of the fully connected layer, that is, the loss function of the CNN model, can be expressed as: L = L + λ * ∑(θ 2 ) L is the loss value, and λ is the weight decay coefficient, which determines the strength of regularization. By adding the sum of squares of the model parameters θ as a penalty term in the loss function, the model can be made to preferentially select smaller model parameter values during training to reduce the risk of overfitting.
[0069] In some embodiments, in order to identify as many generic subdomains as possible, refer to Figure 4 , which shows a schematic flowchart of the method for determining generic subdomains, specifically including: Step 410, the electronic device determines at least one combination method among the target lexical features.
[0070] As can be seen from the foregoing embodiments, each basic term may correspond to at least one target lexical feature. Based on this, for the target lexical features corresponding to a subdomain, there may be at least one combination method. In order to identify as many generic subdomains as possible, the electronic device can determine at least one combination method based on the target lexical features.
[0071] Example 1, assume that the basic term in a set of word segmentation results is "test". Based on matching in the lexical feature library, it is determined that the target lexical features corresponding to "test" are "ceshi" and "cs". Based on these two target lexical features, it can be determined that there are two combination methods, namely "ceshi" and "cs".
[0072] Example 2, assume that the basic terms in a set of word segmentation results include "moni" and "gongji". After matching, it is determined that the target lexical feature corresponding to "moni" is "m", and the target lexical feature corresponding to "gongji" is "g". It can be determined that there is one combination method for the two target lexical features, which is "mg".
[0073] Step 420: The electronic device combines each target vocabulary feature based on each combination method to obtain a generic sub-domain name.
[0074] It can be understood that for one combination method, the electronic device can determine a generic sub-domain name. Assuming that the main domain name is "helloworld.com", for the first example above, two generic sub-domain names can be determined, namely "ceshi.helloworld.com" and "cs.helloworld.com". For the second example above, one generic sub-domain name can be determined, which is "mg.helloworld.com".
[0075] In the embodiment of the present application, for each target vocabulary feature corresponding to a sub-domain name, the electronic device can determine the feasible combination methods of each target vocabulary feature; for each combination method, a generic sub-domain name can be determined. That is to say, the number of generic sub-domain names that can be determined is equal to the number of combination methods of a set of target vocabulary features, which can provide a solid foundation for exploring sub-domain names under the subsequent main domain name.
[0076] In some embodiments, in step 110, the electronic device may determine at least two sub-domain names with common features. For different sub-domain names with common features, after extracting features, the determined generic sub-domain names are repeated. To reduce redundant data, after obtaining all the generic sub-domain names under the main domain name, the electronic device can remove duplicates from all the generic sub-domain names to obtain the final detection result. That is to say, all the generic sub-domain names after deduplication are the detection results of the main domain name.
[0077] In some embodiments, in order to make the generic sub-domain name more widely used, as an example only, the generic sub-domain name can be used as the derivation basis of each sub-domain name. The electronic device can combine with automation tools to generate possible sub-domain names by adding prefixes, suffixes, and / or numbers to the features, so as to flexibly and widely explore the sub-domain names under the main domain name.
[0078] As another example, the sub-domain names used by malicious actors for deception, phishing, or attacks generally also have certain features. By combining this feature with the generic sub-domain name, the electronic device can predict possible malicious sub-domain names in advance.
[0079] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0080] Corresponding to the domain name detection method in the above embodiments Figure 5The structural block diagram of the domain name detection device 5 provided by the embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.
[0081] Referring to Figure 5 , the domain name detection device 5 includes: The first determination module 51 is used to determine at least one sub-domain name under the main domain name based on at least one preset tool; The word segmentation module 52 is used to perform word segmentation processing on each sub-domain name to obtain a corresponding set of word segmentation results, and each set of word segmentation results includes at least one basic vocabulary; The second determination module 53 is used to determine the target vocabulary feature of each basic vocabulary in each set of word segmentation results from the pre-constructed vocabulary feature library; multiple preset vocabularies and the corresponding vocabulary features of each preset vocabulary are recorded in the vocabulary feature library, and the target vocabulary feature is the vocabulary feature corresponding to the preset vocabulary that matches the basic vocabulary among the preset vocabularies; The third determination module 54 is used to, for each set of word segmentation results, determine at least one general sub-domain name based on the target vocabulary features corresponding to the basic vocabularies in the word segmentation results, and obtain the detection result of the main domain name.
[0082] Optionally, the preset vocabulary includes general vocabulary, and correspondingly, the vocabulary feature library includes a general vocabulary feature library; the detection device 5 may further include: The first acquisition module is used to acquire initial general vocabulary from the pre-acquired dataset; The preprocessing module is used to preprocess the initial general vocabulary to obtain general vocabulary; The first construction module is used to construct a general vocabulary feature library based on the general vocabulary, where the vocabulary feature of the general vocabulary is the general vocabulary itself.
[0083] Optionally, the preset vocabulary includes professional vocabulary, and correspondingly, the vocabulary feature library includes a professional vocabulary feature library; the detection device 5 may further include: The second acquisition module is used to acquire multiple sub-domain name samples under the preset domain; The first extraction module is used to extract the corresponding professional vocabulary from each sub-domain name sample; The second extraction module is used to input each professional vocabulary into the pre-trained feature extraction network model to obtain the vocabulary feature corresponding to each professional vocabulary; The second construction module is used to construct a professional vocabulary feature library based on each professional vocabulary and the vocabulary feature corresponding to each professional vocabulary.
[0084] Optionally, the third determination module 54 may include: The determination unit is used to determine at least one combination method between the target vocabulary features; A combination unit for combining each target vocabulary feature based on each combination method to obtain a general sub-domain name.
[0085] Optionally, the detection device 5 may further include: A deduplication module for deduplicating all the general sub-domain names corresponding to the main domain name, and determining the deduplicated general sub-domain names as the detection results.
[0086] Optionally, the preset tool includes record information of the domain name system and / or a domain name dictionary. The first determination module 51 is specifically configured to: match at least one sub-domain name for the main domain name based on the record information of the domain name system and / or the domain name dictionary, so as to determine at least one sub-domain name under the main domain name.
[0087] It should be noted that the content such as the information interaction and execution process between the above-mentioned devices / units, due to being based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details are not described herein again.
[0088] Figure 6 This is a schematic structural diagram of the physical layer of an electronic device provided in an embodiment of the present application. As Figure 6 shown, the electronic device 6 in this embodiment includes: at least one processor 60 ( Figure 6 only one is shown in the figure), a processor, a memory 61, and a computer program 62 stored in the memory 61 and executable on at least one processor 60. When the processor 60 executes the computer program 62, the steps in the method embodiment for detecting any domain name are implemented, such as Figure 1 the steps 110-140 shown.
[0089] The so-called processor 60 may be a central processing unit (CPU), and this processor 60 may also be other general processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0090] The memory 61 may be an internal storage unit of the electronic device 6 in some embodiments, such as the hard disk or memory of the electronic device 6. In some other embodiments, the memory 61 may also be an external storage device of the electronic device 6, such as a plug-in hard disk equipped on the electronic device 6, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0091] Furthermore, the memory 61 may also include both the internal storage unit of the electronic device 6 and the external storage device. The memory 61 is used to store an operating device, application programs, a BootLoader, data, and other programs, such as the program code of a computer program, etc. The memory 61 may also be used to temporarily store the data that has been output or will be output.
[0092] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the above-mentioned device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0093] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.
[0094] The embodiments of the present application provide a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the foregoing method embodiments when executed.
[0095] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct relevant hardware to complete. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the above computer program includes computer program code, and the above computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The above computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0096] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0097] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0098] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the above division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0099] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0100] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for detecting domain names, characterized in that, it includes: Determining at least one sub-domain name under the main domain name based on at least one preset tool; Performing word segmentation processing on each of the sub-domain names to obtain a corresponding set of word segmentation results, and each set of the word segmentation results includes at least one basic vocabulary; Determining the target vocabulary feature of each of the basic vocabularies in each set of the word segmentation results from a pre-constructed vocabulary feature library; multiple preset vocabularies are recorded in the vocabulary feature library, and the vocabulary feature corresponding to each of the preset vocabularies; the target vocabulary feature is the vocabulary feature corresponding to the preset vocabulary that matches the basic vocabulary among the preset vocabularies; For each set of the word segmentation results, determining at least one general sub-domain name based on the target vocabulary features corresponding to the basic vocabularies in the word segmentation results to obtain the detection result of the main domain name.
2. The method for detecting domain names according to claim 1, characterized in that, The preset vocabulary includes general vocabulary, and correspondingly, the vocabulary feature library includes a general vocabulary feature library; the general vocabulary feature library is constructed based on the following steps: Obtaining initial general vocabulary from a pre-acquired dataset; Performing preprocessing on the initial general vocabulary to obtain the general vocabulary; Constructing the general vocabulary feature library based on the general vocabulary, wherein the vocabulary feature of the general vocabulary is the general vocabulary itself.
3. The method for detecting domain names according to claim 2, characterized in that, The preset vocabulary includes professional vocabulary, and correspondingly, the vocabulary feature library includes a professional vocabulary feature library; the professional vocabulary feature library is constructed based on the following steps: Obtaining a plurality of sub-domain name samples under a preset domain; Extracting the corresponding professional vocabulary from each of the sub-domain name samples; Inputting each of the professional vocabularies into a pre-trained feature extraction network model to obtain the vocabulary feature corresponding to each of the professional vocabularies; Constructing the professional vocabulary feature library based on each of the professional vocabularies and the vocabulary features corresponding to each of the professional vocabularies.
4. The method for detecting domain names according to claim 3, characterized in that, The feature extraction model is a convolutional neural network model, and the convolutional neural network model includes a convolutional layer and a fully connected layer; The convolutional layer is used to extract the local features of the professional vocabulary; The fully connected layer is used to predict the vocabulary feature of the professional vocabulary according to the local features.
5. The method for detecting domain names according to claim 3, characterized in that, The determining at least one general sub-domain name based on the target vocabulary features corresponding to the basic vocabularies in the word segmentation results includes: Determining at least one combination method among the target vocabulary features; Combining the target vocabulary features based on each of the combination methods to obtain the general sub-domain name.
6. The method for detecting domain names according to any one of claims 1 to 5, characterized in that, After determining at least one general sub-domain name based on the target vocabulary features corresponding to the basic vocabularies in the word segmentation results, it further includes: Deduplicate all the general sub-domains corresponding to the main domain name, and determine each of the deduplicated general sub-domains as the detection result.
7. The domain name detection method according to any one of claims 1 to 5, wherein, the preset tool includes record information of the domain name system and / or a domain name dictionary, and the determining of at least one sub-domain under the main domain name based on at least one preset tool includes: Based on the record information of the domain name system and / or the domain name dictionary, match at least one of the sub-domains for the main domain name to determine at least one of the sub-domains under the main domain name.
8. A domain name detection device, wherein, it includes: A first determination module, configured to determine at least one sub-domain under the main domain name based on at least one preset tool; A word segmentation module, configured to perform word segmentation processing on each of the sub-domains to obtain a corresponding set of word segmentation results, and each set of the word segmentation results includes at least one basic vocabulary; A second determination module, configured to determine the target vocabulary feature of each of the basic vocabularies in each set of the word segmentation results from a pre-constructed vocabulary feature library; The vocabulary feature library records a plurality of preset vocabularies and the vocabulary features corresponding to each of the preset vocabularies, and the target vocabulary feature is the vocabulary feature corresponding to the preset vocabulary that matches the basic vocabulary among the preset vocabularies; A third determination module, configured to, for each set of the word segmentation results, determine at least one general sub-domain based on the target vocabulary features corresponding to the basic vocabularies in the word segmentation results, to obtain the detection result of the main domain name.
9. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, it implements the domain name detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, it implements the domain name detection method according to any one of claims 1 to 7.