A method, apparatus, device and medium for determining subdomains

By generating and encoding business description text in a large model, and combining enterprise identifiers and website information, the subdomain prefixes of other main domains similar to the target main domain are determined. This solves the problem of high computational resource and time consumption in existing subdomain determination methods, and achieves efficient subdomain generation and brute-force attacks.

CN119449375BActive Publication Date: 2025-11-14CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411433275.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-11-14
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

Existing methods for determining subdomains require significant computational resources and time, resulting in low processing efficiency.

Method used

By inputting the target main domain name, enterprise identifier, and website information into a large model, business description text is generated and encoded into a vector. Combined with pre-saved candidate main domain name vectors, other main domain names that meet the similarity requirements are determined, and the target subdomain is generated using the subdomain prefixes of these main domain names.

Benefits of technology

It improves the efficiency of subdomain determination, reduces computing resources and time, and can more accurately capture the business information behind the main domain, providing precise data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119449375B_ABST
    Figure CN119449375B_ABST
Patent Text Reader

Abstract

This application relates to the field of network security technology, and in particular to a method, apparatus, device, and medium for determining subdomains. In the embodiments of this application, by using a large model and combining enterprise identifiers, target main domains, and website information, business description text is generated and encoded into a target feature vector. This not only accurately captures the business information behind the target main domain but also provides more precise data support for subsequent subdomain generation and brute-force attacks. Furthermore, the embodiments of this application also disclose constructing target subdomains corresponding to the target main domain based on subdomain prefixes corresponding to other main domains similar to the target main domain, improving the efficiency of subdomain determination and reducing computational resources and time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to a method, apparatus, device and medium for determining subdomains. Background Technology

[0002] Subdomain brute-force attacks are a critical task in cybersecurity, primarily aimed at discovering and identifying website subdomains. This is essential for attack surface analysis, vulnerability scanning, and information gathering. With the increasing complexity of network environments and the growing number of security threats, subdomain brute-force attacks have become particularly important. Subdomains often host different applications and services, and discovering these subdomains helps security experts identify potential security vulnerabilities and risks.

[0003] Existing methods for determining subdomains typically rely on a large static dictionary that stores multiple subdomain prefixes. Electronic devices construct each subdomain based on each subdomain prefix, which not only requires a lot of computing resources and time but is also inefficient. Summary of the Invention

[0004] This application provides a method, apparatus, device, and medium for determining subdomains, which solves the problems of existing subdomain determination processes requiring a large amount of computing resources and time, as well as low processing efficiency.

[0005] In a first aspect, embodiments of this application provide a method for determining a subdomain, the method comprising:

[0006] The target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name are input into the large model. The business description text corresponding to the target main domain name output by the large model is obtained, and the target vector corresponding to the business description text is determined by using a preset text embedding method.

[0007] Based on the vector corresponding to each candidate main domain name that is saved in advance and the target vector, other main domain names that meet the first similarity requirement with the target main domain name are determined.

[0008] Based on the saved subdomain prefixes corresponding to the other main domains, generate the target subdomain corresponding to the target main domain.

[0009] Secondly, embodiments of this application also provide a subdomain determination device, the device comprising:

[0010] The processing module is used to input the target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name into the large model, obtain the business description text corresponding to the target main domain name output by the large model, and determine the target vector corresponding to the business description text using a preset text embedding method; and determine other main domain names that meet the first similarity requirement with the target main domain name based on the vector corresponding to each candidate main domain name that is saved in advance and the target vector.

[0011] The building module is used to generate the target subdomain corresponding to the target main domain based on the saved subdomain prefixes corresponding to the other main domains.

[0012] Thirdly, embodiments of this application also provide an electronic device, which includes at least a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the steps of the subdomain determination method as described above.

[0013] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the subdomain determination method as described above.

[0014] In this embodiment, the electronic device inputs the target domain name to be processed, the enterprise identifier corresponding to the target domain name, and the website information corresponding to the target domain name into a large model. It then obtains the business description text corresponding to the target domain name output by the large model and uses a preset text embedding method to determine the target vector corresponding to the business description text. Based on the vectors corresponding to each candidate domain name and the target vector, it determines other domain names that meet the first similarity requirement with the target domain name. Finally, based on the subdomain prefixes corresponding to the other domain names, it generates the target subdomain corresponding to the target domain name. In this embodiment, by generating business description text based on the large model, combined with the enterprise identifier, target domain name, and website information, and encoding the business description text into a target feature vector, it can not only accurately capture the business information behind the target domain name but also provide more accurate data support for subsequent subdomain generation and brute-force attacks. Furthermore, this embodiment also discloses constructing the target subdomain corresponding to the target domain name based on the subdomain prefixes corresponding to other domain names similar to the target domain name, improving the efficiency of subdomain determination and reducing computational resources and time. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram illustrating a subdomain determination process provided in an embodiment of this application;

[0017] Figure 2 This is a schematic diagram of the Bert model provided in the embodiments of this application;

[0018] Figure 3 A flowchart for subdomain brute-force attacks provided in this application embodiment;

[0019] Figure 4 A flowchart of the online process for brute-forcing subdomains provided in this application embodiment;

[0020] Figure 5 This is a schematic diagram of a subdomain determination device provided in an embodiment of this application;

[0021] Figure 6 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0023] The terms "first" and "second" in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. The term "multiple" in this application can mean at least two, for example, two, three, or more, and the embodiments of this application do not impose limitations.

[0024] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These embodiments should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that in the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solutions of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0025] Subdomain brute-force attacks are of great significance in the field of cybersecurity, as they help identify and defend against potential cyberattacks. Current subdomain brute-force methods typically require a comprehensive brute-force attack on a large dictionary, which consumes significant computational resources and time, resulting in low efficiency.

[0026] Based on this, embodiments of this application provide a method, apparatus, device, and medium for determining subdomains. In embodiments of this application, an electronic device inputs the target main domain to be processed, the enterprise identifier corresponding to the target main domain, and the website information corresponding to the target main domain into a large model, obtains the business description text corresponding to the target main domain output by the large model, and uses a preset text embedding method to determine the target vector corresponding to the business description text; based on the vector corresponding to each candidate main domain and the target vector that are saved in advance, other main domains that meet the first similarity requirement with the target main domain are determined; and based on the subdomain prefixes corresponding to the other main domains that are saved, the target subdomain corresponding to the target main domain is generated.

[0027] Example 1:

[0028] Figure 1 This application provides a schematic diagram of a subdomain determination process, which includes:

[0029] S101: Input the target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name into the large model, obtain the business description text corresponding to the target main domain name output by the large model, and determine the target vector corresponding to the business description text using a preset text embedding method.

[0030] The subdomain determination method provided in this application is applied to electronic devices, such as servers or PCs.

[0031] Subdomain brute-force is crucial in cybersecurity, aiding in the identification and defense of potential network attacks. Current subdomain brute-force methods typically require comprehensive brute-forcing of massive dictionaries, consuming significant computational resources and time, resulting in inefficiency. To address this issue, this application proposes an innovative subdomain brute-force method (DomainCF) that combines large-scale modeling and collaborative filtering.

[0032] Specifically, this application embodiment designs a domain name business encoding model based on a large model (DNBC-LM), which inputs the target main domain name, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name into the large model. The large model generates detailed business description text, and uses a text embedding method to encode these business description texts into target vectors.

[0033] Generally, in this embodiment, the enterprise identifier is the enterprise name. The website information corresponding to the target main domain can be obtained through web crawling technology or it can be pre-saved.

[0034] Specifically, the electronic device is pre-configured with an ICP database, which stores each main domain name and its corresponding enterprise identifier. The electronic device can determine the enterprise identifier corresponding to the target main domain name based on this ICP database. Then, using web crawling technology, it identifies the website information corresponding to the target main domain name. The electronic device inputs the target main domain name, enterprise identifier, and website information into a Large Language Model (LLM), which then determines and outputs the business description text corresponding to the target main domain name.

[0035] In the application embodiment, after determining the business description text corresponding to the target main domain name, the electronic device can use text embedding technology to convert the business description text into a target vector.

[0036] Specifically, the electronic device uses a preset text embedding method to determine the target vector corresponding to the business description text. This preset text embedding method can be achieved by processing the text using a preset embedding model, such as the BERT model.

[0037] For example, electronic devices input business description text into the BERT model, and then perform average pooling on the vector output from the last layer of the BERT model to obtain a 1*768 dimension vector. The BERT model converts business information in text form into vector form, which facilitates fast retrieval later.

[0038] Figure 2 This is a schematic diagram of the Bert model provided in the embodiments of this application. Figure 2 As shown, the Bert model includes an encoding layer, an attention layer, and an average pooling layer, wherein the attention layer consists of multiple multi-head self-attention modules and a feedforward neural network.

[0039] S102: Based on the vector corresponding to each candidate main domain name that is saved in advance and the target vector, determine other main domain names that meet the first similarity requirement with the target main domain name; generate the target subdomain name corresponding to the target main domain name based on the subdomain prefixes corresponding to the other main domain names that are saved in advance.

[0040] In this embodiment, the electronic device pre-stores the feature vector corresponding to each candidate main domain name and the subdomain prefix corresponding to each candidate main domain name. After determining the target vector of the target main domain name, the electronic device can filter other main domain names similar to the target main domain name based on the target vector and the stored vectors corresponding to each candidate main domain name, thereby using the subdomain prefixes corresponding to the other main domain names to construct the target subdomain corresponding to the target main domain name, and then perform subdomain brute-force attack on the target subdomain.

[0041] Specifically, in this embodiment, the electronic device calculates the first similarity between the target vector and the vector corresponding to each candidate main domain, and determines other main domains corresponding to the first similarity that meets the requirements. The first similarity that meets the requirements can be a first similarity exceeding a first preset threshold, or it can be the first similarity with the largest first preset number of values. The electronic device obtains the stored subdomain prefixes corresponding to the other main domains, and generates the target subdomain corresponding to the target main domain based on the subdomain prefixes.

[0042] The number of other primary domains identified by the electronic device can be one or more, and there is no limit to this.

[0043] In this embodiment, a business description text is generated based on a large model, combining enterprise identifiers, target main domains, and website information. This business description text is then encoded into a target feature vector. This not only accurately captures the business information behind the target main domain but also provides more precise data support for subsequent subdomain generation and brute-force attacks. Furthermore, this embodiment also discloses constructing target subdomains corresponding to the target main domain based on subdomain prefixes corresponding to other main domains similar to the target main domain. This improves the efficiency of subdomain determination and reduces computational resources and time.

[0044] Example 2:

[0045] To improve the efficiency of subdomain determination and reduce computational resources and time, based on the above embodiments, in this embodiment, the process of determining the candidate primary domain name includes:

[0046] Obtain the categories corresponding to each main domain name that are saved in advance, and the candidate vectors corresponding to the category centers of each category;

[0047] Determine the target classification center with the highest second similarity between the target vector and each candidate vector;

[0048] The main domain name belonging to the category corresponding to the target classification center is determined as the candidate main domain name.

[0049] In this embodiment of the application, when the electronic device determines other main domains that are similar to the target main domain among the candidate main domains, the candidate main domains can be all the main domains stored in the electronic device. However, since the electronic device stores many main domains, in order to improve the efficiency of subdomain determination, the electronic device can first filter out a portion of all the main domains stored as candidate main domains.

[0050] Specifically, when the electronic device saves each main domain name, it can classify each main domain name. The electronic device stores the candidate vector of the category center for each category. The electronic device first determines the target category to which the target main domain name belongs based on the candidate vector of the category center for each category and the target vector of the target main domain name, and determines the subdomains contained in the target category as candidate subdomains.

[0051] For example, an electronic device calculates the similarity between the target vector of the target domain name and the candidate vectors of the category centers of each category. It determines the second similarity between the target vector and each candidate vector, and identifies the category corresponding to the highest second similarity value as the target category to which the target domain name belongs. The electronic device can then select the N domain names with the highest scores as other domain names similar to the target domain name based on the first similarity between the target vector and the vectors corresponding to each candidate domain name in that category. This method utilizes clustering with two layers of similarity calculation. The first layer calculates similarity with K category centers, and the second layer calculates similarity with Q nodes in a given category. The total computational cost K+Q is much less than m, significantly reducing the computational burden.

[0052] In this process, when the electronic device categorizes each main domain name, it first collects a large number of enterprise identifiers. Then, it inputs these enterprise identifiers into an ICP database to obtain the main domain names registered by each enterprise identifier. The electronic device determines the vector corresponding to each main domain name. The process of determining the vector corresponding to each main domain name is consistent with the process of determining the target vector of the target main domain name in the above embodiment, and will not be repeated here. The electronic device sets a parameter K and uses the K-means algorithm to cluster the vectors corresponding to all the obtained main domain names. The number of clusters is K, resulting in K categories. For each category, the electronic device determines the mean vector of the vectors of each main domain name contained in that category as the category center vector.

[0053] Example 3:

[0054] To improve the efficiency of subdomain determination and reduce computational resources and time, based on the above embodiments, in this embodiment, determining other main domains that meet the first similarity requirement with the target main domain includes:

[0055] The cosine similarity algorithm is used to determine the first similarity between the target vector and the vector corresponding to each candidate main domain name;

[0056] The first preset number of candidate main domains with the highest first similarity are determined as the other main domains.

[0057] In this embodiment, the electronic device can use a preset similarity algorithm to determine the first similarity between the vector corresponding to the candidate main domain name and the target vector. This preset similarity algorithm can be a cosine similarity algorithm. The electronic device can determine the first similarity using the following formula:

[0058]

[0059] Where, w Av The first similarity between the target domain A and the candidate domain v; q Aq is the target vector corresponding to the target main domain. v This is the vector corresponding to the candidate main domain name.

[0060] Example 4:

[0061] To improve the efficiency of subdomain determination and reduce computational resources and time, based on the above embodiments, in this embodiment, generating the target subdomain corresponding to the target main domain based on the stored subdomain prefixes corresponding to the other main domains includes:

[0062] For each subdomain prefix corresponding to the other main domains, a preset connector is used to concatenate the subdomain prefix with the target main domain to obtain candidate subdomains;

[0063] Each candidate subdomain is used as the target subdomain.

[0064] In this embodiment, the subdomain prefix corresponding to each main domain name stored in the electronic device is determined based on the historical subdomains corresponding to each main domain name. The historical subdomains of each main domain name can be obtained from the DNS database. Since there are many historical subdomains corresponding to each main domain name, that is, many subdomain prefixes corresponding to each main domain name, in order to improve the subdomain brute-force efficiency of the electronic device, the electronic device can generate candidate subdomains based on the subdomain prefixes corresponding to other main domain names stored in the device, and then select the target subdomains for subdomain brute-force from the candidate subdomains.

[0065] Specifically, for each subdomain prefix corresponding to other main domains, a preset connector is used to concatenate the subdomain prefix with the target main domain to obtain candidate subdomains; each candidate subdomain is then used as the target subdomain. The preset connector is ".".

[0066] For example, if a subdomain of another main domain has the prefix A and the target main domain is aaa.com, then the candidate main domain determined by the electronic device is A.aaa.com.

[0067] Example 5:

[0068] To improve the efficiency of subdomain determination and reduce computational resources and time, based on the above embodiments, in this embodiment, after concatenating the subdomain prefix with the target main domain to obtain candidate subdomains, and before using each candidate subdomain as the target subdomain, the method further includes:

[0069] The score of the candidate subdomain is determined based on the first similarity to the other main domains to which the subdomain prefix belongs;

[0070] Based on the score corresponding to each candidate subdomain, determine the second preset number of candidate subdomains with the highest scores;

[0071] The candidate subdomain is updated using the second preset number of candidate subdomains with the highest scores.

[0072] In this embodiment of the application, when the electronic device filters out the target subdomain for subdomain brute-force from the candidate subdomains, the electronic device can obtain a score for each candidate subdomain based on collaborative filtering. The higher the score, the more likely the candidate subdomain is to be a subdomain of the target main domain.

[0073] Specifically, the electronic device can determine the score of the candidate subdomain based on the first similarity to the other main domains to which the subdomain prefix belongs; determine the second preset number of candidate subdomains with the highest scores based on the scores of each candidate subdomain; and update each candidate subdomain using the second preset number of candidate subdomains with the highest scores.

[0074] To improve the efficiency of subdomain determination and reduce computational resources and time, based on the above embodiments, in this embodiment, determining the score of the candidate subdomain according to the first similarity to other main domains to which the subdomain prefix belongs includes:

[0075] If the subdomain prefix belongs to at least two other main domains, then the sum of the first similarity values ​​corresponding to the at least two other main domains is determined as the score of the candidate subdomain.

[0076] If the subdomain prefix belongs to another subdomain, then the first similarity corresponding to that other subdomain is determined as the score of the candidate subdomain.

[0077] In practical applications, a subdomain prefix may belong to one or more other main domains. For example, in A.aaa.com and A.bbb.com, the subdomain prefix A belongs to both the main domain aaa.com and the main domain bbb.com. Therefore, when determining the score of each candidate subdomain, the electronic device needs to determine the score of the candidate subdomains formed by the subdomain prefix based on the first similarity score corresponding to each other main domain to which the subdomain prefix belongs.

[0078] Specifically, for each subdomain prefix, if the subdomain prefix belongs to at least two other main domains, the electronic device determines the score of the candidate subdomain by the sum of the first similarity values ​​corresponding to the at least two other main domains; if the subdomain prefix belongs to one other subdomain, the electronic device determines the score of the candidate subdomain by the first similarity value corresponding to the other subdomain.

[0079] Electronic devices can determine the score of candidate subdomains based on the following formula:

[0080]

[0081] Among them, w Av p(A,i) represents the first similarity between the target domain A and the candidate domain v; p(A,i) is the score, i is the subdomain prefix, and r is the subdomain prefix. vi This indicates whether the subdomain prefix belongs to the candidate main domain v. If it does, then r... vi =1, if not belonging, then r vi =0.

[0082] Example 6:

[0083] To improve the efficiency of subdomain determination and reduce computational resources and time, based on the above embodiments, in this embodiment, the step of inputting the target main domain to be processed, the enterprise identifier corresponding to the target main domain, and the website information corresponding to the target main domain into the large model includes:

[0084] Obtain a pre-saved prompt word template, wherein the prompt word template includes a first field for writing the main domain name, a second field for writing the enterprise identifier, a third field for writing website information, and prompt words for the prompt word big model to describe the business;

[0085] Add the target main domain name to the first field, add the enterprise identifier to the second field, and add the website information to the third field;

[0086] Input the added prompt word template into the large model.

[0087] In this embodiment, the electronic device pre-stores a prompt word template, which includes a first field for writing the main domain name, a second field for writing the enterprise identifier, a third field for writing website information, and prompt words for the prompt word big model to describe the business. The electronic device can input the target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name into the big model according to the prompt word template.

[0088] Specifically, in this embodiment, the electronic device obtains a pre-saved prompt word template, adds the target main domain name to the first field, adds the enterprise identifier to the second field, and adds the website information to the third field; the electronic device then inputs the added prompt word template into the large model.

[0089] The embodiments of this application have the following advantages over the prior art:

[0090] 1. Existing methods for subdomain mining typically require a comprehensive brute-force attack on a massive subdomain dictionary, which is not only computationally expensive and time-consuming but also inefficient. This application's embodiment, by combining a collaborative filtering algorithm, can recommend the most effective subdomains for brute-force attacks on different domains. In the offline stage, this application's embodiment uses a large model and clustering algorithm to preprocess and cluster domains; in the online stage, it only needs to calculate similarity for specific clusters, thus significantly reducing the number of subdomains requiring brute-force attacks. Through this optimization, this application's embodiment significantly reduces computational load and resource consumption, improving the efficiency of subdomain mining.

[0091] 2. Current methods fail to effectively uncover the business information behind domain names. Existing domain subdomain mining methods largely rely on static rules and dictionary matching, failing to deeply understand the business meaning behind the domain names, resulting in low accuracy and relevance of recommendation results. This application's embodiment borrows a large model, combining company name, main domain name, and website information to generate business description text, and uses the BERT model to encode this textual business description text into a vector. This innovative domain business encoding method not only accurately captures the business information behind the main domain name but also utilizes these business features for similarity calculation and clustering during the recommendation process, thereby recommending more relevant and effective subdomains. This business information-based recommendation mechanism significantly improves the success rate and accuracy of subdomain brute-force attacks.

[0092] Example 7:

[0093] Based on the above embodiments, in practice, the embodiments of this application can be implemented and applied in multiple scenarios, and specific embodiments are listed below:

[0094] 1. Attack surface management.

[0095] Subdomain brute-force attacks play a crucial role in attack surface management. By systematically generating and testing a large number of potential subdomains, this technique helps security experts gain a comprehensive understanding of an organization's internet assets. Many organizations may host various services, including websites, APIs, and mail servers, under different subdomains. These subdomains often become potential security vulnerabilities due to their inconspicuousness or being overlooked. Subdomain brute-force attacks can identify those undocumented or ignored subdomains, thus creating a complete attack surface map. Security teams can then evaluate these subdomains to ensure each one is adequately protected, preventing security incidents caused by oversights. This not only improves overall security but also prevents attackers from discovering sensitive information or vulnerabilities within the organization through subdomains, thus hindering further attacks.

[0096] Step 1: Collect the main domain name, ICP data, and website crawler data.

[0097] Step 2: parse and encode each main domain name to obtain a vector for each main domain name. Then, cluster the vectors to obtain the cluster centers and their corresponding vectors. Store the cluster centers and their corresponding vectors, as well as the vectors of the main domain names, in the database.

[0098] Step 3: For each main domain that needs to be brute-forced, use the method provided in the above embodiments to determine the target subdomain corresponding to the main domain, and brute-force the target subdomain to obtain the final existing subdomain.

[0099] Step 4: Provide the company's security experts with the actual subdomains. The experts will analyze any undocumented or overlooked subdomains to create a complete attack surface map. Evaluate the subdomains to ensure each one receives appropriate security protection and prevent security incidents caused by oversights.

[0100] 2. Penetration testing and vulnerability discovery

[0101] Subdomain brute-force techniques are indispensable tools in penetration testing and vulnerability discovery. Through subdomain brute-force, penetration testers can discover hidden subdomains within the target system. These subdomains may be running different applications or services, providing attackers with more potential entry points. For example, a forgotten development or testing environment may lack robust security measures, becoming an easy entry point for attackers. After discovering these subdomains, penetration testers can conduct further security testing to look for potential vulnerabilities, such as outdated services, default passwords, and publicly available sensitive information. These methods allow for the early detection and remediation of security vulnerabilities, preventing attackers from using these subdomains for unauthorized intrusion and data theft, ultimately improving the overall security of the system.

[0102] Step 1: Collect the main domain name, ICP data, and website crawler data.

[0103] Step 2: parse and encode each main domain name to obtain a vector for each main domain name. Then, cluster the vectors to obtain the cluster centers and their corresponding vectors. Store the cluster centers and their corresponding vectors, as well as the vectors of the main domain names, in the database.

[0104] Step 3: For each main domain that needs to be brute-forced, use the method provided in the above embodiments to determine the target subdomain corresponding to the main domain, and brute-force the target subdomain to obtain the final existing subdomain.

[0105] Step Four: By brute-forcing subdomains, penetration testers can discover hidden subdomains within the target system. Once discovered, these subdomains can be used for further security testing to identify potential vulnerabilities, such as outdated services, default passwords, or publicly available sensitive information. Afterward, the penetration testers will disable the asset, patch the vulnerabilities, and improve the system's security.

[0106] Figure 3 This is a flowchart of subdomain brute-force provided in an embodiment of this application, such as... Figure 3 As shown, the process includes an offline process and an online process. The offline process includes encoding a large number of main domains into a feature vector. The online process includes inputting the main domain, obtaining the feature vector of the main domain, calculating the target subdomain corresponding to the main domain, and performing subdomain brute-force attack on the target subdomain.

[0107] exist Figure 3 On this basis, Figure 4 The flowchart of the subdomain brute-force online process provided in this application embodiment is as follows: Figure 4 As shown, inputting a main domain name A yields its feature vector A. This feature vector is then used to calculate similarity with each cluster center to determine the cluster Ki to which the main domain belongs. All feature vectors within cluster Ki are extracted and their similarity is calculated again with the feature vector A of main domain name A. The N main domains with the highest scores are selected as similar domains to main domain name A. Subdomain prefixes corresponding to these N main domains are retrieved from the DNS database and used to form candidate subdomains. The target subdomain is then selected from the candidate subdomains and subjected to subdomain brute-force attack.

[0108] Example 8:

[0109] Based on the above embodiments, Figure 5 This application provides a schematic diagram of a subdomain determination device, which includes:

[0110] The processing module 501 is used to input the target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name into the large model, obtain the business description text corresponding to the target main domain name output by the large model, and determine the target vector corresponding to the business description text using a preset text embedding method; and determine other main domain names that meet the first similarity requirement with the target main domain name based on the vector corresponding to each candidate main domain name that is saved in advance and the target vector.

[0111] The construction module 502 is used to generate the target subdomain corresponding to the target main domain based on the saved subdomain prefixes corresponding to the other main domains.

[0112] In one possible implementation, the processing module 501 is specifically used to obtain the category corresponding to each main domain name that is pre-saved, and the candidate vector corresponding to the category center of each category; determine each target category center whose second similarity with each candidate vector meets the requirements; and determine the main domain name belonging to the category corresponding to each target category center as the candidate main domain name.

[0113] In one possible implementation, the processing module 501 is specifically used to use a cosine similarity algorithm to determine the first similarity between the target vector and the vector corresponding to each candidate main domain name; and to determine the first preset number of candidate main domain names with the highest first similarity as the other main domain names.

[0114] In one possible implementation, the construction module 502 is specifically used to concatenate the subdomain prefix with the target main domain using a preset connector for each subdomain prefix corresponding to the other main domain to obtain candidate subdomains; and to use each candidate subdomain as the target subdomain.

[0115] In one possible implementation, the construction module 502 is further configured to determine the score of the candidate subdomain based on the first similarity to the other main domains to which the subdomain prefix belongs; determine a second preset number of candidate subdomains with the highest scores based on the scores corresponding to each candidate subdomain; and update each candidate subdomain using the second preset number of candidate subdomains with the highest scores.

[0116] In one possible implementation, the construction module 502 is specifically configured to: if the subdomain prefix belongs to at least two other main domains, determine the sum of the first similarity values ​​corresponding to the at least two other main domains as the score of the candidate subdomain; if the subdomain prefix belongs to one other subdomain, determine the first similarity value corresponding to the one other subdomain as the score of the candidate subdomain.

[0117] In one possible implementation, the processing module 501 is specifically used to obtain a pre-saved prompt word template, wherein the prompt word template includes a first field for writing the main domain name, a second field for writing the enterprise identifier, a third field for writing website information, and prompt words for the prompt word big model to describe the business; add the target main domain name to the first field, add the enterprise identifier to the second field, add the website information to the third field; and input the added prompt word template into the big model.

[0118] Example 9:

[0119] Based on the above embodiments, this application also provides an electronic device. Figure 6 This application provides a schematic diagram of an electronic device structure, such as... Figure 6 As shown, it includes: processor 601, communication interface 602, memory 603 and communication bus 604, wherein processor 601, communication interface 602 and memory 603 communicate with each other through communication bus 604.

[0120] The memory 603 stores a computer program, which, when executed by the processor 601, causes the processor 601 to perform the following steps:

[0121] The target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name are input into the large model. The business description text corresponding to the target main domain name output by the large model is obtained, and the target vector corresponding to the business description text is determined by using a preset text embedding method.

[0122] Based on the vector corresponding to each candidate main domain name that is saved in advance and the target vector, other main domain names that meet the first similarity requirement with the target main domain name are determined.

[0123] Based on the saved subdomain prefixes corresponding to the other main domains, generate the target subdomain corresponding to the target main domain.

[0124] In one possible implementation, the process of determining the candidate primary domain name includes:

[0125] Obtain the categories corresponding to each main domain name that are saved in advance, and the candidate vectors corresponding to the category centers of each category;

[0126] Determine each target classification center whose second similarity to each candidate vector satisfies the requirement;

[0127] The main domain name belonging to the category corresponding to each target classification center is determined as the candidate main domain name.

[0128] In one possible implementation, determining other main domains that meet the first similarity requirement with the target main domain includes:

[0129] The cosine similarity algorithm is used to determine the first similarity between the target vector and the vector corresponding to each candidate main domain name;

[0130] The first preset number of candidate main domains with the highest first similarity are determined as the other main domains.

[0131] In one possible implementation, generating the target subdomain corresponding to the target main domain based on the stored subdomain prefixes corresponding to the other main domains includes:

[0132] For each subdomain prefix corresponding to the other main domains, a preset connector is used to concatenate the subdomain prefix with the target main domain to obtain candidate subdomains;

[0133] Each candidate subdomain is used as the target subdomain.

[0134] In one possible implementation, after concatenating the subdomain prefix with the target main domain to obtain candidate subdomains, and before using each candidate subdomain as the target subdomain, the method further includes:

[0135] The score of the candidate subdomain is determined based on the first similarity to the other main domains to which the subdomain prefix belongs;

[0136] Based on the score corresponding to each candidate subdomain, determine the second preset number of candidate subdomains with the highest scores;

[0137] The candidate subdomain is updated using the second preset number of candidate subdomains with the highest scores.

[0138] In one possible implementation, determining the score of the candidate subdomain based on the first similarity to other main domains to which the subdomain prefix belongs includes:

[0139] If the subdomain prefix belongs to at least two other main domains, then the sum of the first similarity values ​​corresponding to the at least two other main domains is determined as the score of the candidate subdomain.

[0140] If the subdomain prefix belongs to another subdomain, then the first similarity corresponding to that other subdomain is determined as the score of the candidate subdomain.

[0141] In one possible implementation, inputting the target domain name to be processed, the enterprise identifier corresponding to the target domain name, and the website information corresponding to the target domain name into the large model includes:

[0142] Obtain a pre-saved prompt word template, wherein the prompt word template includes a first field for writing the main domain name, a second field for writing the enterprise identifier, a third field for writing website information, and prompt words for the prompt word big model to describe the business;

[0143] Add the target main domain name to the first field, add the enterprise identifier to the second field, and add the website information to the third field;

[0144] Input the added prompt word template into the large model.

[0145] Since the principle of the above-mentioned electronic device in solving the problem is similar to that of the subdomain determination method, the implementation of the above-mentioned electronic device can be found in the embodiments of the method, and repeated parts will not be described again.

[0146] The communication bus mentioned in the above-mentioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus. Communication interface 602 is used for communication between the above-mentioned electronic device and other devices. The memory can include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.

[0147] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0148] Example 10:

[0149] Based on the above embodiments, this invention also provides a computer-readable storage medium storing a computer program executable by a processor. When the program is run on the processor, it causes the processor to perform the following steps:

[0150] The target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name are input into the large model. The business description text corresponding to the target main domain name output by the large model is obtained, and the target vector corresponding to the business description text is determined by using a preset text embedding method.

[0151] Based on the vector corresponding to each candidate main domain name that is saved in advance and the target vector, other main domain names that meet the first similarity requirement with the target main domain name are determined.

[0152] Based on the saved subdomain prefixes corresponding to the other main domains, generate the target subdomain corresponding to the target main domain.

[0153] In one possible implementation, the process of determining the candidate primary domain name includes:

[0154] Obtain the categories corresponding to each main domain name that are saved in advance, and the candidate vectors corresponding to the category centers of each category;

[0155] Determine each target classification center whose second similarity to each candidate vector satisfies the requirement;

[0156] The main domain name belonging to the category corresponding to each target classification center is determined as the candidate main domain name.

[0157] In one possible implementation, determining other main domains that meet the first similarity requirement with the target main domain includes:

[0158] The cosine similarity algorithm is used to determine the first similarity between the target vector and the vector corresponding to each candidate main domain name;

[0159] The first preset number of candidate main domains with the highest first similarity are determined as the other main domains.

[0160] In one possible implementation, generating the target subdomain corresponding to the target main domain based on the stored subdomain prefixes corresponding to the other main domains includes:

[0161] For each subdomain prefix corresponding to the other main domains, a preset connector is used to concatenate the subdomain prefix with the target main domain to obtain candidate subdomains;

[0162] Each candidate subdomain is used as the target subdomain.

[0163] In one possible implementation, after concatenating the subdomain prefix with the target main domain to obtain candidate subdomains, and before using each candidate subdomain as the target subdomain, the method further includes:

[0164] The score of the candidate subdomain is determined based on the first similarity to the other main domains to which the subdomain prefix belongs;

[0165] Based on the score corresponding to each candidate subdomain, determine the second preset number of candidate subdomains with the highest scores;

[0166] The candidate subdomain is updated using the second preset number of candidate subdomains with the highest scores.

[0167] In one possible implementation, determining the score of the candidate subdomain based on the first similarity to other main domains to which the subdomain prefix belongs includes:

[0168] If the subdomain prefix belongs to at least two other main domains, then the sum of the first similarity values ​​corresponding to the at least two other main domains is determined as the score of the candidate subdomain.

[0169] If the subdomain prefix belongs to another subdomain, then the first similarity corresponding to that other subdomain is determined as the score of the candidate subdomain.

[0170] In one possible implementation, inputting the target domain name to be processed, the enterprise identifier corresponding to the target domain name, and the website information corresponding to the target domain name into the large model includes:

[0171] Obtain a pre-saved prompt word template, wherein the prompt word template includes a first field for writing the main domain name, a second field for writing the enterprise identifier, a third field for writing website information, and prompt words for the prompt word big model to describe the business;

[0172] Add the target main domain name to the first field, add the enterprise identifier to the second field, and add the website information to the third field;

[0173] Input the added prompt word template into the large model.

[0174] Since the principle of solving the problem by the above computer program product is similar to that of the subdomain determination method, the implementation of the above computer program product can be referred to the implementation of the method, and the repeated parts will not be repeated.

[0175] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0176] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0177] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0178] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0179] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for determining subdomains, characterized in that, The method includes: The target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name are input into the large model. The business description text corresponding to the target main domain name output by the large model is obtained, and the target vector corresponding to the business description text is determined by using a preset text embedding method. Based on the vector corresponding to each candidate main domain name that is saved in advance and the target vector, other main domain names that meet the first similarity requirement with the target main domain name are determined. Based on the saved subdomain prefixes corresponding to the other main domains, generate the target subdomain corresponding to the target main domain; The process of determining the candidate primary domain name includes: Obtain the categories corresponding to each main domain name that are saved in advance, and the candidate vectors corresponding to the category centers of each category; Determine the target classification center with the highest second similarity between the target vector and each candidate vector; The main domain name belonging to the category corresponding to the target classification center is determined as the candidate main domain name; The other main domains that meet the first similarity requirement with the target main domain include: The cosine similarity algorithm is used to determine the first similarity between the target vector and the vector corresponding to each candidate main domain name; The first preset number of candidate main domains with the highest first similarity are determined as the other main domains.

2. The method according to claim 1, characterized in that, The step of generating the target subdomain corresponding to the target main domain based on the saved subdomain prefixes corresponding to the other main domains includes: For each subdomain prefix corresponding to the other main domains, a preset connector is used to concatenate the subdomain prefix with the target main domain to obtain candidate subdomains; Each candidate subdomain is used as the target subdomain.

3. The method according to claim 2, characterized in that, After concatenating the subdomain prefix with the target main domain to obtain candidate subdomains, and before using each candidate subdomain as the target subdomain, the method further includes: The score of the candidate subdomain is determined based on the first similarity to the other main domains to which the subdomain prefix belongs; Based on the score corresponding to each candidate subdomain, determine the second preset number of candidate subdomains with the highest scores; The candidate subdomain is updated using the second preset number of candidate subdomains with the highest scores.

4. The method according to claim 3, characterized in that, The step of determining the score of the candidate subdomain based on the first similarity to other main domains to which the subdomain prefix belongs includes: If the subdomain prefix belongs to at least two other main domains, then the sum of the first similarity values ​​corresponding to the at least two other main domains is determined as the score of the candidate subdomain. If the subdomain prefix belongs to another subdomain, then the first similarity corresponding to that other subdomain is determined as the score of the candidate subdomain.

5. The method according to claim 1, characterized in that, The step of inputting the target domain name to be processed, the enterprise identifier corresponding to the target domain name, and the website information corresponding to the target domain name into the large model includes: Obtain a pre-saved prompt word template, wherein the prompt word template includes a first field for writing the main domain name, a second field for writing the enterprise identifier, a third field for writing website information, and prompt words for the prompt word big model to describe the business; Add the target main domain name to the first field, add the enterprise identifier to the second field, and add the website information to the third field; Input the added prompt word template into the large model.

6. A subdomain determination device, characterized in that, The device includes: The processing module is used to input the target main domain name to be processed, the enterprise identifier corresponding to the target main domain name, and the website information corresponding to the target main domain name into the large model, obtain the business description text corresponding to the target main domain name output by the large model, and determine the target vector corresponding to the business description text using a preset text embedding method; and determine other main domain names that meet the first similarity requirement with the target main domain name based on the vector corresponding to each candidate main domain name that is saved in advance and the target vector. The building module is used to generate the target subdomain corresponding to the target main domain based on the saved subdomain prefixes corresponding to the other main domains; Specifically, the processing module is used to obtain the category corresponding to each main domain name and the candidate vector corresponding to the category center of each category; determine the target category center with the highest second similarity between the target vector and each candidate vector; and determine the main domain name belonging to the category corresponding to the target category center as the candidate main domain name. Specifically, the processing module is used to use a cosine similarity algorithm to determine the first similarity between the target vector and the vector corresponding to each candidate main domain name; and to determine the first preset number of candidate main domain names with the highest first similarity as the other main domain names.

7. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, wherein the processor is used to implement the steps of the subdomain determination method as described in any one of claims 1-5 when executing a computer program stored in the memory.

8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the subdomain determination method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Network data acquisition method, equipment and related equipment thereof

    CN111556077A

  • Domain name detection method and device, equipment and storage medium

    CN114818689A