IP address owner industry classification system for unlabeled data scene
The system addresses the challenge of accurate and extensive IP address industry classification by using PTR and DNS records, web page analysis, and deep learning models to classify IP addresses without prior labeling, achieving high precision and recall rates.
Patent Information
- Application Number
- CN202510302784.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-15
AI Technical Summary
The existing IP address industry classification method cannot achieve wide coverage and high accuracy in the unlabeled data scenario, and the existing technology cannot effectively identify the industry attributes of critical infrastructure in the Internet, especially the IP addresses of web assets.
The IP address domain name reverse search module, organization-based industry classification module and web page-based industry classification module are adopted, combined with large language model and noise-based learning methods, domain names are obtained through PTR records and Passive DNS records, and the organization and web page text information are used for industry classification, label noise is corrected, and high accuracy and wide coverage are achieved without pre-labeling data.
It has achieved high accuracy and wide coverage of IP addresses on the entire network, identified industries with more than 1 billion IP addresses, covered more than 20,000 organizations, and reached a high level of classification accuracy and recall, meeting the needs of network security and network management.
Smart Images

Figure CN120316639A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of network security and network mapping, and particularly to an IP address owner industry classification system for unlabeled data scenarios. Background Art
[0002] With the development of the Internet, the number of assets in the cyberspace is increasing continuously. IP addresses play an important role in network security and network management. By classifying the industries of IP address owners, the structure of the Internet can be deeply understood, which can help identify critical infrastructures as well as the sources of network security threats and attacks, providing important basis and support for network security management. From the perspective of security departments and security companies, how to discover the assets owned by critical infrastructure industries in the Internet and conduct corresponding protection is also a key issue. The IP addresses on which Web assets are deployed are particularly important, that is, the IP addresses that have opened domain name services. These addresses account for a large proportion among IP addresses and are also the key targets of attackers. Relevant security reports show that the attacks on Web applications and APIs are increasing month by month. Therefore, from the perspectives of network management and network attack and defense, a method for large-scale identifying the assets belonging to critical infrastructure industries in the Internet is needed.
[0003] In order to identify the industry attribution of Internet assets, in recent years, the academic and industrial communities have proposed a series of methods to identify the industries to which web assets belong from different perspectives. However, some of the existing methods focus on identifying the IP addresses of specific organizations and do not have a greater coverage of the perception of the attributes of all network IP address owners. On the other hand, some of the current work can only achieve organization identification and industry classification at the AS level, with insufficient granularity. Since ASs and IP address owners are inconsistent, the organization of the IP address cannot be accurately determined. At the same time, the accuracy of the current work in industry classification is not high, and there are false positives or false negatives. And for large-scale classification of the assets belonging to critical infrastructure industries in the Internet, a wide-coverage and accurate industry classification system is needed. However, the existing methods cannot meet these requirements. In addition, the classification work in the field of artificial intelligence is usually oriented to recommendation systems, and the special requirements in the field of network security require corresponding adjustment of the classification categories to meet the needs of identifying critical infrastructure industries. Currently, the amount of labeled data in this regard is small, and it is required that the industry classification system can achieve classification effects with a small amount or lack of labeled data. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems in the related technologies to some extent.
[0005] The present invention proposes an IP address owner industry classification system for unlabeled data scenarios, which makes up for the deficiencies of existing methods and meets the special needs of industry classification in the field of network security. The present invention can classify the IP addresses with deployed Web assets without pre-labeling data and can cover a wide range of IP addresses in the entire network, with good accuracy and wide coverage.
[0006] Another object of the present invention is to propose an IP address owner industry classification method for unlabeled data scenarios.
[0007] To achieve the above object, on the one hand, the present invention proposes an IP address owner industry classification system for unlabeled data scenarios, including:
[0008] An IP address domain name reverse lookup module, configured to obtain the domain name corresponding to the Web asset deployed on the IP address based on the PTR record and the Passive DNS record;
[0009] An organization-based industry classification module, configured to identify the organization of the domain name owner based on the consistency result of the organization of the domain name owner and the organization of the IP address owner and the corresponding relationship between the organization and the industry, and determine the industry to which the IP address belongs based on the organization-based industry classification algorithm;
[0010] A web-based industry classification module, configured to perform industry classification on the industry to which the IP address belongs based on the web-based industry classification algorithm and according to the web page text to obtain an industry classification result.
[0011] The IP address owner industry classification system for unlabeled data scenarios according to the embodiments of the present invention may also have the following additional technical features:
[0012] In an embodiment of the present invention, the IP address domain name reverse lookup module is further configured to:
[0013] Scan the PTR record to cover IPv4 and IPv6 addresses by actively sending requests to the DNS resolver, and based on the obtained list of IPv6 response addresses, send requests to the addresses therein to obtain the PTR record query response probability;
[0014] Use the Passive DNS record to supplement the domain name information corresponding to the IP address without a configured PTR record;
[0015] Obtain the fully qualified domain name retrieved from the DNS PTR record and the Passive DNS record, and extract the second-level domain name from the fully qualified domain name for classification work.
[0016] In one embodiment of the present invention, the industry classification module based on organizations includes an organization identification module, a large model-based industry classification module, and a data label noise correction module; wherein, the organization identification module is used for:
[0017] Extract the organization name of the registrant from the registrant field of the WHOIS record through the open API of the domain name registrar, use the extracted organization name as the clustering keyword, and cluster the domain names belonging to the same organization into the same cluster;
[0018] Obtain the SSL / TLS certificate by requesting port 443 of the main domain name, extract the key fields in the certificate. For organization validation certificates and extended validation certificates, the Organization field represents the organization holding the issued certificate and is used as the first keyword for organization clustering; distinguish CDN and cloud service providers through manual identification and large language model judgment, and construct a list of secondary domain names of CDN providers. If the secondary domain name falls into the list of secondary domain names of CDN providers, it is retained in the CDN cluster, otherwise it is excluded from the cluster;
[0019] After deleting the secondary domain names belonging to CDN providers, cluster the remaining domain names based on the Common Name and Subject Alternative Name fields. The union of the secondary domain names in these two fields is used as the second feature for clustering; based on the two data sources of WHOIS and certificates, cluster the domain names into their respective clusters according to the organization name. For data lacking the organization field, cluster based on the Common Name and Subject Alternative Name fields.
[0020] In one embodiment of the present invention, the large model-based industry classification module is further used for:
[0021] Crawl the organization description from online encyclopedias and extract the industry to which it belongs through natural language processing techniques;
[0022] Construct a pre-trained large language model, and input the organization description crawled from online encyclopedias and the definitions of each industry into the pre-trained large language model; wherein, provide examples for the large language model through the chain of thought prompting technique, and the chain of thought includes summarizing the organization description, determining the main business, checking the industry definition, and providing reasons for matching or not matching the industry;
[0023] Use the pre-trained large language model to select from all categories and give the potential categories to which the organization may belong, and verify whether the organization meets the proposed potential categories.
[0024] In one embodiment of the present invention, the data label noise correction module is further used for:
[0025] Use the noise screening algorithm in noisy learning to regard the labels output by the large language model as the true values, and input them as training data into the long short-term memory network to fit the labels and calculate the loss; use the Gaussian mixture model to cluster the loss distribution of the training data to distinguish clean data and noisy data;
[0026] Train LSTM-1 and LSTM-2 with the labels before and after the secondary confirmation of the large language model respectively. Among them, each LSTM independently uses GMM to model the training noise to distinguish noisy data and clean data; for noisy data, when re-labeling, use LSTM-1 and LSTM-2 for joint labeling, and process the positive and negative labels separately. In each round, LSTM-1 first models and corrects the positive label loss, and then models and corrects the negative label loss; LSTM-2 is the opposite;
[0027] After multiple rounds of iteration, LSTM-1 and LSTM-2 respectively perform multiple loss modeling and correction on the positive and negative labels. After the last round of iteration, the labels corrected by LSTM-1 and LSTM-2 are integrated through a voting mechanism to obtain the finally optimized labels.
[0028] In an embodiment of the present invention, the web-based industry classification module includes a web page text acquisition module and a web page text classification module based on noisy learning; wherein, the web page text acquisition module is used for:
[0029] Input the domain name obtained by the IP address domain name reverse lookup module, access the home page of the domain name through the browser, and obtain the web page title, the home page text content, and the meta-description element attributes in the HTML; recursively track the links containing specific keywords in the home page, and obtain the titles, texts, and meta-description attributes in these sub-pages, and splice all the obtained text contents to form the original text data;
[0030] Unify the translation of the original text data into English, and use the TF-IDF method to calculate the importance of the words in the text, and sum the importance of all the words in the sentence to obtain the importance of the sentence;
[0031] Sort the sentences in the text data corresponding to each domain name from largest to smallest according to the importance, and splice the sorted sentences in turn until the set text length limit is exceeded to obtain the processed web page text data.
[0032] In an embodiment of the present invention, the web page text classification module based on noisy learning is further used for:
[0033] Use the processed web page text data as input data to train the neural network, and use the output of the organization-based industry classification algorithm as the annotation data;
[0034] LSTM is used to distinguish the noise data in the input data, and LSTM and Roberta are used to label them together. The corrected data is used to train the LSTM and Roberta models. For the labeled data, a random data in the category with the least amount of data is selected and concatenated with another data that does not belong to the category to form new input data. The union of the two labels is taken as the label of the new data until the difference in data volume between the category with the largest amount of data and the category with the least amount of data is less than the set threshold.
[0035] After multiple rounds of iterative training, the LSTM and Roberta models are gradually optimized, and the trained Roberta model is deployed in practical applications for industry classification of web page text data.
[0036] To achieve the above object, the present invention proposes, on the other hand, a method for classifying IP address owners by industry for unlabeled data scenarios, comprising:
[0037] Obtain the domain name corresponding to the Web asset deployed on the IP address based on the PTR record and Passive DNS record;
[0038] According to the consistency results of the domain name owner's organization and the IP address owner's organization and the correspondence between the organization and the industry, the domain name owner's organization is identified based on the organization's industry classification algorithm to determine the industry to which the IP address belongs;
[0039] Based on the industry classification algorithm of the web page and according to the web page text, the industry to which the IP address belongs is classified to obtain the industry classification result.
[0040] The IP address owner industry classification system and method for unlabeled data scenarios of the embodiments of the present invention are applicable to the fields of network security and network management, which are reflected in the fact that the designed system needs to meet the requirements of accuracy, accurately determine the industry to which the network asset owner at the IP address level belongs, unlabeled, no need to collect labeled data in advance, and wide coverage, to solve the needs of IP address industry classification problems with different characteristics of different organizations. Due to the complex relationships, mutual constraints and even contradictions between these conditions, it is easy to lose sight of one thing while focusing on another. The present invention combines the large language model annotation method and the classification method with noise learning through in-depth analysis of the actual situation of the problem to be solved and the relationship between the requirements of each part, so as to comprehensively meet various requirements.
[0041] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:
[0043] Figure 1 is a structural diagram of an IP address owner industry classification system for an unlabeled data scenario according to an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of a multi-label data correction model for dual-model collaborative training according to an embodiment of the present invention;
[0045] Figure 3 is a schematic diagram of a web page text classification model based on noisy learning according to an embodiment of the present invention;
[0046] Figure 4 is a flowchart of an IP address owner industry classification method for an unlabeled data scenario according to an embodiment of the present invention. Detailed Embodiments
[0047] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0048] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] The IP address owner industry classification system and method for an unlabeled data scenario according to an embodiment of the present invention will be described below with reference to the accompanying drawings.
[0050] Figure 1 is a structural diagram of an IP address owner industry classification system for an unlabeled data scenario according to an embodiment of the present invention, as Figure 1 shown, including:
[0051] An IP address domain name reverse lookup module, configured to obtain the domain name corresponding to the Web asset deployed on the IP address based on the PTR record and the Passive DNS record;
[0052] An organization-based industry classification module, configured to identify the organization of the domain name owner based on the consistency result of the organization of the domain name owner and the organization of the IP address owner and the corresponding relationship between the organization and the industry, and determine the industry to which the IP address belongs based on the organization-based industry classification algorithm;
[0053] The web page-based industry classification module is used to classify the industry to which the IP address belongs based on the web page-based industry classification algorithm and according to the web page text to obtain the industry classification result.
[0054] It can be understood that the present invention determines the web assets deployed on the IP address through the IP address domain name reverse query module, uses the industry classification algorithm based on the organization to realize the automatic acquisition of label data, and uses this data as training data to train the industry classification algorithm based on the web page to realize efficient and accurate industry classification of the IP addresses of the web assets deployed on the whole network. Figure 1 shown.
[0055] In one embodiment of the present invention, the IP address domain name reverse lookup module uses the PTR record and the Passive DNS record to obtain the domain name corresponding to the Web asset deployed on the IP address.
[0056] The PTR record is configured by the IP address owner and stored in a reverse zone file. The record can be queried through a reverse DNS request and provides a correspondence between the IP address and the domain name. First, the present invention scans the PTR record by actively initiating a request to the DNS resolver, covering both IPv4 and IPv6 addresses. Among them, IPv4 addresses can be scanned across the entire network, while the IPv6 address space is too large to be scanned across the entire network. Since active addresses only account for a small portion of IPv6 addresses, the present invention sends a request to the addresses in the IPv6 response address list based on the IPv6 response address list obtained in previous work. Since the addresses therein are verified to respond to ICMP requests, querying these addresses can obtain a higher probability of a PTR record query response.
[0057] Since PTR records need to be actively configured by the IP address owner, there is a phenomenon that some IP addresses with web assets deployed do not have PTR records configured. Passive DNS records can supplement the domain name information corresponding to the IP address in this case. However, due to the existence of CDN, the results of Passive DNS queries may not necessarily maintain the consistency between the IP address owner attribute and the domain name owner attribute. Depending on the implementation method of CDN, the following two cases are discussed. First, CDN implemented based on CNAME in DNS queries. Its working mode can be roughly summarized as: original user domain name -> CNAME CDN domain name -> CDN IP address. In this case, since what is recorded in Passive DNS is the result of the last A record or AAAA record query, that is, CNAME CDN domain name -> CDN IP address, the owner attribute of the IP address and the owner attribute of the domain name are consistent. Second, CDN implemented based on anycast. Its working mode can be roughly summarized as user domain name -> CDN anycast address. At this time, the result obtained by Passive DNS reverse lookup is the user domain name corresponding to the CDN anycast address. The owner attribute of the IP address is the CDN manufacturer, while the owner attribute of the user domain name is the CDN user, resulting in the query result at this time breaking the consistency between the IP address and the domain name owner attributes. Based on the above observations, after removing the anycast address in the present invention, Passive DNS records can be used to determine the correspondence between the IP address and the domain name to supplement the results obtained through PTR records.
[0058] The domain names retrieved from DNS PTR records and Passive DNS records are fully qualified domain names (FQDNs), which represent the complete domain names of specific hosts. Since it is challenging and labor-intensive to identify each FQDN in the industry, to meet the need for efficiency, the present invention extracts the second-level domain name from the FQDN, that is, the main domain name registered by the user, and conducts subsequent classification work at this level.
[0059] In an embodiment of the present invention, according to the results of the IP address domain name reverse lookup module, it can be known that the organization of the domain name owner is consistent with the organization of the IP address owner, and there is a relatively clear corresponding relationship between the organization and the industry. Therefore, the industry classification module based on the organization can determine the industry classification to which the IP address belongs by identifying the organization of the domain name owner, mainly including an organization identification module, an industry classification module, and a data correction module. The working processes of each module are introduced in detail below.
[0060] The organization identification module obtains the domain name owner organization through WHOIS information and certificates, and clusters the domain names of the same organization into the same cluster.
[0061] Domain name registrars disclose the registrant's organizational information in fields such as registrant in the WHOIS record. This information can be obtained through the registrar's open API, and the organizational names extracted from fields such as registrant are used as clustering keywords.
[0062] Due to GDPR and privacy considerations, the organizational information corresponding to a domain name may not necessarily be retrievable from WHOIS, and the present invention supplements it with the information in the certificate. The certificate is obtained by requesting port 443 of the main domain name. The certificate contains basic information about the controlling organization of the domain name. Important fields include Common Name, Organization, and SubjectAlternative Name in the extension. The Common Name field represents the single server name protected by the certificate, while the SubjectAlternative Name field represents an enumerated list of server names authorized by the certificate. The Organization field represents the organization that holds the issued certificate. For organization validation certificates and extended validation certificates, this field is verified by the issuing authority to ensure its reliability and credibility related to the identified organization. However, for domain validation certificates and self-signed certificates, this field is empty or not verified. Since the organization field clearly indicates the owner organization, this module uses it as the first keyword for organization clustering. However, some domain names hosted on a content delivery network (CDN) may use certificates issued by the CDN provider, resulting in this field incorrectly indicating the CDN organization rather than the actual domain name owner. To address this misalignment issue, the present invention distinguishes between CDNs and cloud service providers through manual identification based on the descriptions in online encyclopedias. Due to the large number of organizations, the judgment of large language models is provided as a reference for manual work. After distinguishing the CDN cluster, a list of secondary domain names of CDN providers is constructed by calculating the occurrence frequencies of secondary domain names within the cluster. For the CDN cluster, if a secondary domain name falls into the list of secondary domain names of CDN providers, it is retained in the cluster; otherwise, it is excluded from the cluster. For the remaining data lacking the organization field, the present invention further clusters the domain names based on the Common Name and Subject Alternative Name fields. After deleting the domain names in the CDN secondary domain name list, the union of the secondary domain names in the above two fields is used as the second feature for clustering.
[0063] Through the above two data sources, the domain names are clustered into their respective clusters according to the organization.
[0064] In one embodiment of the present invention, for the industry classification module based on the large model, after determining the organization, the present invention crawls available organization descriptions from online encyclopedias for industry classification. Although some descriptions in the encyclopedia directly indicate the relevant industries, most descriptions only outline the main business activities. It is necessary to further obtain the industry to which it belongs through natural language processing. The complexity of natural language processing tasks makes it infeasible to manually process a large amount of data. Therefore, the present invention designs a zero-shot organization-industry classification system based on the large model.
[0065] Large language models show superior performance in dealing with complex natural language processing tasks. The pre-training process on a wide range of general corpora and the supervised fine-tuning process with human annotations enable large language models to obtain natural language processing capabilities. These models can be adapted to the specific tasks of the present invention through zero-shot learning, meeting the requirements of no annotation and accuracy. The present invention proposes an industry classification system based on the large model to determine the industry of the organization. There is no requirement for the selection of specific models. According to the actual business needs, existing open-source and closed-source models can be embedded into the system of the present invention by writing prompt words. The organization descriptions crawled from the online encyclopedia and the definitions of each industry in the category are input into the large language model. To enhance the logical thinking ability of the large language model and improve the classification effect, the present invention uses the chain-of-thought prompting technique to provide examples for the large model. The chain of thought includes summarizing the organization description, determining the main business, checking the industry definition, and providing reasons for matching or not matching the industry, and finally giving the industry classification. To address the potential uncertainty and inconsistency in the large model's predictions, the present invention introduces a double confirmation mechanism. Initially, the large model is prompted to select from all categories and give the potential categories to which the organization may belong. Subsequently, the large model is repeatedly asked to verify whether the organization actually conforms to the potential categories it recommends. If the potential category is confirmed at least once in the second stage, it will be retained in the final result.
[0066] Data label noise correction module. Since there is still a certain amount of noise in the data labels annotated by the large model, the present invention uses the noise screening algorithm in learning with noise, takes the labels output by the large model as the ground truth, constructs training data and inputs it into the long short-term memory network (LSTM). This network fits the labels, calculates the loss, and uses the Gaussian mixture model (GMM) to cluster the distribution of the training loss of the data to screen out the noise data and optimize it. Since it is a multi-label classification task, different from the previous noise data screening work for multi-classification at the sample granularity, the loss modeling in this module is carried out at the sample-category binary tuple granularity. Since the training losses of clean sample-category binary tuples and noise follow different Gaussian distributions, clean data and noise data can be distinguished.
[0067] The loss function is calculated through binary cross - entropy loss (BCELoss), and the formula is as follows:
[0068]
[0069] Among them, x is the sample, c is the category, (x, c) is the sample - category binary tuple, y is the label given by the large model, taking values of 0 or 1. If the label output by the large model for sample x contains category c, it is labeled 1, otherwise it is labeled 0. is the label probability predicted by LSTM, and its value range is [0, 1].
[0070] The Gaussian mixture model is a probability model used to represent a mixture of multiple Gaussian distributions. The GMM assumes that data points are generated by multiple Gaussian distributions, and the probability density function formula of GMM is as follows:
[0071]
[0072] Among them, K is the number of Gaussian distributions generating the data, which is set to 2 in this module, namely the distribution of noise data loss and the distribution of clean data loss; π k is the weight of the k - th Gaussian distribution, satisfying and π k ≥0; is the probability density function of the k - th Gaussian distribution, with mean μ k and covariance matrix Σ k .
[0073] The parameters of the Gaussian mixture model include the weight π k , the mean μ k and the covariance matrix Σ k , and they are estimated using the Expectation - Maximization (EM) algorithm. The EM algorithm maximizes the log - likelihood function of the data through iterative optimization steps to find the optimal parameters and solve the distributions of clean data loss and noise data loss. Since the neural network fits clean data faster than noise data, it is considered that the distribution with a smaller mean is the distribution of clean data loss.
[0074] The probability that the loss of each sample - category binary tuple in training follows the clean data loss distribution is predicted through the Gaussian mixture model, that is k min = arg min k μ k .
[0075] Set the threshold t. If p > t, it is noise data; if p < t, it is clean data. For clean data, keep the original label, and for noise data, relabel it using the neural network.
[0076] Due to the self-supervised bias in using a single model to label its own training data, the present invention designs a multi-label data correction method for collaborative training of two models, and the process is as follows Figure 2 As shown. For the labels before the second confirmation of the large model and the labels after the second confirmation of the large model, an LSTM is used for training respectively, denoted as LSTM-1 and LSTM-2, and a GMM is independently used to model the training noise to distinguish between noise data and clean data. For noise data, when re-labeling, LSTM-1 and LSTM-2 are used for joint labeling to avoid self-supervised bias. The formula is as follows:
[0077]
[0078] Where k is the current round, is the pseudo-label output in the current round, is the pseudo-label labeled in the previous round, is the label labeled by the large model, and are the labels predicted by LSTM-1 and LSTM-2 respectively, p is the probability that follows the loss distribution of clean data, and t is the threshold.
[0079] Since the positive label and the negative label, that is, y (x,c) = 0 and y (x,c) = 1, have different loss distributions during training. To improve the ability of the model to distinguish clean data and noise data, in the above process, the positive label and the negative label are considered separately. In each round, LSTM-1 first models and corrects the positive label loss, and then models and corrects the negative label loss, while LSTM-2 does the opposite. After multiple rounds of iteration, the labels corrected by LSTM-1 and LSTM-2 in the last round are obtained through a voting mechanism to get the finally optimized label.
[0080] Furthermore, the industry classification algorithm based on the organization has high accuracy, but since it is difficult to determine the organization of the Web assets deployed by some IP addresses, using only this algorithm cannot meet the requirement of wide coverage. Therefore, the present invention proposes an industry classification module based on web pages, which classifies the IP address according to the web page text and can be widely applied to the IP addresses with Web assets deployed. The process includes a web page text acquisition module and a web page text classification module based on noisy learning. The working processes of each module are introduced in detail below.
[0081] In an embodiment of the present invention, the input of the web page text acquisition module is the domain name obtained by the IP address domain name reverse lookup module. The home page of the domain name is accessed through a browser to obtain the web page text, including the web page title,
[0082] Homepage text content and the element attributes of meta-description in HTML. Given that there may also be text information helpful for classification in internal pages, this module recursively tracks the links containing specific keywords on the homepage and obtains the titles, texts, and meta-description attributes in sub-pages. All the obtained text content is concatenated to get the original text data.
[0083] The original text data is in multiple languages. For the convenience of model processing, it is first uniformly translated into English. Since there is a lot of redundant text and noise in the data, the present invention uses a key sentence extraction algorithm based on word frequency to retain the relatively important sentences in the original text. The specific approach is to use the TF-IDF method to calculate the importance of words in the text, and sum up the importance of all words in the sentence to obtain the importance degree of the sentence. The formula is as follows:
[0084]
[0085] TF-IDF (t,d,D) =TF (t,d) ×IDF (t,D)
[0086]
[0087] Where t represents a word, s represents a sentence, d represents a document, that is, the original text data of a web page, D represents a document set, that is, the set of web page original text data. f t,d is the number of times the word t appears in d. n d is the total number of all words in d. N is the total number of documents in the set D, and |d∈D:t∈d| is the number of documents containing the word t.
[0088] Sort the sentences in the text data corresponding to each domain name in descending order according to the importance degree K s and concatenate the sorted results until the set text length limit is exceeded to obtain the processed web page text data.
[0089] In an embodiment of the present invention, for the web page text classification module based on noisy learning, using the processed web page text data as input to train a neural network and output an industry classification result. Due to the special requirements of the network security field for industry classification leading to scarce labeled data, and web page annotation requires a large amount of manual work, there is a need for classification without annotation in this scenario. The present invention uses the output of an industry classification algorithm based on organizations as labeled data to train the web page text classification algorithm without additional labeled data.
[0090] Although noise correction has been performed, there is still some noise in the data labels. In addition, due to the heterogeneity of the input of the organization-based industry classification algorithm and the web-based industry classification method, that is, for organization description classification and web text classification, noise is introduced into the labeled data, but at the same time, it provides the feasibility for further correcting the noise not found in the data label noise correction module. The specific process is as Figure 3 shown.
[0091] Considering the superiority of LSTM in distinguishing noise and the superiority of the Transformer-based model (Roberta) in classification, similar to the web-based industry classification module, this module uses LSTM to distinguish the noise in the input data. For the noisy data, LSTM and Roberta are used together for annotation, and the corrected data is used to train the LSTM and Roberta models. Due to the problem of uneven data volume among different categories in the labeled data, this module adopts the data augmentation method in semi-supervised learning, continuously selects a random data from the category with the least amount of data, and splices it with another data that does not belong to this category as the new input data, and the labels of the two are taken as the union as the label of the new data until the difference in the data volume between the category with the largest amount of data and the category with the least amount of data is less than the threshold.
[0092] After multiple rounds of iterative training, the trained Roberta model is deployed to actual applications for industry classification of web text data.
[0093] Furthermore, the effect of the present invention has been tested in the mapping application of the real network.
[0094] In terms of wide coverage, the present invention has successfully identified the industry classifications of more than 1 billion IP addresses, accounting for 20% of the active addresses, covering more than 20,000 organizations, proving that the present invention can be used for the industry classification problem in large-scale mapping of the existing network and can classify addresses of different organizations and different characteristics.
[0095] In terms of accuracy, the data marked for the present invention was evenly sampled by category and manually marked to construct an evaluation dataset. First, the evaluation task was transformed into multiple binary classification tasks, and the effects were evaluated separately for each category. The overall metric was the result of weighted averaging of each category's metric by the number of samples. The evaluation metrics included accuracy (the number of samples correctly predicted for this category / all samples predicted as this category), recall (the number of samples correctly predicted for this category / all samples belonging to this category), and F1-score. Based on the industry classification algorithm for organizations, the overall accuracy on the dataset was 96%, the overall recall was 83%, and the average F1-score was 0.89. Based on the industry classification algorithm for web pages, the overall accuracy on the evaluation dataset was 89%, the overall recall was 80%, and the average F1-score was 0.84. In addition, as a multi-label classification task, the present invention was also evaluated from the perspective of samples. For the industry classification algorithm based on organizations, the proportion of samples with completely correct labels (identical to the manual marking) was 85%, the proportion of samples with less than one missing label was 94%, and the proportion of samples with one more or one less label was 97%. For the industry classification algorithm based on web pages, the proportion of samples with completely correct labels was 82%, the proportion of samples with less than one missing label was 91%, and the proportion of samples with one more or one less label was 94%.
[0096] In terms of non-annotation, the classification process of the present invention does not require prior annotation of training data.
[0097] According to the IP address owner industry classification system for unlabeled data scenarios according to the embodiments of the present invention, it is possible to classify the IP addresses on which Web assets are deployed without pre-labeling data and can cover a wide range of IP addresses in the entire network. At the same time, by deeply analyzing the actual situation of the problems to be solved and the relationship between the requirements of each part, the large language model annotation method and the classification method with noisy learning are combined to comprehensively meet various requirements.
[0098] To implement the above embodiments, as Figure 4 shown, the present embodiment also provides an IP address owner industry classification method for unlabeled data scenarios, including:
[0099] S1, obtaining the domain name corresponding to the Web asset deployed on the IP address based on the PTR record and the Passive DNS record;
[0100] S2, based on the consistency result of the organization of the domain name owner and the organization of the IP address owner and the corresponding relationship between the organization and the industry, and based on the industry classification algorithm for organizations, identifying the organization of the domain name owner to determine the industry to which the IP address belongs;
[0101] S3, based on the industry classification algorithm for web pages and according to the web page text, performing industry classification on the industry to which the IP address belongs to obtain an industry classification result.
[0102] The IP address owner industry classification method for unlabeled data scenarios according to the embodiments of the present invention can classify the IP addresses with deployed Web assets without pre-labeling the data and can cover a wide range of IP addresses in the whole network. At the same time, by deeply analyzing the actual situation of the problems to be solved and the relationship between the requirements of each part, the large language model labeling method and the classification method with noisy learning are combined to comprehensively meet various requirements.
[0103] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0104] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
Claims
1. An IP address owner industry classification system for unlabeled data scenarios, characterized in that, Including: An IP address domain name reverse lookup module, which is used to obtain the domain name corresponding to the Web assets deployed on the IP address based on the PTR record and the Passive DNS record; An organization-based industry classification module, which is used to identify the industry to which the IP address belongs by determining the consistency result of the organization of the domain name owner and the organization of the IP address owner, the corresponding relationship between the organization and the industry, and based on the organization-based industry classification algorithm; A web page-based industry classification module, which is used to classify the industry to which the IP address belongs based on the web page-based industry classification algorithm and according to the web page text to obtain an industry classification result.
2. The system according to claim 1, wherein The IP address domain name reverse lookup module is also used for: Scanning the PTR record to cover IPv4 and IPv6 addresses by actively sending requests to the DNS resolver, and based on the obtained IPv6 response address list, sending requests to the addresses therein to obtain the PTR record query response probability; Using the Passive DNS record to supplement the domain name information corresponding to the IP address without a configured PTR record; Obtaining the fully qualified domain names retrieved from the DNS PTR record and the Passive DNS record, and extracting the second-level domain names from the fully qualified domain names for classification work.
3. The system according to claim 1, characterized in that, The organization-based industry classification module includes an organization identification module, a large model-based industry classification module, and a data label noise correction module; among them, the organization identification module is used for: Extracting the organization name of the registrant from the registrant field of the WHOIS record through the open API of the domain name registrar, using the extracted organization name as a clustering keyword, and clustering the domain names belonging to the same organization into the same cluster; Obtaining the SSL / TLS certificate by requesting port 443 of the main domain name, and extracting the key fields in the certificate. For the organization validation certificate and the extended validation certificate, the Organization field represents the organization holding the issued certificate and is used as the first keyword for organization clustering; distinguishing CDN and cloud service providers through manual identification and large language model judgment, and constructing a list of second-level domain names of CDN providers. If the second-level domain name falls into the list of second-level domain names of CDN providers, it is retained in the CDN cluster, otherwise it is excluded from the cluster; After deleting the second-level domain names belonging to CDN providers, clustering the remaining domain names based on the Common Name and Subject Alternative Name fields. The union of the second-level domain names in these two fields is used as the second feature for clustering; clustering the domain names into their respective clusters according to the organization name based on the two data sources of WHOIS and the certificate. For the data lacking the organization field, clustering is performed based on the Common Name and Subject Alternative Name fields.
4. The system according to claim 3, characterized in that, The large model-based industry classification module is also used for: Crawling the organization description from the online encyclopedia and extracting the industry to which it belongs through natural language processing technology; Build a pre-trained large language model and input the organization descriptions crawled from online encyclopedias and the definitions of various industries into the pre-trained large language model; among them, provide examples for the large language model through the chain of thought prompting technique, and the chain of thought includes summarizing the organization description, determining the main business, checking the industry definition, and providing reasons for matching or not matching the industry. Use the pre-trained large language model to select from all categories and give the potential categories that the organization may belong to, and verify whether the organization meets the proposed potential categories.
5. The system according to claim 3, wherein The data label noise correction module is also used for: Use the noise screening algorithm in noisy learning to regard the labels output by the large language model as the true values and input them as training data into the long short-term memory network to fit the labels and calculate the loss. Use the Gaussian mixture model to cluster the loss distribution of the training data to distinguish clean data and noisy data. Train LSTM-1 and LSTM-2 with the labels before and after the large language model's secondary confirmation respectively. Among them, each LSTM independently uses the GMM to model the training noise to distinguish noisy data and clean data; for noisy data, when re-labeling, use LSTM-1 and LSTM-2 for joint labeling, and process the positive and negative labels separately. In each round, LSTM-1 first models and corrects the positive label loss, and then models and corrects the negative label loss; LSTM-2 is the opposite. After multiple rounds of iteration, LSTM-1 and LSTM-2 respectively perform multiple loss modeling and correction on the positive and negative labels. After the last round of iteration, the labels corrected by LSTM-1 and LSTM-2 are integrated through a voting mechanism to obtain the finally optimized labels.
6. The system according to claim 1, characterized in that, The web-based industry classification module includes a web page text acquisition module and a web page text classification module based on noisy learning; among them, the web page text acquisition module is used for: Input the domain name obtained by the IP address domain name reverse lookup module, access the home page of the domain name through the browser, and obtain the web page title, the home page text content, and the meta-description element attributes in the HTML; recursively track the links containing specific keywords in the home page and obtain the titles, texts, and meta-description attributes in these sub-pages, and splice all the obtained text contents to form the original text data. Unify the translation of the original text data into English, and use the TF-IDF method to calculate the importance of the words in the text, and sum the importance of all the words in the sentence to obtain the importance degree of the sentence. Sort the sentences in the text data corresponding to each domain name from largest to smallest according to the importance degree, and splice the sorted sentences in turn until the set text length limit is exceeded to obtain the processed web page text data.
7. The system according to claim 6, wherein The web page text classification module based on noisy learning is also used for: Use the processed web page text data as input data to train the neural network, and use the output of the industry classification algorithm based on the organization as the labeled data. Use LSTM to distinguish noise data in the input data, and use LSTM and Roberta for joint annotation. The corrected data is used to train the LSTM and Roberta models. For the labeled data, select a random data from the category with the least amount of data, and splice it with another data that does not belong to this category to form new input data. The labels of the two are taken as the union as the label of the new data until the difference in the amount of data between the category with the most data and the category with the least data is less than the set threshold. Gradually optimize the LSTM and Roberta models through multiple rounds of iterative training, and deploy the trained Roberta model to actual applications for industry classification of web page text data.
8. An industry classification method for IP address owners in the scenario of unlabeled data, characterized in that, Including: Obtain the domain names corresponding to the Web assets deployed on the IP address based on the PTR records and Passive DNS records; Based on the consistency result of the organization of the domain name owner and the organization of the IP address owner and the corresponding relationship between the organization and the industry, and based on the industry classification algorithm of the organization, identify the organization of the domain name owner to determine the industry to which the IP address belongs; Based on the industry classification algorithm of the web page and according to the web page text, conduct industry classification on the industry to which the IP address belongs to obtain the industry classification result.
Citation Information
Cited By
IP address application scene identification method and system and electronic equipment
CN121030471A