A method for identifying a storage bucket
By collecting enterprise domain name information and crawling and analyzing information, combining regular matching and feature matching, the enterprise's storage address is identified, which solves the problem of insufficient bucket recognition capability in the prior art, and improves recognition efficiency and comprehensiveness.
Patent Information
- Application Number
- CN202410836490.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-06-26
AI Technical Summary
The prior art is difficult to effectively identify all bucket addresses of enterprises, especially in multi-cloud environments, making it difficult for enterprises to manage and monitor their sensitive data.
By collecting enterprise name information, obtaining enterprise domain name information, and crawling and parsing information through various methods, combining regular matching, feature matching and behavioral rules to determine the bucket address.
It improves the efficiency and comprehensiveness of bucket identification, and can identify multiple bucket addresses related to the enterprise, providing a comprehensive data source for subsequent sensitive data identification.
Smart Images

Figure CN118714111B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cloud storage, and particularly relates to a method for identifying storage buckets. Background Art
[0002] A bucket is a concept in cloud storage services, especially widely used in object storage services. It is a container for storing objects, which can be various data types such as files, pictures, videos, etc. Each bucket has a unique name, and a region is usually specified when it is created to determine the location where the data is stored.
[0003] The existing technologies mainly focus on the discovery method of subdomains. The existing methods mainly focus on the discovery of enterprise-related domain names. An enterprise may have multiple cloud service platform accounts. If the management is not standardized, it is very difficult to collect all the accounts and obtain the corresponding AK / SK. The method of only querying and obtaining the storage bucket address through the API provided by the public cloud is not sufficient to discover all the storage bucket addresses. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the purpose of the invention is to provide a method for identifying storage buckets, which improves the efficiency of storage bucket identification. This method solves the problem of insufficient enterprise storage bucket identification ability, and can help an enterprise discover as many storage bucket addresses related to it as possible, providing a comprehensive data source for subsequent sensitive data identification.
[0005] In the first aspect of the present invention, a method for identifying storage buckets is proposed, including:
[0006] S1, collecting enterprise name information, and obtaining enterprise domain name information based on the enterprise name information and according to a preset data acquisition method;
[0007] S2, performing information crawling based on the enterprise domain name information to obtain the web page content within the corresponding domain name;
[0008] S3, parsing the web page content to obtain the web page content parsing result, and determining the storage bucket address based on the web page content parsing result;
[0009] S4, screening the storage bucket addresses based on a preset screening rule to generate a final set of storage bucket addresses.
[0010] Further, in S1, the preset data acquisition method at least includes retrieving enterprise domain names on a network space mapping platform, retrieving enterprise domain names on an attack surface management platform, discovering enterprise subdomains using open source projects, collecting domain names through security devices deployed by the company, retrieving enterprise domain names in a search engine, obtaining from company materials, and obtaining domain names using the API of a cloud service provider.
[0011] Further, in S1, all the obtained domain names are de-duplicated, and the de-duplicated domain names are stored in a file or a database.
[0012] Further, the domain names are saved in JSON format, and each domain name includes at least two fields, which are respectively used to record domain name information and domain name source information.
[0013] Further, in S3, the web content corresponding to the enterprise domain name information is retrieved, and the enterprise domain name information is analyzed based on regular matching to determine whether it contains a bucket address. If it contains, the bucket address is recorded.
[0014] Further, in S3, the web content corresponding to the enterprise domain name information is retrieved to determine whether the enterprise domain name is in XML format. If the web content corresponding to the enterprise domain name is in XML format, the corresponding bucket address is recorded.
[0015] Further, in S3, the web content corresponding to the enterprise domain name information is retrieved, the header information is retrieved, the bucket features are extracted from the header information, and feature matching is performed based on the extracted bucket features. If the matching is successful, the corresponding bucket address is extracted.
[0016] Further, in S3, if the bucket features extracted from the header information fail to match, the body information is retrieved, the recognition features are extracted from the body information, and feature matching is performed according to the recognition features. If the matching is successful, the corresponding bucket address is extracted.
[0017] Further, in S3, if the matching fails according to both the header information and the body information, it is determined that the current enterprise domain name does not contain a bucket address.
[0018] Further, the preset screening rules at least include determining the derived source of the domain name, verifying the API of the cloud service provider, verifying whether the domain name comes from enterprise materials, and determining whether the bucket usage behavior matches the enterprise.
[0019] The beneficial effects of the present invention are as follows:
[0020] The method of the present invention can propose a method for automatically identifying a bucket according to the URL format and web content, can identify the cloud service provider providing the bucket service, and can judge whether a bucket address is related to the company according to the bucket usage behavior rules. Even if a bucket address is not operated by the company but comes from a third-party SaaS service, if this address is accessed multiple times, it can also be identified, greatly improving the identification efficiency and comprehensiveness of the bucket. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are only for the purpose of showing specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0022] Figure 1 It is a flowchart of a storage bucket identification method provided by an embodiment of the present invention;
[0023] Figure 2 It is a schematic diagram of obtaining enterprise domain name information provided by an embodiment of the present invention;
[0024] Figure 3 It is a flowchart of storage bucket identification provided by an embodiment of the present invention;
[0025] Figure 4 It is a flowchart for determining whether the storage bucket address is related to the company provided by an embodiment of the present invention. Detailed Embodiments
[0026] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. It should be understood that these descriptions are only exemplary and are not used to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts disclosed in the present invention.
[0028] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. The terms "installation", "connection", "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0029] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are only examples of methods and systems consistent with some aspects of the present invention as detailed in the appended claims.
[0030] The present invention provides a method for identifying storage buckets, which solves the problem of insufficient storage bucket identification ability of enterprises, and can help an enterprise discover as many storage bucket addresses related to it as possible, providing a comprehensive data source for subsequent sensitive data identification.
[0031] Method Embodiment
[0032] A method for identifying storage buckets provided by the present invention includes:
[0033] S1, collecting enterprise name information, and obtaining enterprise domain name information based on the enterprise name information and according to a preset data acquisition method.
[0034] In this step, as Figure 1 and Figure 2 shown, enterprise-related domain names are obtained through various methods. As shown in the appendix Figure 2As shown in the figure. It is necessary to collect information such as the main domain names and names of the company / sub - company / holding company. This can be collected through the domain name information filing management system of the Ministry of Industry and Information Technology, or through commercial query platforms. These information are the basic information for subsequent sub - domain name discovery. Taking Company A as an example, the full name of Company A is A Cloud Technology Co., Ltd. By searching in the domain name information filing management system of the Ministry of Industry and Information Technology (https: / / beian.miit.gov.cn / ), 46 main domain name records can be found. In the commercial query platform, it can be seen that Company A has 83 branches, which are the information of its subsidiaries and branch companies. It can also be seen that Company A has invested in three companies externally, Company B, Company C, and Company D. These three are the holding companies of Company A, and domain name collection also needs to be carried out in the domain name information filing management system later.
[0035] The present invention proposes to obtain enterprise - related domain names through 7 methods. Since the domain names obtained by a single method may all be incomplete, it is necessary to collect and summarize domain names through multiple methods.
[0036] (1) Conduct enterprise - related domain name retrieval on the network space mapping platform.
[0037] Currently, there are many domestic and foreign network space mapping platforms, such as Shodan, ZoomEye, FOFA, etc. Taking FOFA as an example, when it is known that ctyun.cn is one of the main domain names of China Telecom Cloud, all sub - domain name information under this main domain name can be obtained through the search syntax host = "ctyun.cn".
[0038] (2) Conduct enterprise - related domain name retrieval on the attack surface management platform.
[0039] Currently, the attack surface management platform also provides enterprise domain name discovery services. Taking 00 Information Security (https: / / 0.zone / ) as an example, by searching for China Telecom Cloud Technology Co., Ltd. on it, 4244 domain name records can be seen.
[0040] (3) Use open - source projects to discover enterprise sub - domain names.
[0041] There are also some current open-source projects for subdomain discovery, such as OneForAll (https: / / github.com / shmilylty / OneForAll), subfinder (https: / / github.com / projectdiscovery / subfinder), subDomainsBrute (https: / / github.com / lijiejie / subDomainsBrute), etc. Taking OneForAll as an example, if you want to discover the subdomains of ctyun.cn, you can use the following command: python3 oneforall.py --target ctyun.cn run.
[0042] (4) Conduct domain name collection through the security devices deployed by the company.
[0043] Companies generally deploy many security devices, and some security devices will also record domain name access records, such as Internet behavior management devices. However, the domain name records of these devices cannot directly distinguish whether the domain names are related to the company itself, nor can they determine whether they are related to the company's storage buckets. In the third step, this patent will handle this situation.
[0044] (5) Conduct enterprise-related domain name retrieval in search engines.
[0045] Subdomains can be collected in common search engines, such as Google, Baidu, and Bing. Taking Google as an example, you can obtain the subdomains related to ctyun.cn by retrieving site:ctyun.cn.
[0046] (6) Query company information.
[0047] There will be domain name records in the company information, but this information may not be complete. Sometimes, due to departmental barriers, it may not be possible to obtain all the data. In addition, storage buckets are generally not recorded by enterprise asset administrators. Therefore, the present invention combines multiple means to discover enterprise-related domain names and obtain enterprise domain name information.
[0048] (7) Utilize the APIs of cloud service providers.
[0049] Cloud service providers all provide APIs for obtaining the storage bucket addresses corresponding to accounts. If you can obtain the AK / SK of the account for which an enterprise purchases public cloud services, you can obtain them in this way. However, different departments of an enterprise may have different accounts, and it is difficult to ensure that you can ask all relevant people and obtain the AK / SK of all accounts.
[0050] Finally, save the collected domain names after removing duplicates to a file or database. For example, they can be saved as a JSON-formatted file. Each domain name contains two fields. One is "domain" for recording domain name information, and the other is "from" for recording the source of the domain name. Note that there may be multiple sources for the domain name. A record is as follows:
[0051] {
[0052] "domain":"www.ctyun.cn",
[0053] "from":["FOFA","OneForAll"]
[0054] }
[0055] S2, perform information crawling based on the enterprise domain name information to obtain the web page content within the corresponding domain name.
[0056] In this step, write a crawler to crawl the web page content corresponding to the domain name. There are multiple ways to implement the crawler. Taking the Python language as an example, it can be implemented through the requests library. The crawling results are saved to a file or database, and both the HTTP header and body need to be saved. The following format can be referred to. Since the content of the header and body is very long, only examples are given here.
[0057] {
[0058] "domain":"www.ctyun.cn",
[0059] "header":"HTTP / 1.1 200OK…",
[0060] "body":"<!doctype html>…",
[0061] "from":["FOFA","OneForAll"]
[0062] }
[0063] S3, parse the web page content to obtain the web page content parsing result, and determine the storage bucket address based on the web page content parsing result.
[0064] In this step, as Figure 3 shown, (1) the storage bucket addresses of some public clouds have a relatively fixed format, and the storage bucket addresses can be identified by regular matching of this format. Taking Alibaba Cloud as an example, the address format of the Alibaba Cloud storage bucket is as follows:
[0065] https: / / [bucket-name].oss-[region].aliyuncs.com
[0066] Among them, [bucket-name] is the name of the bucket, and [region] is the region where the Alibaba Cloud bucket is located. Therefore, the bucket addresses that conform to this format can be identified through regular expressions. However, the address formats of not all public cloud buckets are very standardized, and there are also cases of private clouds and dedicated clouds. In addition, there are many public cloud providers, and it is difficult to avoid omission. Therefore, only the address formats of several mainstream public clouds can be sorted out. Other bucket addresses can be identified in the following way.
[0067] (2) Determine whether the web page of the domain name is in XML format. It can be judged through the Content-Type field in the header. If the web page is in XML format, the content of this field will contain application / XML. Through the analysis of the buckets of multiple cloud providers, these buckets will follow the bucket protocol format of Amazon, and the web page content is in XML format. Therefore, if the web page is not in XML format, it must not be a bucket.
[0068] (3) There are some characteristics in the header part of the bucket address to indicate that it is a bucket. Taking Huawei Cloud as an example, the following content will appear in the header:
[0069] Server:OBS
[0070] x-obs-request-id:000001***6F696
[0071] x-obs-bucket-location:cn-east-3
[0072] x-obs-id-2:32AAAQ****k4NPt4e
[0073] Therefore, the characteristics that can be extracted are Server:OBS, x-obs-request-id, x-obs-bucket-location, and x-obs-id-2. These can be used as the characteristics of Huawei Cloud buckets. If they are matched, it means it is a Huawei Cloud bucket. Similar characteristic information can also be extracted for other clouds.
[0074] (4) If the characteristics of the bucket cannot be matched in the header part, the body matching process will be entered. If the files in the bucket can be accessed without authentication, at this time, the body will contain a string with the word ListBucketResult, which can be used as the characteristics of the bucket. However, if the bucket requires a key to access, since there is no key in the current detection, the access will be denied. Taking Amazon Cloud as an example, the content returned when it denies access is
[0075] <?XML version="1.0" encoding="UTF-8"?>
[0076] <error>
[0077] <code>AccessDenied< / code>
[0078] <message>Access Denied< / message>
[0079] <requestid>8H9B6J2K0V3B7CWK< / requestid>
[0080] <hostid>oc6HOMY8***jKQ=< / hostid>
[0081] < / error>
[0082] It is also possible to extract features (such as <code>AccessDenied< / code> ) as the identification features of the bucket. Note that the features may not be unique. The core is to be able to distinguish this type of bucket.
[0083] (5) If the characteristics of both the header and the body are not matched, then this domain name is not the bucket address.
[0084] Save the recognition result to a file or database. The format can refer to the following content. domain is the domain name address, isBucket indicates whether this domain name is a bucket, bucketProvider indicates which cloud service provider provides this bucket. from is used to record the source of the domain name.
[0085] {
[0086] "domain":"wb.oss-cn-hangzhou.aliyuncs.com",
[0087] "isBucket":True,
[0088] "bucketProvider":"Alibaba Cloud",
[0089] "from":["FOFA","OneForAll"]
[0090] }
[0091] S4. Filter the bucket addresses based on preset filtering rules to generate the final set of bucket addresses.
[0092] In this step, filter the bucket addresses to obtain the final set of bucket addresses. According to the following judgment method, the priority decreases step by step. As shown in the appendix Figure 4 shown.
[0093] (1) If the domain name is derived from the main domain name of the company / sub - company / branch / holding company, then this domain name is the company - related bucket address.
[0094] (2) If the domain name comes from the API of a cloud service provider, then this domain name is the company - related bucket address.
[0095] (3) If the domain name comes from the company's internal asset record form, then this domain name is the company - related bucket address.
[0096] (4) If the domain name comes from the security devices deployed by the company and can match the company's bucket usage behavior rules at the same time, then this domain name is the company - related bucket address.
[0097] A bucket address recorded by the security devices deployed by the company may also be used by a third - party website when a company employee downloads external materials. However, in this case, the bucket address will not be accessed by multiple people frequently. Therefore, if a bucket address is accessed by multiple people and multiple times within the company within a certain period of time, it can be considered that this address is related to the company. The specific time, number of people, and number of times can be adjusted according to the actual situation of the company. For example, it can be that within one month, there are more than 10 people and more than 20 visits, then it is considered a company - related bucket address.
[0098] (5) Otherwise, this domain name has nothing to do with the company's bucket address.
[0099] The present invention proposes a method for automatically identifying a bucket according to the URL format and web page content, and can identify the cloud service provider providing the bucket service; a method for judging whether a bucket address is related to the company according to the bucket usage behavior rules. Even if a bucket address is not operated by the company but comes from a third - party SaaS service, if this address is accessed multiple times, it can also be identified.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for identifying a storage bucket, characterized in that: include: S1, collecting enterprise name information, and obtaining enterprise domain name information based on the enterprise name information and in accordance with a preset data acquisition method; S2, crawling information based on the enterprise domain name information to obtain the web page content within the corresponding domain name; S3, parsing the webpage content to obtain the webpage content parsing result, and determining the storage bucket address based on the webpage content parsing result; S4, filtering the bucket addresses based on preset filtering rules to generate a final bucket address set; In S1, the preset data acquisition methods at least include searching for enterprise domain names on the cyberspace mapping platform, searching for enterprise domain names on the attack surface management platform, discovering enterprise subdomain names using open source projects, collecting domain names through security devices deployed by the company, searching for enterprise domain names on search engines, obtaining from company materials, and obtaining domain names using the API of a cloud service provider; In S3, retrieve the web page content corresponding to the enterprise domain name information, retrieve the header information, extract the bucket features from the header information, perform feature matching based on the extracted bucket features, and if the match is successful, extract the corresponding bucket address.
2. A method for identifying a storage bucket according to claim 1, characterized in that: In S1, all the acquired domain names are deduplicated and the deduplicated domain names are stored in a file or a database.
3. A storage bucket identification method according to claim 2, characterized in that: The domain name is saved in JSON format. Each domain name includes at least two fields, which are used to record the domain name information and the domain name source information respectively.
4. A method for identifying a storage bucket according to claim 1, characterized in that: In S3, the web page content corresponding to the enterprise domain name information is retrieved, and the enterprise domain name information is analyzed based on regular matching to determine whether it contains the storage bucket address. If it does, the storage bucket address is recorded.
5. A method for identifying a storage bucket according to claim 1, characterized in that: In S3, the web page content corresponding to the enterprise domain name information is retrieved to determine whether the enterprise domain name is in XML format. If the web page content corresponding to the enterprise domain name is in XML format, the corresponding storage bucket address is recorded.
6. A method for identifying a storage bucket according to claim 5, characterized in that: In S3, if the bucket feature extracted from the header information fails to match, the body information is retrieved, and the identification features are extracted from the body information. Feature matching is performed based on the identification features. If the match is successful, the corresponding bucket address is extracted.
7. A method for identifying a storage bucket according to claim 1, characterized in that: In S3, if no match is found based on the header information or the body information, it is determined that the current enterprise domain name does not contain the bucket address.
8. A method for identifying a storage bucket according to claim 1, characterized in that: The preset screening rules at least include determining the derived source of the domain name, verifying the API of the cloud service provider, verifying whether the domain name comes from corporate data, and determining whether the bucket usage behavior matches the enterprise.
Citation Information
Patent Citations
Key information extraction method and device, computer storage medium and electronic equipment
CN114238733A