Enterprise asset classification model training method and device

By building a knowledge graph of enterprise website and domain name system data and adjusting the parameters of the enterprise asset classification model, the problem of insufficient utilization of enterprise asset sequence information in the existing technology is solved, and the accuracy and comprehensiveness of asset classification are improved.

CN120217162APending Publication Date: 2025-06-27CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510381417.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When identifying and classifying enterprise assets, the existing technology lacks in-depth utilization of asset sequence information, resulting in limited scope and accuracy of classification, making it difficult to deal with complex and diverse enterprise asset data.

Method used

By building a knowledge graph based on enterprise website data and domain name system data, the path prediction value of the enterprise asset classification model is obtained, and the loss values ​​of edges and sub-paths are calculated to adjust model parameters and enhance the utilization of node association relationships and path information.

Benefits of technology

It improves the accuracy of enterprise asset classification, can more comprehensively identify enterprise asset types, enhance the model's learning ability to edges, and deeply utilize path information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217162A_ABST
    Figure CN120217162A_ABST
Patent Text Reader

Abstract

The invention provides an enterprise asset classification model training method and device, is applied to the technical field of data processing, and is used for improving the accuracy of enterprise asset classification. The method comprises the following steps: constructing a first knowledge graph according to website data and domain name system data of each enterprise in a plurality of enterprises; obtaining at least one path corresponding to any enterprise from the first knowledge graph, and inputting each path into an enterprise asset classification model for prediction; obtaining a predicted value of the enterprise asset classification model for each edge in each path, and calculating a first loss value between the predicted value of each edge and the true value of each edge; obtaining a predicted value of the enterprise asset classification model for each sub-path in each path, and calculating a second loss value between the predicted value of each sub-path and the true value of each sub-path; and obtaining a total loss value according to the first loss value and the second loss value, and adjusting parameters of the enterprise asset classification model with the purpose that the total loss value is smaller than a first threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a method and device for training an enterprise asset classification model. Background Art

[0002] Accurately and comprehensively identifying enterprise assets is of great significance for optimizing resource allocation and preventing security risks. However, using a simple classification model to identify and classify enterprise assets lacks in-depth utilization of asset sequence information, thus limiting the scope and accuracy of asset classification. When facing relatively complex and diverse enterprise asset data, it is difficult to achieve an ideal effect.

[0003] In view of this, how to improve the accuracy of enterprise asset classification is an urgent problem to be solved. Summary of the Invention

[0004] This application provides a method and device for training an enterprise asset classification model, which is used to improve the accuracy of enterprise asset classification.

[0005] In a first aspect, an embodiment of this application provides a method for training an enterprise asset classification model, which can be applied to any electronic device with processing capabilities. The method includes: Construct a first knowledge graph according to the website data and domain name system data of each enterprise among multiple enterprises; wherein, each node in the first knowledge graph corresponds one-to-one with each parameter of the website data and each parameter of the domain name system data, and the parameter is used to indicate the virtual assets of the enterprise website; the path relationship between nodes in the first knowledge graph corresponds to the association relationship between each parameter in the website data, the association relationship between each parameter in the domain name system data, and the association relationship between each parameter in the website data and each parameter in the domain name system data; Obtain at least one path corresponding to any enterprise from the first knowledge graph, and input each path in the at least one path into the enterprise asset classification model for prediction; the parameter types corresponding to the starting nodes of the at least one path are the same, and the parameter types corresponding to the ending nodes of the at least one path are also the same; Obtain the predicted value of each edge in each path by the enterprise asset classification model, where the predicted value of each edge is used to indicate the probability that the parameter corresponding to the node included in each edge belongs to the corresponding enterprise, and calculate the first loss value between the predicted value of each edge and the true value of each edge; Obtain the predicted value of each sub-path in each path by the enterprise asset classification model, where each sub-path is a path from the starting node to any node on the corresponding path, and the predicted value of each sub-path is used to indicate the probability that the parameter corresponding to the node included in each sub-path belongs to the corresponding enterprise, and calculate the second loss value between the predicted value of each sub-path and the true value of each sub-path; Obtain the total loss value based on the first loss value and the second loss value, and adjust the parameters of the enterprise asset classification model with the aim of the total loss value being less than the first threshold.

[0006] In this method, a first knowledge graph is constructed based on the website data and domain name system data of the enterprise, which is convenient for more intuitive analysis of the association relationships between various virtual assets of the enterprise in the follow-up; at least one path of each enterprise is obtained from the first knowledge graph and input into the enterprise asset classification model, the predicted value of each edge in each path by the enterprise asset classification model is obtained, the first loss value between the predicted value of each edge and the true value of each edge is calculated, and the predicted value of each sub-path in each path by the enterprise asset classification model is obtained, the second loss value between the predicted value of each sub-path and the true value of each sub-path is calculated, and the parameters of the model are adjusted by integrating the first loss value and the second loss value; on the one hand, by analyzing and predicting each edge, the learning ability of the enterprise asset classification model for the association relationships between nodes (i.e., edges) can be enhanced, so that the output prediction results can more accurately reflect the associations between nodes; on the other hand, by comprehensively analyzing all the information from the start node to the current node at any node of the path, the in-depth utilization of path (sequence) information is realized, and the enterprise asset types can be more comprehensively identified; combining these two aspects can improve the accuracy of enterprise asset classification.

[0007] Optionally, each parameter in the website data includes: Uniform Resource Locator (URL), first domain name, certificate, IP address; each parameter in the domain name system data includes: first domain name, second domain name and / or IP address, the second domain name is the domain name mapped by the first domain name on other services, and the other services are used to indicate services different from the service of the first domain name; constructing the first knowledge graph according to the website data and domain name system data of each enterprise in multiple enterprises includes: taking each item of URL, first domain name, second domain name, certificate, and IP address of each enterprise as a node; creating an edge between the node corresponding to the URL of each enterprise and the node corresponding to the first domain name, the node corresponding to the certificate, and the node corresponding to the IP address of the enterprise; creating an edge between the node corresponding to the first domain name and the node corresponding to the second domain name, and / or, between the node corresponding to the first domain name and the node corresponding to the IP address of each enterprise; constructing the first knowledge graph according to the nodes corresponding to each parameter and all the created edges.

[0008] Optionally, obtaining at least one path corresponding to each enterprise from the first knowledge graph includes: obtaining all paths from each first domain name to the IP address connected to each first domain name from the first knowledge graph, and determining at least one first domain name corresponding to any IP address; querying at least one enterprise corresponding to at least one first domain name; determining that at least one enterprise is the same enterprise, then obtaining the path from the first domain name connected to the any IP address to the any IP address from all paths, and obtaining at least one path of the enterprise corresponding to the any IP address.

[0009] Optionally, the method further includes: determining that at least one enterprise includes at least two different enterprises, then obtaining any n paths from the first domain name connected to the any IP address to the any IP address from all paths as negative samples, where n is a positive integer less than the number of paths from the first domain name connected to the any IP address to the any IP address; inputting each path in the negative samples into the enterprise asset classification model for prediction; obtaining the predicted value of each edge in each path by the enterprise asset classification model, where the predicted value of each edge is used to indicate the probability that the parameter corresponding to the node included in each edge belongs to the corresponding enterprise, and calculating the third loss value between the predicted value of each edge and the true value of each edge; obtaining the predicted value of each sub-path in each path by the enterprise asset classification model, where each sub-path is the path from the start node to any node on the corresponding path, and the predicted value of each sub-path is used to indicate the probability that the parameter corresponding to the node included in each sub-path belongs to the corresponding enterprise, and calculating the fourth loss value between the predicted value of each sub-path and the true value of each sub-path; obtaining the total loss value of the negative samples according to the third loss value and the fourth loss value, and adjusting the parameters of the enterprise asset classification model with the aim that the total loss value of the negative samples is greater than the second threshold.

[0010] Optionally, obtaining the predicted value of each edge in each path by the enterprise asset classification model includes: any one edge in each edge includes a first node and a second node, and the first node is the previous node of the second node. The following operations are performed on the any one edge by the enterprise asset classification model: calculating the conditional probability of the appearance of the second node in the presence of the first node to obtain the feature vector of the second node; splicing the feature vector of the second node and the feature vector of the first node to obtain a spliced vector; the feature vector of the first node is the conditional probability of the appearance of the first node in the presence of the previous node of the first node; integrating and non-linearly transforming the spliced vector according to a preset rule to obtain the predicted value of the any one edge.

[0011] Optionally, obtain the predicted values of the enterprise asset classification model for each sub-path in each path, including: Each sub-path includes m nodes, where m is a positive integer. For any sub-path, the enterprise asset classification model performs the following operations: Calculate the feature vectors of each of the m nodes, where the feature vector is the probability of each node occurring given the previous node of each node; Calculate the mean of all the feature vectors of the m nodes; Integrate and perform non-linear transformation processing on the mean according to a preset rule to obtain the predicted value of any sub-path.

[0012] In a second aspect, an embodiment of the present application provides an enterprise asset classification method, including: Construct a sub-graph based on the website data and domain name system data of the enterprise to be detected; Each node of the sub-graph corresponds one-to-one with each parameter in the website data and each parameter in the domain name system data, where the parameter is used to indicate the virtual assets of the website of the enterprise to be detected; The path relationship between the nodes in the sub-graph corresponds to the association relationship between each parameter in the website data, the association relationship between each parameter in the domain name system data, and the association relationship between each parameter in the website data and each parameter in the domain name system data; Input the sub-graph into the enterprise asset classification model, and use the enterprise asset classification model to predict each edge and each sub-path in the sub-graph. Each sub-path is a path from the starting node to any node on the sub-graph, and obtain the prediction result of the enterprise asset classification model; The prediction result is used to indicate whether the parameter corresponding to each node in the sub-graph belongs to the enterprise to be detected.

[0013] In this method, by constructing a sub-graph based on the website data and domain name system data of the enterprise to be detected, it is convenient for the enterprise asset classification model to more clearly and accurately analyze each virtual asset of the enterprise to be detected; By comprehensively analyzing and predicting each edge and each sub-path in the sub-graph of the enterprise to be detected according to the enterprise asset classification model, the path information and the association relationship between the nodes can be deeply utilized to more comprehensively identify each node and improve the accuracy of enterprise asset classification.

[0014] In a third aspect, an embodiment of the present application provides an enterprise asset classification model training device, including: A construction module, configured to: Construct a first knowledge graph based on the website data and domain name system data of each enterprise in a plurality of enterprises; Wherein, each node in the first knowledge graph corresponds one-to-one with each parameter in the website data and each parameter in the domain name system data, and the parameter is used to indicate the virtual assets of the enterprise website; The path relationship between the nodes in the first knowledge graph corresponds to the association relationship between each parameter in the website data, the association relationship between each parameter in the domain name system data, and the association relationship between each parameter in the website data and each parameter in the domain name system data; An identification module, configured to: obtain at least one path corresponding to any enterprise from a first knowledge graph, and input each path in the at least one path into an enterprise asset classification model for prediction; parameter types corresponding to start nodes of the at least one path are the same, and parameter types corresponding to end nodes of the at least one path are also the same; obtain prediction values of each edge in each path by the enterprise asset classification model, where the prediction value of each edge is used to indicate the probability that the parameter corresponding to the node included in each edge belongs to the corresponding enterprise, and calculate a first loss value between the prediction value of each edge and the true value of each edge; obtain prediction values of each sub-path in each path by the enterprise asset classification model, where each sub-path is a path from the start node to any node on the corresponding path, and the prediction value of each sub-path is used to indicate the probability that the parameter corresponding to the node included in each sub-path belongs to the corresponding enterprise, and calculate a second loss value between the prediction value of each sub-path and the true value of each sub-path. An optimization module, configured to: obtain a total loss value according to the first loss value and the second loss value, and adjust parameters of the enterprise asset classification model with the aim of making the total loss value less than a first threshold.

[0015] Optionally, each parameter in the website data includes: a Uniform Resource Locator (URL), a first domain name, a certificate, and an IP address; each parameter in the Domain Name System (DNS) data includes: the first domain name, a second domain name, and / or an IP address, where the second domain name is a domain name mapped by the first domain name in other services, and the other services are used to indicate services different from the service of the first domain name; when constructing the first knowledge graph according to the website data and DNS data of each enterprise among multiple enterprises, the construction module is specifically configured to: use each item of the URL, the first domain name, the second domain name, the certificate, and the IP address of each enterprise as a node; create an edge between the node corresponding to the URL of each enterprise and the nodes corresponding to the first domain name, the certificate, and the IP address of the enterprise respectively; create an edge between the node corresponding to the first domain name of each enterprise and the node corresponding to the second domain name, and / or, between the node corresponding to the first domain name and the node corresponding to the IP address; construct the first knowledge graph according to the nodes corresponding to each parameter and all the created edges.

[0016] Optionally, when obtaining at least one path corresponding to each enterprise from the first knowledge graph, the identification module is specifically configured to: obtain all paths from each first domain name to the IP address connected to each first domain name from the first knowledge graph, and determine at least one first domain name corresponding to any IP address; query at least one enterprise corresponding to the at least one first domain name; if it is determined that the at least one enterprise is the same enterprise, obtain the path from the first domain name connected to the any IP address to the any IP address from all paths, and obtain at least one path corresponding to the enterprise corresponding to the any IP address.

[0017] Optionally, the recognition module is further configured to: determine that at least one enterprise includes at least two different enterprises, and then obtain any n paths from the first domain name connected to any IP address to the any IP address as negative samples from all paths, where n is a positive integer less than the number of paths from the first domain name connected to the any IP address to the any IP address; input each path in the negative samples into the enterprise asset classification model for prediction; obtain the prediction value of each edge in each path by the enterprise asset classification model, where the prediction value of each edge is used to indicate the probability that the parameter corresponding to the node included in each edge belongs to the corresponding enterprise, and calculate the third loss value between the prediction value of each edge and the true value of each edge; obtain the prediction value of each sub-path in each path by the enterprise asset classification model, where each sub-path is a path from the starting node to any node on the corresponding path, and the prediction value of each sub-path is used to indicate the probability that the parameter corresponding to the node included in each sub-path belongs to the corresponding enterprise, and calculate the fourth loss value between the prediction value of each sub-path and the true value of each sub-path; the optimization module is further configured to: obtain the total loss value of the negative samples according to the third loss value and the fourth loss value, and adjust the parameters of the enterprise asset classification model with the aim that the total loss value of the negative samples is greater than the second threshold.

[0018] Optionally, when the recognition module obtains the prediction value of each edge in each path by the enterprise asset classification model, it is specifically configured to: any one of the edges in each edge includes a first node and a second node, the first node is the previous node of the second node, and the following operations are performed on the any one of the edges by the enterprise asset classification model: calculate the conditional probability of the appearance of the second node in the presence of the first node to obtain the feature vector of the second node; splice the feature vector of the second node and the feature vector of the first node to obtain a spliced vector; the feature vector of the first node is the conditional probability of the appearance of the first node in the presence of the previous node of the first node; perform integration and non-linear transformation processing on the spliced vector according to a preset rule to obtain the prediction value of the any one of the edges.

[0019] Optionally, when the recognition module obtains the prediction value of each sub-path in each path by the enterprise asset classification model, it is specifically configured to: any one of the sub-paths in each sub-path includes m nodes, where m is a positive integer, and the following operations are performed on any one of the sub-paths by the enterprise asset classification model: calculate the feature vector of each of the m nodes, where the feature vector is the probability of the appearance of each node in the case of the previous node of each node; calculate the mean value of all the feature vectors of the m nodes; perform integration and non-linear transformation processing on the mean value according to a preset rule to obtain the prediction value of any one of the sub-paths.

[0020] Fourthly, an embodiment of the present application provides an enterprise asset classification device, including: A building module, configured to: construct a sub-graph based on the website data and domain name system data of an enterprise to be detected; each node of the sub-graph corresponds one-to-one to each parameter in the website data and each parameter in the domain name system data, and the parameter is used to indicate the virtual assets of the website of the enterprise to be detected; the path relationship between the nodes in the sub-graph corresponds to the association relationship between the parameters in the website data, the association relationship between the parameters in the domain name system data, and the association relationship between the parameters in the website data and the parameters in the domain name system data; An identification module, configured to: input the sub-graph into an enterprise asset classification model, predict each edge and each sub-path in the sub-graph through the enterprise asset classification model, each sub-path is a path from the starting node to any node on the sub-graph, and obtain the prediction result of the enterprise asset classification model; the prediction result is used to indicate whether the parameter corresponding to each node in the sub-graph belongs to the enterprise to be detected.

[0021] In a fifth aspect, an embodiment of the present application provides an electronic device, including at least one processor, and when the at least one processor executes a computer program stored in a memory, the method in the first aspect or any optional implementation manner of the first aspect or the method in the second aspect or any optional implementation manner of the second aspect is implemented.

[0022] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store instructions, and when the instructions are executed, the method in the first aspect or any optional implementation manner of the first aspect or the method in the second aspect or any optional implementation manner of the second aspect is implemented.

[0023] In a seventh aspect, an embodiment of the present application provides a computer program product, including computer program code, and when the computer program code runs on a computer, the method in the first aspect or any optional implementation manner of the first aspect or the method in the second aspect or any optional implementation manner of the second aspect is implemented.

[0024] The technical effects or advantages of one or more technical solutions provided in the second, third, fourth, fifth, sixth, and seventh aspects in the embodiments of the present application can be correspondingly explained by the technical effects or advantages of the corresponding one or more technical solutions provided in the first aspect. Description of the Drawings

[0025] Figure 1 It is a flowchart of a method for training an enterprise asset classification model provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the principle of an enterprise asset classification model provided by an embodiment of the present application; Figure 3Structural diagram of an enterprise asset classification model training device provided by an embodiment of the present application; Figure 4 Structural diagram of an enterprise asset classification device provided by an embodiment of the present application; Figure 5 Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0026] In the technical solution of the present application, the collection, dissemination, use, etc. of data all comply with the requirements of relevant national laws and regulations.

[0027] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0028] The technical solution of the present application will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0029] It should be understood that "a plurality of" in the description of the embodiments of the present application means two or more. The "first", "second", etc. in the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. The term "and / or" in the embodiments of the present application is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. The modules in the embodiments of the present application refer to parts with independent functions in a software system.

[0030] Enterprise asset identification plays a crucial role in enterprise security. Accurately and comprehensively identifying enterprise assets is of great significance for optimizing resource allocation and preventing security risks. However, current enterprise asset identification methods are mainly based on simple classification networks, lacking in-depth utilization of sequence information, thus limiting the scope and accuracy of asset identification. Secondly, such classification networks usually ignore the relationships between the edges of nodes, resulting in insufficient sequence structure expression ability, lacking in-depth exploration of the context associations between nodes, making it difficult to accurately depict the mutual influences between nodes, and being able to identify only single-type assets, with the application scope and accuracy being restricted. This also makes it difficult for traditional asset classification models to achieve ideal results when facing complex and diverse enterprise asset data.

[0031] In view of this, the embodiments of the present application are provided. A knowledge graph is constructed based on the website data and domain name system data of each enterprise, facilitating a more intuitive and clear analysis of the relationships between each enterprise and its virtual assets; at least one path corresponding to any enterprise is obtained from the knowledge graph and input into the enterprise asset classification model for prediction, obtaining the prediction values of each edge in each path and the prediction values of each sub-path in the enterprise asset classification model, and calculating the first loss value and the second loss value respectively according to each prediction value and the corresponding true value, obtaining the total loss value by synthesizing the first loss value and the second loss value, and adjusting the parameters of the enterprise asset classification model according to the total loss value until the total loss value is less than the first threshold; through the analysis and prediction of each edge, the learning ability of the enterprise asset classification model for the edges (i.e., the association relationships between nodes) in each path can be enhanced; and through the analysis and prediction of each sub-path, in-depth utilization of path information can be achieved, enabling the trained enterprise asset classification model to more comprehensively identify enterprise asset types, thereby improving the accuracy of enterprise asset classification.

[0032] It can be understood that the embodiments of the present application can be applied to any asset classification and identification scenarios, including but not limited to the above classification of enterprise virtual assets.

[0033] See Figure 1 , which is a flowchart of a method for training an enterprise asset classification model provided by the embodiments of the present application. Taking enterprise assets as virtual assets as an example, the method includes the following steps S101 to S106: S101. Construct a first knowledge graph according to the website data and domain name system data of each enterprise among multiple enterprises.

[0034] Among them, each node in the first knowledge graph corresponds one by one to each parameter of the website data and each parameter of the domain name system data, and the parameter is used to indicate the virtual assets of the enterprise website; the path relationship between the nodes in the first knowledge graph corresponds to the association relationship between each parameter in the website data, the association relationship between each parameter in the domain name system data, and the association relationship between each parameter in the website data and each parameter in the domain name system data.

[0035] In a possible example, the embodiment of the present application further provides a method for obtaining the website data of each enterprise. The specific implementation manner of this method is as follows: First, collect multiple enterprise names.

[0036] It can be understood that the enterprise name is the name of the enterprise that needs to perform asset management.

[0037] Secondly, query the main domain name of the enterprise corresponding to each enterprise name in the Internet Content Provider (ICP) database based on each enterprise name.

[0038] Then, query at least one historical sub-domain name corresponding to each enterprise in the Domain Name System (DNS) database according to the main domain name of each enterprise.

[0039] Then, construct at least one Uniform Resource Locator (URL) of the enterprise based on at least one historical sub-domain name.

[0040] Exemplarily, adding the prefix "http: / / " or "https: / / " to each historical sub-domain name can obtain the URL corresponding to each historical sub-domain name.

[0041] Finally, perform web crawling according to the obtained URL of each enterprise to obtain the website data of each enterprise.

[0042] In this way, all the website data of each enterprise can be obtained more comprehensively, which is convenient for more accurately identifying the assets of each enterprise subsequently.

[0043] In a possible embodiment, each parameter in the website data includes: Uniform Resource Locator URL, first domain name, certificate, IP address; each parameter in the domain name system data includes: first domain name, second domain name and / or IP address, and the second domain name is the domain name mapped by the first domain name in other services, and the other service is used to indicate a service different from the service of the first domain name; the specific implementation manner of step S101 is as follows: Take each item among the URL, the first domain name, the second domain name, the certificate, and the IP address of each enterprise as a node; Create an edge between the node corresponding to the URL of each enterprise and the nodes corresponding to the first domain name, the certificate, and the IP address of that enterprise respectively; Create an edge between the node corresponding to the first domain name and the node corresponding to the second domain name of each enterprise, and / or between the node corresponding to the first domain name and the node corresponding to the IP address; Construct a first knowledge graph based on the nodes corresponding to each parameter and all the created edges.

[0044] Exemplarily, for website data, since the website data of an enterprise such as the first domain name, the certificate, the IP address, etc. are all obtained by web crawling based on the URL, that is, each parameter in the website data is related to the URL. Therefore, taking the URL of each enterprise as the basic node, create an undirected edge (i.e., an edge without direction) between the URL and the first domain name, that is, the first domain name - URL (or URL - the first domain name); create an undirected edge between the URL and the certificate, that is, URL - certificate; create an edge between the URL and the IP address, that is, URL - IP address.

[0045] If all parameters of the website data of an enterprise include the URL, the first domain name, the certificate, and the IP address, that is, there are edges between the URL and the first domain name, the certificate, and the IP address, that is, the certificate, the IP address, and the first domain name are all connected.

[0046] For the DNS database, it can be considered that there is an association relationship between the various parameters included in each record in the DNS database, and each record must include the first domain name. Therefore, taking the first domain name as the basic point, create an edge between the first domain name and other parameters in each record respectively. For example, if the record in the DNS database is a CNAME record, that is, a record that maps the first domain name to the second domain name, then the various parameters obtained in the domain name system data include the first domain name and the second domain name, so create an undirected edge between the first domain name and the second domain name, that is, the first domain name - the second domain name; if the record in the DNS database is an A / AAAA record, then the various parameters obtained in the domain name system data include the first domain name and the IP address, so create an edge between the first domain name and the IP address, that is, the first domain name - IP address. If both records exist, then the various parameters obtained in the domain name system data include the first domain name, the second domain name, and the IP address, so create an edge between the first domain name and the second domain name, and between the first domain name and the IP address respectively, to get the first domain name - the second domain name - IP address, etc. The embodiments of the present application do not limit this.

[0047] Since the first domain name exists in both the website data and the DNS database, a first knowledge graph can be constructed based on the first domain name and each parameter connected to the first domain name. For example, the paths included in the first knowledge graph can be: first domain name - IP address, first domain name - URL - IP address, first domain name - second domain name - URL - IP address, first domain name - second domain name - URL - certificate - IP address, etc. Only some examples are listed in the embodiments of the present application, and the actual situation is not limited thereto.

[0048] S102. Obtain at least one path corresponding to any enterprise from the first knowledge graph.

[0049] Among them, the parameter types corresponding to the starting nodes of at least one path are the same, and the parameter types corresponding to the ending nodes of at least one path are also the same.

[0050] In a possible embodiment, the specific implementation manner of step S102 is as follows: Obtain all paths from each first domain name to the IP address connected to each first domain name from the first knowledge graph, and determine at least one first domain name corresponding to any IP address; Query at least one enterprise corresponding to at least one first domain name; If it is determined that at least one enterprise is the same enterprise, then obtain the path from the first domain name connected to the any IP address to the any IP address from all paths, and obtain at least one path of the enterprise corresponding to the any IP address.

[0051] Exemplarily, the starting node of at least one path is the first domain name, and the ending node is the IP address, such as the first domain name - second domain name - URL - certificate - IP address in the above example.

[0052] It can be understood that an IP address can only belong to one enterprise, and an enterprise may have multiple domain names. Therefore, the enterprise corresponding to each first domain name (query the name, identifier, etc. of the enterprise) can be queried. If the enterprises corresponding to at least one first domain name are the same, it means that the any IP address only belongs to the enterprise, and all paths of the any IP address are correct and can be used as positive samples for training the enterprise asset classification model.

[0053] In another possible embodiment, if it is determined that at least one enterprise includes at least two different enterprises, then obtain any n paths from the first domain name connected to the any IP address to the any IP address from all paths as negative samples, where n is a positive integer less than the number of paths from the first domain name connected to the any IP address to the any IP address.

[0054] It can be understood that if the enterprise corresponding to any one of the IP addresses is not unique, it means that there is an incorrect path among all the paths from the first domain name connected to the any one of the IP addresses to the any one of the IP addresses. Then, randomly obtain n paths from these paths as negative samples. The specific value of n can be determined according to the actual situation. For example, if the number of paths from the first domain name connected to the any one of the IP addresses to the any one of the IP addresses is large, the value of n can be appropriately increased; conversely, the value of n can be appropriately decreased. The embodiments of the present application do not limit this. In this way, the accuracy of the negative samples can be improved, and further, when training the enterprise asset classification model based on the negative samples later, the accuracy of the enterprise asset classification model can be further improved.

[0055] S103. Input each path in at least one path into the enterprise asset classification model for prediction.

[0056] Exemplarily, usually, since the edges in the first knowledge graph can only identify that there is an association between two nodes, but it is not necessarily possible to determine that both nodes belong to the same enterprise. For example: Enterprise A deploys a website on the server of Enterprise B. It can be determined that the domain name of the website belongs to Enterprise A, but the server (i.e., the IP address) belongs to Enterprise B. Therefore, the embodiments of the present application can use the enterprise asset classification model to determine whether the parameters corresponding to each node in the path belong to each enterprise.

[0057] The enterprise asset classification model provided by the embodiments of the present application can be a StackLSTM model, that is, a deep learning model that enhances the model's expression ability by stacking multiple Long Short-Term Memory (LSTM) layers. Other types of models can also be selected according to actual needs. The embodiments of the present application do not limit this.

[0058] S104. Obtain the predicted value of each edge in each path by the enterprise asset classification model, and calculate the first loss value between the predicted value of each edge and the true value of each edge.

[0059] In a possible embodiment, the specific implementation manner of obtaining the predicted value of each edge in each path by the enterprise asset classification model is as follows: Any one of the edges in each edge includes a first node and a second node. The first node is the previous node of the second node. Perform the following operations on the any one of the edges through the enterprise asset classification model: Calculate the conditional probability of the appearance of the second node when the first node exists to obtain the feature vector of the second node; Concatenate the feature vector of the second node and the feature vector of the first node to obtain a concatenated vector; the feature vector of the first node is the conditional probability of the appearance of the first node when the previous node of the first node exists. Integrate and perform non - linear transformation on the spliced vector according to the preset rules to obtain the predicted value of any one of the edges.

[0060] Exemplarily, the order of each node in each path starts with the first domain name as the starting node and ends with the IP address as the ending node. The order of the nodes between the starting node and the ending node can be arranged according to the degree of association between the parameters corresponding to each node and the first domain name in practice, or arranged according to other rules, or randomly arranged. The embodiments of the present application do not limit this.

[0061] Exemplarily, since the starting node of each path is the first domain name, the corresponding enterprise can be confirmed according to the first domain name, that is, it can be determined that the first domain name belongs to the enterprise. Therefore, a conditional probability model is introduced, and by calculating the conditional probabilities of each node, it is predicted whether each node belongs to the enterprise.

[0062] See Figure 2 , which is a schematic diagram of the principle of an enterprise asset classification model provided by the embodiments of the present application.

[0063] Assume that any selected path is the first domain name - the second domain name - the certificate - the URL - the IP address. As Figure 2 shown, the first domain name is sport.sohu.com, the second domain name is sohu.com, the certificate is *.sohu.com, the URL is https: / / sohu.com, and the IP address is 1.1.1.1.

[0064] First, convert each parameter into an input feature vector, that is, a feature vector that can be recognized by the enterprise asset classification model.

[0065] Then, input this path into the enterprise asset classification model. The enterprise asset classification model uses LSTM to identify and predict each node in the input path, calculates the conditional probability of each node appearing when the previous node exists through the hidden layer, and outputs the conditional probabilities of each node.

[0066] Taking the first edge (i.e., the first domain name - the second domain name) as an example, the following operations are performed on the first edge through the enterprise asset classification model: 1. Calculate the conditional probability corresponding to the first domain name and the conditional probability corresponding to the second domain name.

[0067] Since the first domain name necessarily belongs to the corresponding enterprise, the conditional probability corresponding to the first domain name is 1; the conditional probability of the second domain name is the predicted conditional probability of the second domain name appearing when the first domain name exists (i.e., the probability that the second domain name also belongs to the corresponding enterprise when the first domain name belongs to the corresponding enterprise).

[0068] 2. Concatenate the conditional probability of the first domain name and the conditional probability of the second domain name to obtain a concatenated vector.

[0069] Exemplarily, the specific concatenation method can be selected according to actual needs, and the embodiments of the present application do not limit this.

[0070] 3. Input the concatenated vector into the fully connected layer and the activation function layer for processing such as integration, feature transformation, feature mapping, and non-linear transformation to obtain the predicted value of the first edge.

[0071] In a possible embodiment, after obtaining the predicted value of each edge based on the above method, the method for calculating the first loss value can be determined according to actual needs. The embodiments of the present application provide an example of a method for obtaining the first loss value by calculating the cross-entropy loss value, as shown in the following formula: ; ; ; where is the feature vector of the first node, is the feature vector of the second node, v is the concatenated vector of the first node and the second node is the predicted value of any edge; the value of y can be determined according to actual needs, such as always being 1, etc.; N represents the number of nodes, represents the first loss value.

[0072] It can be understood that the above is only a possible example of calculating the first loss value given by the embodiments of the present application, and the actual situation is not limited to this.

[0073] S105. Obtain the predicted value of each sub-path in each path by the enterprise asset classification model, and calculate the second loss value between the predicted value of each sub-path and the true value of each sub-path.

[0074] In a possible embodiment, the specific implementation manner of obtaining the predicted value of each sub-path in each path by the enterprise asset classification model is as follows: Any sub-path in each sub-path includes m nodes, where m is a positive integer. Perform the following operations on any sub-path through the enterprise asset classification model: Calculate the feature vector of each of the m nodes. The feature vector is the probability of each node appearing under the condition of the previous node of each node; Calculate the mean value of all the feature vectors of the m nodes; Integrate and perform non-linear transformation processing on the mean value according to a preset rule to obtain the predicted value of any sub-path.

[0075] Exemplarily, such as Figure 2As shown, any sub-path can be any one of: the first domain name - the second domain name, the first domain name - the second domain name - the certificate, the first domain name - the second domain name - the certificate - the URL, and the first domain name - the second domain name - the certificate - the URL - the IP address.

[0076] Taking any sub-path as the first domain name - the second domain name - the certificate as an example, that is, the value of m is 3: First, calculate the feature vectors of the first domain name, the second domain name, and the certificate respectively. That is, the feature vector of the first domain name, that is, the conditional probability is 1, the feature vector of the second domain name is the conditional probability of the second domain name appearing in the presence of the first domain name, and the feature vector of the certificate is the conditional probability of the certificate appearing in the presence of the second domain name.

[0077] Then, calculate the mean of the feature vectors of these three nodes. That is Figure 2 As shown, first calculate the sum of the conditional probabilities of these three nodes, and then calculate the mean according to the sum.

[0078] Finally, input the obtained mean into the fully connected layer of the enterprise asset classification model for integration, feature mapping and other processing, and then input the processed mean output by the fully connected layer into the activation function layer for feature transformation, non-linear transformation and other processing to obtain the predicted value of this sub-path.

[0079] In a possible example, the method for calculating the second loss value can also be selected according to actual needs. In this embodiment of the application, taking calculating the cross-entropy loss value to obtain the second loss value as an example, the specific formula is as follows: ; ; Among them, represents any one of the m nodes, represents the predicted value of any sub-path; the value of y can be determined according to actual needs, such as always being 1, etc.; N represents the number of nodes, represents the second loss value.

[0080] It can be understood that the above is only a possible example of calculating the second loss value given in this embodiment of the application, and the actual situation is not limited to this.

[0081] S106. Obtain the total loss value according to the first loss value and the second loss value, and adjust the parameters of the enterprise asset classification model with the aim that the total loss value is less than the first threshold.

[0082] It can be understood that in practical applications, the proportion or weight of the first loss value and the second loss value in the total loss value can be adjusted according to actual needs, such as the accuracy requirement and training cost requirement of the enterprise asset classification model, or only the first loss value or only the second loss value can be used to adjust the parameters of the enterprise asset classification model.

[0083] In addition, the first threshold can also be set according to actual needs, such as adjusting the parameters of the model until the total loss value converges (that is, it can be understood that the first threshold is the critical value at convergence), etc. The embodiments of the present application do not limit this.

[0084] In a possible embodiment, if the training sample used by the enterprise asset classification model is a negative sample, that is, the negative sample obtained in the above step S102, then each path in the negative sample is input into the enterprise asset classification model for prediction; the prediction value of each edge in each path obtained by the enterprise asset classification model is obtained, and the prediction value of each edge is used to indicate the probability that the parameters corresponding to the nodes included in each edge belong to the corresponding enterprise, and the third loss value between the prediction value of each edge and the true value of each edge is calculated; the prediction value of each sub-path in each path obtained by the enterprise asset classification model is obtained, each sub-path is a path from the start node to any node on the corresponding path, and the prediction value of each sub-path is used to indicate the probability that the parameters corresponding to the nodes included in each sub-path belong to the corresponding enterprise, and the fourth loss value between the prediction value of each sub-path and the true value of each sub-path is calculated; the total loss value of the negative sample is obtained according to the third loss value and the fourth loss value, and the parameters of the enterprise asset classification model are adjusted with the aim that the total loss value of the negative sample is greater than the second threshold.

[0085] It can be understood that when training with negative samples, the calculation methods of the third loss value and the fourth loss value, as well as the second threshold, can be selected according to actual needs, and the embodiments of the present application do not limit this.

[0086] In this method, a first knowledge graph is constructed based on the website data and domain name system data of an enterprise, facilitating subsequent more intuitive analysis of the association relationships between various virtual assets of the enterprise website; at least one path of each enterprise is obtained from the first knowledge graph and input into the enterprise asset classification model, the predicted value of each edge in each path by the enterprise asset classification model is obtained, the first loss value between the predicted value of each edge and the true value of each edge is calculated, and the predicted value of each sub-path in each path by the enterprise asset classification model is obtained, the second loss value between the predicted value of each sub-path and the true value of each sub-path is calculated, and the parameters of the model are adjusted by integrating the first loss value and the second loss value; on the one hand, by analyzing and predicting each edge, the learning ability of the enterprise asset classification model for the relationships (i.e., edges) between nodes can be enhanced, enabling the output prediction results to more accurately reflect the associations between nodes; on the other hand, by comprehensively analyzing all the information from the starting node to a node at any node of the path, the in-depth utilization of path (sequence) information is realized, and the enterprise asset types can be more comprehensively identified; combining these two aspects can improve the accuracy of enterprise asset classification.

[0087] In a possible design, an embodiment of the present application further provides an enterprise asset classification method, which is based on the above-trained enterprise asset classification model, and the specific implementation manner of the method is as follows: First, a sub-graph is constructed according to the website data and domain name system data of the enterprise to be detected.

[0088] Among them, each node of the sub-graph corresponds one-to-one with each parameter in the website data and each parameter in the domain name system data, and the parameter is used to indicate the virtual assets of the website of the enterprise to be detected; the path relationship between the nodes in the sub-graph corresponds to the association relationship between each parameter in the website data, the association relationship between each parameter in the domain name system data, and the association relationship between each parameter in the website data and each parameter in the domain name system data.

[0089] Exemplarily, the acquisition method of the website data of the enterprise to be detected can also refer to the acquisition method of the website data in the above step S101, and the embodiments of the present application will not elaborate here.

[0090] Then, the sub-graph is input into the enterprise asset classification model, and each edge and each sub-path in the sub-graph are predicted by the enterprise asset classification model. Each sub-path is the path from the starting node to any node on the sub-graph, and the prediction result of the enterprise asset classification model is obtained; the prediction result is used to indicate whether the parameter corresponding to each node in the sub-graph belongs to the enterprise to be detected.

[0091] In this way, each edge and each sub-path in the sub-graph of the enterprise to be detected can be comprehensively analyzed and predicted according to the enterprise asset classification model, enabling in-depth utilization of path information and the correlation relationships between nodes, more comprehensive identification of each node, and improvement of the accuracy of enterprise asset classification.

[0092] Optionally, after obtaining the prediction results output by the enterprise asset classification model, the assets (i.e., parameters) belonging to the enterprise to be detected can be provided to the security experts of the enterprise to be detected. The security experts can manage these assets and visualize the communication relationships of the assets to facilitate real-time understanding of the status and distribution of virtual assets, thereby improving the efficiency of security monitoring.

[0093] It can be understood that each of the above embodiments can be implemented independently or combined arbitrarily according to requirements, and the present application places no restrictions on this.

[0094] The method provided by the embodiments of the present application is introduced above. The device provided by the embodiments of the present application is introduced below.

[0095] Based on the same technical concept, the embodiments of the present application provide an enterprise asset classification model training device. The device includes modules / units / means for executing the methods performed by the electronic device in the above method embodiments. These modules / units / means can be implemented by software, or by hardware, or by hardware executing corresponding software.

[0096] Exemplarily, referring to Figure 3 , device 300 includes: A construction module 301, configured to: construct a first knowledge graph according to the website data and domain name system data of each enterprise in a plurality of enterprises; wherein, each node in the first knowledge graph corresponds one-to-one with each parameter of the website data and each parameter in the domain name system data, and the parameter is used to indicate the virtual assets of the enterprise website; the path relationship between nodes in the first knowledge graph corresponds to the association relationship between each parameter in the website data, the association relationship between each parameter in the domain name system data, and the association relationship between each parameter in the website data and each parameter in the domain name system data; The recognition module 302 is configured to: obtain at least one path corresponding to any enterprise from the first knowledge graph, and input each path in the at least one path into the enterprise asset classification model for prediction; the parameter types corresponding to the starting nodes of the at least one path are the same, and the parameter types corresponding to the ending nodes of the at least one path are also the same; obtain the predicted value of each edge in each path by the enterprise asset classification model, where the predicted value of each edge is used to indicate the probability that the parameter corresponding to the node included in each edge belongs to the corresponding enterprise, and calculate the first loss value between the predicted value of each edge and the true value of each edge; obtain the predicted value of each sub-path in each path by the enterprise asset classification model, where each sub-path is a path from the starting node to any node on the corresponding path, and the predicted value of each sub-path is used to indicate the probability that the parameter corresponding to the node included in each sub-path belongs to the corresponding enterprise, and calculate the second loss value between the predicted value of each sub-path and the true value of each sub-path. The optimization module 303 is configured to: obtain the total loss value according to the first loss value and the second loss value, and adjust the parameters of the enterprise asset classification model with the aim of making the total loss value less than the first threshold.

[0097] Optionally, each parameter in the website data includes: Uniform Resource Locator (URL), first domain name, certificate, IP address; each parameter in the Domain Name System (DNS) data includes: first domain name, second domain name, and / or IP address, where the second domain name is the domain name mapped by the first domain name in other services, and the other services are used to indicate services different from the service of the first domain name; when the construction module 301 constructs the first knowledge graph according to the website data and DNS data of each enterprise among multiple enterprises, it is specifically configured to: use each item of the URL, first domain name, second domain name, certificate, and IP address of each enterprise as a node; create an edge between the node corresponding to the URL of each enterprise and the nodes corresponding to the first domain name, certificate, and IP address of the enterprise; create an edge between the node corresponding to the first domain name and the node corresponding to the second domain name, and / or, between the node corresponding to the first domain name and the node corresponding to the IP address of each enterprise; construct the first knowledge graph according to the nodes corresponding to each parameter and all the created edges.

[0098] Optionally, when the recognition module 302 obtains at least one path corresponding to each enterprise from the first knowledge graph, it is specifically configured to: obtain all paths from each first domain name to the IP address connected to each first domain name from the first knowledge graph, and determine at least one first domain name corresponding to any IP address; query at least one enterprise corresponding to the at least one first domain name; if it is determined that the at least one enterprise is the same enterprise, then obtain the path from the first domain name connected to the any IP address to the any IP address from all the paths, and obtain at least one path corresponding to the enterprise corresponding to the any IP address.

[0099] Optionally, the recognition module 302 is further configured to: determine that at least one enterprise includes at least two different enterprises, and then obtain any n paths from the first domain name connected to the any IP address to the any IP address as negative samples from all the paths, where n is a positive integer less than the number of paths from the first domain name connected to the any IP address to the any IP address; input each path in the negative samples into the enterprise asset classification model for prediction; obtain the predicted values of each edge in each path by the enterprise asset classification model, where the predicted value of each edge is used to indicate the probability that the parameter corresponding to the node included in each edge belongs to the corresponding enterprise, and calculate the third loss value between the predicted value of each edge and the true value of each edge; obtain the predicted values of each sub-path in each path by the enterprise asset classification model, where each sub-path is a path from the start node to any node on the corresponding path, and the predicted value of each sub-path is used to indicate the probability that the parameter corresponding to the node included in each sub-path belongs to the corresponding enterprise, and calculate the fourth loss value between the predicted value of each sub-path and the true value of each sub-path; the optimization module 303 is further configured to: obtain the total loss value of the negative samples according to the third loss value and the fourth loss value, and adjust the parameters of the enterprise asset classification model with the aim that the total loss value of the negative samples is greater than the second threshold.

[0100] Optionally, when the recognition module 302 obtains the predicted values of each edge in each path by the enterprise asset classification model, it is specifically configured to: any one of the edges in each edge includes a first node and a second node, where the first node is the previous node of the second node, and perform the following operations on the any one of the edges through the enterprise asset classification model: calculate the conditional probability of the second node appearing in the case where the first node exists to obtain the feature vector of the second node; splice the feature vector of the second node and the feature vector of the first node to obtain a spliced vector; the feature vector of the first node is the conditional probability of the first node appearing in the case where the previous node of the first node exists; perform integration and non-linear transformation processing on the spliced vector according to a preset rule to obtain the predicted value of the any one of the edges.

[0101] Optionally, when the recognition module 302 obtains the predicted values of each sub-path in each path by the enterprise asset classification model, it is specifically configured to: any one of the sub-paths in each sub-path includes m nodes, where m is a positive integer, and perform the following operations on any one of the sub-paths through the enterprise asset classification model: calculate the feature vector of each of the m nodes, where the feature vector is the probability of each node appearing in the case of the previous node of each node; calculate the mean value of all the feature vectors of the m nodes; perform integration and non-linear transformation processing on the mean value according to a preset rule to obtain the predicted value of any one of the sub-paths.

[0102] Based on the same technical concept, an embodiment of the present application provides an enterprise asset classification model training device. Exemplarily, refer to Figure 4 , the device 400 includes: A building block 401, configured to: construct a sub-graph based on website data and domain name system data of an enterprise to be detected; each node in the sub-graph corresponds one-to-one to each parameter in the website data and each parameter in the domain name system data, and the parameter is used to indicate virtual assets of the website of the enterprise to be detected; the path relationship between nodes in the sub-graph corresponds to the association relationship between each parameter in the website data, the association relationship between each parameter in the domain name system data, and the association relationship between each parameter in the website data and each parameter in the domain name system data; An identification module 402, configured to: input the sub-graph into an enterprise asset classification model, predict each edge and each sub-path in the sub-graph through the enterprise asset classification model, each sub-path is a path from a starting node to any node on the sub-graph, and obtain a prediction result of the enterprise asset classification model; the prediction result is used to indicate whether the parameter corresponding to each node in the sub-graph belongs to the enterprise to be detected.

[0103] It should be understood that all relevant contents of each step involved in the above method embodiment can be cited to the function description of the corresponding functional module, and will not be elaborated here.

[0104] Based on the same technical concept, see Figure 5 , an embodiment of the present application further provides an electronic device 500, including: At least one processor 501; and a communication interface 503 communicatively connected to the at least one processor 501; the at least one processor 501 executes instructions stored in a memory 502, so that the electronic device 500 executes the method steps executed by the dashboard in the above method embodiment through the communication interface 503.

[0105] Optionally, the memory 502 is located outside the electronic device 500.

[0106] Optionally, the electronic device 500 includes the memory 502, the memory 502 is connected to the at least one processor 501, and the memory 502 has instructions executable by the at least one processor 501. Attached Figure 5 The memory 502 is optionally represented by a dashed line for the electronic device 500.

[0107] Wherein, the at least one processor 501 and the memory 502 can be coupled through an interface circuit or integrated together, which is not limited here.

[0108] In the embodiment of the present application, the specific connection medium between the at least one processor 501, the memory 502, and the communication interface 503 is not limited. In the embodiment of the present application Figure 5 it is shown that the at least one processor 501, the memory 502, and the communication interface 503 are connected through a bus 504, and the bus is inFigure 5 is represented by a thick line. The connection manners between other components are only for illustrative purposes and are not restrictive. This bus portion can be an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 it is only represented by a thick line, but it does not mean that there is only one bus or one type of bus.

[0109] It should be understood that the processor mentioned in the embodiments of the present application can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented by software, the processor can be a general-purpose processor that realizes by reading software code stored in a memory.

[0110] Exemplarily, the processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0111] It should be understood that the memory mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which serves as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0112] It should be noted that when the processor is a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, the memory (storage module) may be integrated in the processor.

[0113] It should be noted that the memory described herein is intended to include but not be limited to these and any other suitable types of memory.

[0114] Based on the same technical concept, the embodiments of the present application also provide a computer-readable storage medium, which is used to store instructions. When the instructions are executed, the computer executes the method steps executed by any device in the above method embodiments.

[0115] Based on the same technical concept, the embodiments of the present application also provide a computer program product, including computer program code. When the computer program code runs on the computer, the method steps executed by any device in the above method embodiments are implemented.

[0116] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0117] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0118] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0120] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A method for training an enterprise asset classification model, characterized in that: include: A first knowledge graph is constructed based on the website data and domain name system data of each enterprise in a plurality of enterprises; wherein each node in the first knowledge graph corresponds to each parameter of the website data and each parameter in the domain name system data, and the parameter is used to indicate the virtual assets of the enterprise website; and the path relationship between the nodes in the first knowledge graph corresponds to the association relationship between each parameter in the website data, the association relationship between each parameter in the domain name system data, and the association relationship between each parameter in the website data and each parameter in the domain name system data; Obtain at least one path corresponding to any enterprise from the first knowledge graph, and input each of the at least one path into the enterprise asset classification model for prediction; the parameter types corresponding to the starting nodes of the at least one path are the same, and the parameter types corresponding to the ending nodes of the at least one path are also the same; Obtaining a predicted value of each edge in each path by the enterprise asset classification model, wherein the predicted value of each edge is used to indicate a probability that a parameter corresponding to a node included in each edge belongs to a corresponding enterprise, and calculating a first loss value between the predicted value of each edge and a true value of each edge; Obtaining a predicted value of each subpath in each path by the enterprise asset classification model, wherein each subpath is a path from the starting node to any node on the corresponding path, and the predicted value of each subpath is used to indicate a probability that a parameter corresponding to a node included in each subpath belongs to a corresponding enterprise, and calculating a second loss value between the predicted value of each subpath and a true value of each subpath; A total loss value is obtained according to the first loss value and the second loss value, and parameters of the enterprise asset classification model are adjusted with the goal of making the total loss value smaller than a first threshold.

2. The method according to claim 1, characterized in that The parameters in the website data include: a uniform resource locator URL, a first domain name, a certificate, and an IP address; the parameters in the domain name system data include: a first domain name, a second domain name and / or an IP address, wherein the second domain name is a domain name mapped to the first domain name on other services, and the other services are used to indicate services different from the services of the first domain name; the first knowledge graph is constructed according to the website data and domain name system data of each enterprise in the plurality of enterprises, including: Each of the URL, the first domain name, the second domain name, the certificate, and the IP address of each enterprise is taken as a node; Create an edge between the node corresponding to the URL of each enterprise and the node corresponding to the first domain name of the enterprise, the node corresponding to the certificate, and the node corresponding to the IP address; Creating an edge between a node corresponding to the first domain name and a node corresponding to the second domain name of each enterprise, and / or between a node corresponding to the first domain name and a node corresponding to the IP address; The first knowledge graph is constructed according to the nodes corresponding to the various parameters and all the edges created.

3. The method according to claim 1, characterized in that The acquiring at least one path corresponding to each enterprise from the first knowledge graph includes: Acquire all paths from each first domain name to an IP address connected to each first domain name from the first knowledge graph, and determine at least one first domain name corresponding to any IP address; Query at least one enterprise corresponding to the at least one first domain name; If it is determined that the at least one enterprise is the same enterprise, a path from the first domain name connected to any IP address to any IP address is obtained from all the paths to obtain at least one path of the enterprise corresponding to any IP.

4. The method according to claim 3, characterized in that The method further comprises: Determining that the at least one enterprise includes at least two different enterprises, obtaining any n paths from the first domain name connected to the any IP address to the any IP address from all the paths as negative samples, wherein n is a positive integer less than the number of paths from the first domain name connected to the any IP address to the any IP address; Input each path in the negative sample into the enterprise asset classification model for prediction; Obtaining a predicted value of each edge in each path by the enterprise asset classification model, wherein the predicted value of each edge is used to indicate a probability that a parameter corresponding to a node included in each edge belongs to a corresponding enterprise, and calculating a third loss value between the predicted value of each edge and a true value of each edge; Obtaining a predicted value of each subpath in each path calculated by the enterprise asset classification model, wherein each subpath is a path from the starting node to any node on the corresponding path, and the predicted value of each subpath is used to indicate a probability that a parameter corresponding to a node included in each subpath belongs to a corresponding enterprise, and calculating a fourth loss value between the predicted value of each subpath and a true value of each subpath; The total loss value of the negative samples is obtained according to the third loss value and the fourth loss value; and the parameters of the enterprise asset classification model are adjusted with the purpose of making the total loss value of the negative samples greater than a second threshold.

5. The method according to claim 1, characterized in that The obtaining of the predicted value of each edge in each path by the enterprise asset classification model includes: Any one of the edges includes a first node and a second node, the first node is a previous node of the second node, and the following operation is performed on any one of the edges through the enterprise asset classification model: Calculate the conditional probability of the second node appearing when the first node exists, and obtain a feature vector of the second node; concatenating the feature vector of the second node and the feature vector of the first node to obtain a concatenated vector; the feature vector of the first node is the conditional probability of the first node appearing when the previous node of the first node exists; The splicing vectors are integrated and nonlinearly transformed according to preset rules to obtain a predicted value of any edge.

6. The method according to claim 1, characterized in that: The obtaining of the predicted value of each sub-path in each path by the enterprise asset classification model includes: Any subpath in each subpath includes m nodes, where m is a positive integer. The following operations are performed on any subpath through the enterprise asset classification model: Calculate a feature vector of each of the m nodes, where the feature vector is a conditional probability of each node appearing when a previous node of each node exists; Calculate the mean of all eigenvectors of the m nodes; The mean values ​​are integrated and nonlinearly transformed according to preset rules to obtain a predicted value of any sub-path.

7. A method for classifying enterprise assets, characterized in that: include: Constructing a sub-graph according to the website data and the domain name system data of the enterprise to be detected; each node of the sub-graph corresponds to each parameter in the website data and each parameter in the domain name system data, and the parameter is used to indicate the virtual assets of the website of the enterprise to be detected; The path relationship between the nodes in the sub-graph corresponds to the association relationship between the parameters in the website data, the association relationship between the parameters in the domain name system data, and the association relationship between the parameters in the website data and the parameters in the domain name system data; The sub-graph is input into the enterprise asset classification model, and each edge and each sub-path in the sub-graph is predicted through the enterprise asset classification model, where each sub-path is a path from a starting node to any node on the sub-graph, and the prediction result of the enterprise asset classification model is obtained; the prediction result is used to indicate whether the parameters corresponding to each node in the sub-graph belong to the enterprise to be detected.

8. An enterprise asset classification model training device, characterized in that: include: A construction module, used to: construct a first knowledge graph according to the website data and domain name system data of each enterprise in a plurality of enterprises; wherein each node in the first knowledge graph corresponds to each parameter of the website data and each parameter in the domain name system data, and the parameter is used to indicate the virtual assets of the enterprise website; the path relationship between the nodes in the first knowledge graph corresponds to the association relationship between the parameters in the website data, the association relationship between the parameters in the domain name system data, and the association relationship between the parameters in the website data and the domain name system data; An identification module, used to: obtain at least one path corresponding to any enterprise from the first knowledge graph, and input each path in the at least one path into the enterprise asset classification model for prediction; the parameter types corresponding to the starting nodes of the at least one path are the same, and the parameter types corresponding to the ending nodes of the at least one path are also the same; obtain the predicted value of each edge in each path by the enterprise asset classification model, the predicted value of each edge is used to indicate the probability that the parameter corresponding to the node included in each edge belongs to the corresponding enterprise, and calculate the first loss value between the predicted value of each edge and the true value of each edge; obtain the predicted value of each subpath in each path by the enterprise asset classification model, each subpath is a path from the starting node to any node on the corresponding path, the predicted value of each subpath is used to indicate the probability that the parameter corresponding to the node included in each subpath belongs to the corresponding enterprise, and calculate the second loss value between the predicted value of each subpath and the true value of each subpath; The optimization module is used to obtain a total loss value according to the first loss value and the second loss value, and adjust the parameters of the enterprise asset classification model with the goal of making the total loss value less than a first threshold.

9. An enterprise asset classification device, characterized in that: include: A construction module, used to: construct a sub-graph according to the website data and domain name system data of the enterprise to be detected; each node of the sub-graph corresponds to each parameter in the website data and each parameter in the domain name system data, and the parameter is used to indicate the virtual assets of the website of the enterprise to be detected; The path relationship between the nodes in the sub-graph corresponds to the association relationship between the parameters in the website data, the association relationship between the parameters in the domain name system data, and the association relationship between the parameters in the website data and the parameters in the domain name system data; An identification module is used to: input the sub-graph into the enterprise asset classification model, predict each edge and each sub-path in the sub-graph through the enterprise asset classification model, wherein each sub-path is a path from a starting node to any node on the sub-graph, and obtain the prediction result of the enterprise asset classification model; the prediction result is used to indicate whether the parameters corresponding to each node in the sub-graph belong to the enterprise to be detected.

10. An electronic device, characterized in that: include: A memory for storing program instructions; A processor is used to call the program instructions stored in the memory, and execute the method according to any one of claims 1 to 6 or the steps included in the method according to claim 7 according to the obtained program instructions.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to have a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the method according to any one of claims 1 to 6 or the method according to claim 7 is implemented.

Citation Information

Cited By

  • Vulnerability maintenance method for heterogeneous resources in cloud environment and related equipment

    CN121744318A