Network asset identification method and device, electronic equipment and storage medium

By building URLs and calculating editing distances, combining network asset recognition models, low correlation text is filtered, and the problem of low accuracy of network asset recognition is solved and more efficient network asset recognition is achieved.

CN120238342APending Publication Date: 2025-07-01CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510368673.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, network asset identification methods cannot effectively distinguish invalid or redundant information, resulting in low identification accuracy and inaccurate identification of the organization's network assets.

Method used

By receiving the terminal's network asset identification request, using the DNS database to build a URL, extract the keywords of the target organization name and web page data text, calculate the correlation score based on the editing distance, filter the low-correlation text, and combine the network asset identification model to determine whether the web page data text belongs to the target organization.

Benefits of technology

It improves the accuracy and efficiency of network asset identification, reduces interference with invalid information, and ensures the quality and reliability of identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238342A_ABST
    Figure CN120238342A_ABST
Patent Text Reader

Abstract

The invention discloses a network asset identification method and device, electronic equipment and a storage medium. The method comprises the following steps: receiving a network asset identification request sent by a terminal; constructing URLs based on domain names, IP addresses and preset port numbers in a preset DNS database, and obtaining webpage data texts corresponding to the URLs; extracting a first keyword in the target organization name information and a second keyword of each webpage data text; for each second keyword of each webpage data text, determining a correlation score between the second keyword and all the first keywords according to an editing distance between each first keyword and the second keyword; extracting context statements corresponding to target second keywords of which the correlation scores with all the first keywords are greater than a preset score threshold in the webpage data text; and determining whether the webpage data text belongs to the target organization or not based on the target organization name information, the context statement corresponding to the target second keyword and a network asset identification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a network asset identification method, device, electronic device and storage medium. Background Art

[0002] Identifying an organization's network assets is a core task of organizational management and security protection. With the rapid development of the Internet, the online business of many organizations continues to expand, including not only the official website of the organization, but also multiple sub-sites, historical pages, cooperation project websites, etc. These websites may be distributed on different domain names and sub-domains, or form "dark assets" on hosting platforms that the organization cannot fully control. These corporate network assets that are not effectively managed often bring multiple hidden dangers, including potential security vulnerabilities, outdated content or system vulnerabilities, which are easily exploited by attackers. Identifying all the network assets of an organization exposed on the Internet and effectively managing them is crucial to protecting the information security of the organization.

[0003] In the related art, the commonly used network asset identification methods only rely on simple web crawling and static analysis, but cannot distinguish a large amount of invalid or redundant information. Organization websites usually contain non-core content such as advertisements, navigation bars, templated information, etc. The interference of this information often causes valid information to be submerged, thus affecting the accuracy of network asset identification.

[0004] Therefore, how to improve the accuracy of network asset identification is one of the technical problems that need to be urgently solved in the existing technology. Summary of the invention

[0005] In order to solve the problem of low accuracy in network asset identification, embodiments of the present application provide a network asset identification method, device, electronic device, and storage medium.

[0006] In a first aspect, an embodiment of the present application provides a network asset identification method, including:

[0007] Receiving a network asset identification request sent by a terminal, wherein the network asset identification request carries target organization name information;

[0008] Constructing a uniform resource locator URL based on the domain name in a preset domain name system DNS database, the network protocol IP address corresponding to the domain name and the preset port number, and obtaining the web page data text corresponding to each URL;

[0009] Extracting a first keyword from the target organization name information and extracting a second keyword from each web page data text;

[0010] For each second keyword included in each web page data text, determine the correlation score between the second keyword and all first keywords according to the edit distance between each first keyword and the second keyword. The correlation score between the second keyword and all first keywords represents the degree of association between the second keyword and the target organization name information;

[0011] For each web page data text, extract the context statement corresponding to the target second keyword whose correlation score with all first keywords is greater than a preset score threshold;

[0012] Based on the target organization name information, the context statement corresponding to the target second keyword, and a network asset identification model, determine whether the web page data text belongs to the target organization. The network asset identification model is used to identify whether the web page data text is a network asset of the target organization.

[0013] In one implementation, after extracting the second keywords of each web page data text, it further includes:

[0014] Expand the second keyword of each web page data text based on the synonyms of the second keyword of each web page data text.

[0015] In one implementation, constructing a URL based on the domain name, the IP address corresponding to the domain name, and the preset port number in the preset DNS database specifically includes:

[0016] Construct a first URL based on the domain name in the preset DNS database;

[0017] Construct a second URL based on the IP address corresponding to the domain name in the preset DNS database and the preset port number;

[0018] Obtain the web page data text corresponding to each URL, specifically including:

[0019] Obtain the web page data text corresponding to each first URL and each second URL respectively.

[0020] In one implementation, extracting the first keyword in the target organization name information specifically includes:

[0021] Input the target organization name information into a large language model to obtain the first keyword in the target organization name information. The large language model is used to extract the first keyword in the target organization name information.

[0022] In one implementation, extracting the second keyword of each web page data text specifically includes:

[0023] Perform word segmentation on each of the web page data texts to obtain the words contained in each of the web page data texts;

[0024] Determine the weights of the various words contained in each web page data text, where the weight of a word represents the importance of the word in the web page data text;

[0025] Determine the words with weights greater than a preset weight threshold as the second keywords of the corresponding web page data texts.

[0026] In one implementation, determining the weights of the various words contained in each web page data text specifically includes:

[0027] For each word contained in each web page data text, determine the term frequency-inverse document frequency corresponding to the word according to the number of times the word appears in the web page data text, the number of words contained in the web page data text, the number of web page data texts, and the number of web page data texts containing the word;

[0028] Determine the term frequency-inverse document frequency corresponding to the word as the weight of the word.

[0029] In one implementation, for each second keyword contained in each web page data text, determine the correlation score between the second keyword and all first keywords according to the edit distance between each first keyword and the second keyword, specifically including:

[0030] Calculate the correlation score between the second keyword and all first keywords through the following formula:

[0031]

[0032] where S i represents the correlation score between the i-th second keyword in the web page data text and all first keywords;

[0033] o i represents the i-th second keyword in the web page data text;

[0034] c j represents the j-th first keyword in the target organization name information, where j = 1, 2,..., m, and m represents the number of first keywords;

[0035] Levenshtein(c j ,o i ) represents the edit distance between the j-th first keyword c j in the target organization name information and the i-th second keyword o i in the web page data text;

[0036] L(oi ) represents the i-th second keyword o in the web page data text i in length;

[0037] (L(c j ) represents the j-th first keyword c in the target organization name information j in length.

[0038] In one implementation, the network asset identification model includes a text splicing processing layer, a multi-granularity encoding and aggregation layer, a first feed-forward neural network layer, and an output layer;

[0039] Based on the target organization name information, the context statements corresponding to the target second keywords, and the network asset identification model, determining whether the web page data text belongs to the target organization specifically includes:

[0040] Input the target organization name information and the context statements corresponding to each target second keyword into the text splicing processing layer;

[0041] Through the text splicing processing layer, splice the target organization name information with the context statements corresponding to all target second keywords to obtain a spliced text, and splice the target organization name information with the context statements of each target second keyword respectively to obtain corresponding spliced statements, and send the spliced text and each spliced statement to the multi-granularity encoding and aggregation layer;

[0042] Through the multi-granularity encoding and aggregation layer, extract the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement, perform semantic aggregation based on the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement to obtain the target semantic vector of the spliced statement, and send the comprehensive semantic feature vector obtained by splicing the semantic feature vector of the spliced text and the target semantic feature vector of the spliced statement to the first feed-forward neural network layer;

[0043] Through the first feed-forward neural network layer, perform feature transformation on the comprehensive semantic feature vector to obtain a target comprehensive semantic feature vector, and output a target prediction probability value through the output layer, where the target prediction probability value represents the probability that the web page data text belongs to the target organization;

[0044] Determine whether the web page data text belongs to the target organization according to the target prediction probability value.

[0045] In one implementation, the multi-granularity encoding and aggregation layer includes a multi-head attention network layer and a second feed-forward neural network layer;

[0046] Extract the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement through the multi-granularity encoding and aggregation layer, and perform semantic aggregation based on the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement to obtain the target semantic vector of the spliced statement, specifically including:

[0047] Extract the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement through the multi-head attention network layer;

[0048] Perform feature transformation on the semantic feature vector of the spliced text through the second feed-forward neural network layer to obtain the target feature vector of the spliced text;

[0049] Perform dot product of the target feature vector of the spliced text with the semantic feature vectors of each spliced statement respectively to obtain the weights corresponding to the semantic feature vectors of the corresponding spliced statements;

[0050] Normalize the weights corresponding to the semantic feature vectors of each spliced statement, and perform weighted average of the normalized weights of the semantic feature vectors of each spliced statement and the semantic feature vectors of each spliced statement to obtain the target semantic vector of the spliced statement.

[0051] In a second aspect, an embodiment of the present application provides a network asset identification device, including:

[0052] A receiving module, configured to receive a network asset identification request sent by a terminal, where the network asset identification request carries target organization name information;

[0053] An obtaining module, configured to construct a uniform resource locator (URL) based on domain names, network protocol IP addresses corresponding to the domain names, and preset port numbers in a preset domain name system (DNS) database, and obtain web data texts corresponding to each URL;

[0054] A keyword extraction module, configured to extract a first keyword from the target organization name information and a second keyword from each web data text;

[0055] A determination module, configured to, for each second keyword included in each web data text, determine a correlation score between the second keyword and all first keywords according to the edit distance between each first keyword and the second keyword, and the correlation score between the second keyword and all first keywords represents the degree of association between the second keyword and the target organization name information;

[0056] A context extraction module, configured to, for each web data text, extract context statements corresponding to target second keywords in the web data text whose correlation scores with all first keywords are greater than a preset score threshold.

[0057] An identification module, configured to determine whether the web page data text belongs to the target organization based on the target organization name information, the context statement corresponding to the target second keyword, and a network asset identification model, where the network asset identification model is used to identify whether the web page data text is a network asset of the target organization.

[0058] In one implementation, the apparatus further includes:

[0059] An expansion module, configured to expand the second keyword of each web page data text based on a synonym of the second keyword of each web page data text after extracting the second keyword of each web page data text.

[0060] In one implementation, the obtaining module is specifically configured to construct a first URL based on the domain name in the preset DNS database; construct a second URL based on the IP address corresponding to the domain name in the preset DNS database and a preset port number; and obtain the web page data text corresponding to each first URL and each second URL respectively.

[0061] In one implementation, the keyword extraction module is specifically configured to input the target organization name information into a large language model to obtain a first keyword in the target organization name information, where the large language model is used to extract the first keyword in the target organization name information.

[0062] In one implementation, the keyword extraction module is specifically configured to perform word segmentation on each web page data text to obtain the words included in each web page data text; determine the weight of each word included in each web page data text, where the weight of the word represents the importance degree of the word in the web page data text; and determine the word with a weight greater than a preset weight threshold as the second keyword of the corresponding web page data text.

[0063] In one implementation, the keyword extraction module is specifically configured to, for each word included in each web page data text, determine the term frequency-inverse document frequency corresponding to the word according to the number of times the word appears in the web page data text, the number of words included in the web page data text, the number of web page data texts, and the number of web page data texts including the word; and determine the term frequency-inverse document frequency corresponding to the word as the weight of the word.

[0064] In one implementation, the determining module is specifically configured to calculate a correlation score between the second keyword and all first keywords through the following formula:

[0065]

[0066] Where Si Represents the relevance score of the i-th second keyword in the web page data text to all the first keywords;

[0067] o i Represents the i-th second keyword in the web page data text;

[0068] c j Represents the j-th first keyword in the target organization name information, where j = 1, 2, ……, m, and m represents the number of first keywords;

[0069] Levenshtein(c j ,o i ) represents the j-th first keyword c in the target organization name information j and the i-th second keyword o in the web page data text i The edit distance between them;

[0070] L(o i ) represents the length of the i-th second keyword o in the web page data text i ;

[0071] (L(c j ) represents the length of the j-th first keyword c in the target organization name information j ;

[0072] In one implementation, the network asset identification model includes a text splicing processing layer, a multi-granularity encoding and aggregation layer, a first feedforward neural network layer, and an output layer;

[0073] The recognition module is specifically configured to input the target organization name information and the context statements corresponding to each target second keyword into the text splicing processing layer; splice the target organization name information with the context statements corresponding to all target second keywords through the text splicing processing layer to obtain a spliced text, and splice the target organization name information with the context statements of each target second keyword respectively to obtain corresponding spliced statements, and send the spliced text and each spliced statement to the multi-granularity encoding and aggregation layer; extract the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement through the multi-granularity encoding and aggregation layer, perform semantic aggregation based on the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement to obtain the target semantic vector of the spliced statement, and send the comprehensive semantic feature vector obtained by splicing the semantic feature vector of the spliced text and the target semantic feature vector of the spliced statement to the first feed-forward neural network layer; perform feature transformation on the comprehensive semantic feature vector through the first feed-forward neural network layer to obtain a target comprehensive semantic feature vector, and output a target prediction probability value through the output layer, where the target prediction probability value represents the probability that the web page data text belongs to the target organization; determine whether the web page data text belongs to the target organization according to the target prediction probability value.

[0074] In one implementation, the multi-granularity encoding and aggregation layer includes a multi-head attention network layer and a second feed-forward neural network layer;

[0075] The recognition module is specifically configured to extract the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement through the multi-head attention network layer; perform feature transformation on the semantic feature vector of the spliced text through the second feed-forward neural network layer to obtain the target feature vector of the spliced text; perform dot product of the target feature vector of the spliced text with the semantic feature vector of each spliced statement respectively to obtain the weight corresponding to the semantic feature vector of the corresponding spliced statement; normalize the weight corresponding to the semantic feature vector of each spliced statement, and perform weighted average of the normalized weights of the semantic feature vectors of each spliced statement and the semantic feature vectors of each spliced statement to obtain the target semantic vector of the spliced statement.

[0076] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the network asset recognition method described in the present application when executing the program.

[0077] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the network asset identification method described in the present application are implemented.

[0078] The beneficial effects of the present application are as follows:

[0079] The network asset identification method, device, electronic device and storage medium provided by the embodiments of this application. The server receives a network asset identification request sent by a terminal, and the network asset identification request carries target organization name information; constructs a URL (Uniform Resource Locator) based on the domain names, IP (Internet Protocol) addresses corresponding to the domain names, and preset port numbers in a preset DNS (Domain Name System) database, and obtains the web data text corresponding to each URL; extracts the first keywords in the target organization name information and the second keywords in each web data text; for each second keyword included in each web data text, determines the correlation score between the second keyword and all the first keywords according to the edit distance between each first keyword and the second keyword, and the correlation score between the second keyword and all the first keywords represents the degree of association between the second keyword and the target organization name information; for each web data text, extracts the context statements corresponding to the target second keywords whose correlation scores with all the first keywords are greater than a preset score threshold in the web data text; determines whether the web data text belongs to the target organization based on the target organization name information, the context statements corresponding to the target second keywords, and a network asset identification model, and the network asset identification model is used to identify whether the web data text is a network asset of the target organization. In the embodiments of this application, when the server identifies the network assets of the target organization, it first comprehensively constructs possible URLs according to the existing domain names, IP addresses corresponding to the domain names, and preset port numbers in the DNS database, obtains the web data text corresponding to each constructed URL, extracts the first keywords in the target organization name information and the second keywords in each web data text, and then determines the correlation score between each second keyword and all the first keywords according to the edit distance between each first keyword in the target organization name information and each second keyword in each web data text, so as to determine the degree of association between each second keyword and the target organization name information, and only extracts the context statements corresponding to the target second keywords with high correlation with the target organization name information as the analysis text, and predicts whether each web data text is a network asset of the target organization based on the target organization name information, the context statements corresponding to the extracted target second keywords, and the trained network asset identification model. Thus, it effectively filters out the web data text segments with low correlation with the target organization name information, greatly reduces the interference of invalid information, improves the quality of the analysis text, and improves the accuracy of network asset identification. Moreover, based on the pre-trained network asset identification model, the attribution of the web data text is automatically identified, further improving the accuracy and efficiency of network asset identification.

[0080] Other features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present application. The objectives and other advantages of the present application may be realized and attained by the structure particularly pointed out in the written description, claims, as well as the drawings. Description of the Drawings

[0081] The drawings described herein are provided to further understand the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0082] Figure 1 It is a schematic diagram of an application scenario of the network asset identification method provided by an embodiment of the present application;

[0083] Figure 2 It is a schematic diagram of the flow of the network asset identification method provided by an embodiment of the present application;

[0084] Figure 3 It is a schematic diagram of the flow of extracting the second keyword of each web page data text provided by an embodiment of the present application;

[0085] Figure 4 It is a schematic diagram of the structure of the preset network model provided by an embodiment of the present application;

[0086] Figure 5 It is a schematic diagram of the training process of the network asset identification model provided by an embodiment of the present application;

[0087] Figure 6 It is a schematic diagram of the process of obtaining the target semantic vector of the sample splicing statement provided by an embodiment of the present application;

[0088] Figure 7 It is a schematic diagram of the process of determining whether the web page data text belongs to the target organization provided by an embodiment of the present application;

[0089] Figure 8 It is a schematic diagram of the process of obtaining the target semantic vector of the splicing statement provided by an embodiment of the present application;

[0090] Figure 9 It is a schematic diagram of the structure of the network asset identification device provided by an embodiment of the present application;

[0091] Figure 10 It is a schematic diagram of the structure of the electronic device provided by an embodiment of the present application. Detailed Embodiments

[0092] To solve the problem of low accuracy in network asset identification, embodiments of the present application provide a network asset identification method, device, electronic device, and storage medium.

[0093] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0094] First, refer to Figure 1, which is a schematic diagram of an application scenario of the network asset identification method provided in the embodiments of the present application. It may include a terminal 101 and a server 102. The terminal 101 sends a network asset identification request to the server 102. The network asset identification request carries target organization name information. The server 102 receives the network asset identification request sent by the terminal 101, constructs a URL based on the domain names, IP addresses corresponding to the domain names, and preset port numbers in the preset DNS database, and obtains the web page data text corresponding to each URL. Extract the first keywords in the target organization name information and the second keywords in each web page data text. For each second keyword included in each web page data text, determine the correlation score between the second keyword and all the first keywords according to the edit distance between each first keyword and the second keyword. The correlation score between the second keyword and all the first keywords represents the degree of association between the second keyword and the target organization name information. For each web page data text, extract the context statements corresponding to the target second keywords whose correlation scores with all the first keywords are greater than the preset score threshold in the web page data text. Furthermore, the server 102 determines whether the web page data text belongs to the target organization based on the target organization name information, the context statements corresponding to the target second keywords, and the network asset identification model. The network asset identification model is used to identify whether the web page data text is the network asset of the target organization. In the embodiments of the present application, when the server 102 identifies the network assets of the target organization, it first comprehensively constructs possible existing URLs according to the domain names, IP addresses corresponding to the domain names, and preset port numbers in the DNS database, obtains the web page data text corresponding to each constructed URL, extracts the first keywords in the target organization name information and the second keywords in each web page data text, and then determines the correlation score between each second keyword and all the first keywords according to the edit distance between each first keyword in the target organization name information and each second keyword in each web page data text, so as to determine the degree of association between each second keyword and the target organization name information. Only extract the context statements corresponding to the target second keywords with high correlation with the target organization name information as the analysis text, and predict whether each web page data text is the network asset of the target organization based on the target organization name information, the context statements corresponding to the extracted target second keywords, and the trained network asset identification model. Thus, it effectively filters out the web page data text segments with low correlation with the target organization name information, greatly reduces the interference of invalid information, improves the quality of the analysis text, and improves the accuracy of network asset identification. Moreover, based on the pre-trained network asset identification model, the attribution of the web page data text is automatically identified, further improving the accuracy and efficiency of network asset identification.

[0095] The server 102 can be an independent physical server or a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, and cloud storage. The terminal 101 can be, but is not limited to, a smart phone, a tablet computer, a laptop computer, a desktop computer, etc. The server 102 and the terminal 101 can be connected through a network, and the embodiments of the present application do not limit this.

[0096] Based on the above application scenarios, the exemplary embodiments of the present application will be described in more detail below. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited by any of them here. On the contrary, the embodiments of the present application can be applied to any applicable scenario. Figures 2 to 8 As shown in

[0097] FIG. Figure 2 10, which is a schematic flowchart of the implementation process of the network asset identification method provided by the embodiment of the present application. The network asset identification method can be applied to the above-mentioned server 102 and specifically may include the following steps:

[0098] S21. Receive a network asset identification request sent by a terminal. The network asset identification request carries target organization name information.

[0099] Specifically, when implemented, the server receives a network asset identification request sent by a terminal. The network asset identification request carries target organization name information. The target organization is the organization of the network assets to be identified and can be, but is not limited to, any organization such as an enterprise, an institution, or a public institution. The embodiments of the present application do not limit this.

[0100] S22. Construct a URL based on the domain names, the IP addresses corresponding to the domain names, and the preset port numbers in a preset DNS database, and obtain the web page data text corresponding to each URL.

[0101] Specifically, when implemented, the server constructs a first URL based on the domain names in the preset DNS database; constructs a second URL based on the IP addresses corresponding to the domain names in the preset DNS data and the preset port numbers.

[0102] Specifically, the server pre-collects the domain names existing in the entire network and the IP addresses corresponding to the domain names, and stores the domain names and the IP addresses corresponding to the domain names in a preset DNS database. When identifying the network assets of the target organization, the server extracts each domain name in the preset DNS database, and adds prefixes: "http: / / " and "https: / / " to each domain name according to network protocols such as http and https to construct a URL based on the domain name, which can be denoted as the first URL. For example, the domain name domain1 can generate two corresponding first URLs: "http: / / domain1" and "https: / / domain1". The server splices the IP address corresponding to the domain name in the preset DNS database with suspicious web ports (such as: 80, 8000, 8888, 8889, 443, 6443, etc.), and then adds prefixes "http: / / " and "https: / / " to generate a URL based on the IP address and port, which can be denoted as the second URL. For example, the second URLs generated by splicing the IP address "192.168.0.0" and the port "80" and then adding prefixes "http: / / " and "https: / / " are: "http: / / 192.168.0.0:80" and "https: / / 192.168.0.0:80", and the second URLs generated by splicing the IP address "192.168.0.0" and the port "8000" and then adding prefixes "http: / / " and "https: / / " are: "http: / / 192.168.0.0:8000" and "https: / / 192.168.0.0:8000". In this way, the URLs that may exist in the entire network can be constructed more comprehensively based on the domain name, as well as based on the IP address and port, so as to improve the comprehensiveness and accuracy of network asset identification.

[0103] Furthermore, the server obtains the web data texts corresponding to each first URL and each second URL respectively.

[0104] Specifically, the server can obtain the web data texts corresponding to each first URL and each second URL constructed through a crawler tool, that is, the website data texts corresponding to each URL.

[0105] S23. Extract the first keywords in the target organization name information and extract the second keywords in each web data text.

[0106] In specific implementation, the first keyword in the target organization name information can be extracted in the following way: The server inputs the target organization name information into a large language model (LLM) to obtain the keywords in the target organization name information, which can be denoted as the first keyword. The large language model is used to extract the first keyword in the target organization name information. The large language model is a deep learning model that obtains extensive general knowledge through training based on a large amount of text data. By learning the language patterns and semantics of the text data, the large language model can deeply understand the text meaning and can be used to process complex language understanding tasks. The large language model can use, but is not limited to, any model in the following series of models: Qwen series models, GLM series models, DeepSeek series models, etc. The embodiments of the present application do not make any limitations in this regard. Organization names have a general pattern. Based on prompts, the large language model can accurately identify the keywords in the organization name without training.

[0107] Specifically, the server can input the organization name information and the operation instruction text into the large language model, and output each first keyword included in the organization name information. The operation instruction text can be "extract keywords".

[0108] In specific implementation, the server performs word segmentation processing on each obtained web page data text, and determines the keywords of each web page data text based on the weights of the words included in each web page data text, which can be denoted as the second keyword.

[0109] Specifically, it can be in accordance with Figure 3 the shown process to extract the second keyword of each web page data text, which may include the following steps:

[0110] S31. Perform word segmentation processing on each web page data text to obtain the words included in each web page data text.

[0111] In specific implementation, the web page text data includes the website title and the website body. The server can use a word segmentation tool to perform word segmentation processing on the website title and the website body in each web page data text to obtain the words included in each web page data text. Among them, the word segmentation tool can use, but is not limited to, any one of the following word segmentation tools: Jieba word segmentation tool, Stanford tool, etc. Any other word segmentation tool can also be used. The embodiments of the present application do not make any limitations in this regard.

[0112] In one implementation manner, the server can use the first keywords of the target organization name information as the seeds of the word segmentation tool to overcome the problem that the word segmentation tool cannot recognize unseen words.

[0113] S32. Determine the weights of the words included in each web page data text.

[0114] In specific implementation, for each word included in each web page data text, according to the number of times the word appears in the web page data text, the number of words included in the web page data text, the number of web page data texts, and the number of web page data texts including the word, the term frequency-inverse document frequency (TF-IDF) corresponding to the word is determined, and the term frequency-inverse document frequency corresponding to the word is determined as the weight of the word. The weight of the word represents the importance degree of the word in the web page data text, that is: the term frequency-inverse document frequency corresponding to the word represents the importance degree of the word in the web page data text.

[0115] Specifically, for any web page data text seq, its word segmentation result can be marked as [w1,..., w n , where w1,..., w n are the 1st to nth words included in the web page data text seq. The term frequency-inverse document frequency of the ith word w i in the web page data text seq can be calculated through the following formula:

[0116]

[0117] where TF-IDF i represents the term frequency-inverse document frequency corresponding to the ith word w i in the web page data text;

[0118] represents the number of times the ith word w i appears in the web page data text;

[0119] C seq represents the number of words included in the web page data text (C seq =n), and C doc represents the number of web page data texts;

[0120] C wi_doc represents the number of web page data texts including the ith word w i of the web page data text.

[0121] The term frequency-inverse document frequency TF-IDF i corresponding to the ith word w i of the web page data text is used as the weight of the ith word w i of the web page data text. In this way, the weights of each word included in each web page data text can be calculated.

[0122] S33. Determine the words with weights greater than the preset weight threshold as the second keywords of the corresponding web page data text.

[0123] In specific implementation, for each web page data text, the words with weights greater than a preset weight threshold are determined as the second keywords of the web page data text, where the preset weight threshold can be set according to actual needs. For example, it can be set to 0.5, or any value between (0, 1). The embodiments of the present application do not limit this. The higher the weight of a word, the more important it is in the web page data text.

[0124] In one implementation, it is also possible to select a preset proportion of words as the second keywords in the order of decreasing weight. For example, 50% of the words with high weights can be selected as the second keywords. For example, if a web page data text contains n words, the first n / 2 words with high weights can be selected as the second keywords. The embodiments of the present application do not limit this.

[0125] In one implementation, after extracting the second keywords of each web page data text, the server can also expand the second keywords of each web page data text based on the synonyms of the second keywords of each web page data text. In this way, by using words with semantic similarities to the second keywords to expand the second keywords, it is possible to avoid missing the second keywords with a relatively high correlation with the target enterprise name when determining the correlation between the second keywords and the target enterprise name.

[0126] In specific implementation, the server pre-stores a preset thesaurus, and the preset thesaurus contains the corresponding relationship between words and word vectors. The server uses a word vector extraction model to obtain the word vectors of each second keyword of each web page data text. For the word vectors of each second keyword of each web page data text, calculate the similarity between the word vector of the second keyword and each first word vector in the preset thesaurus, and determine the word corresponding to the first word vector with a similarity greater than the set threshold as the synonym of the second keyword. In implementation, the cosine similarity between the word vector of the second keyword and the first word vector can be calculated to obtain the cosine similarity between the two, or the Euclidean distance between the word vector of the second keyword and the first word vector can be calculated to obtain the similarity between the two. Any other method for calculating similarity can also be used to calculate the similarity between the two. The embodiments of the present application do not limit this. The set threshold can be set according to actual needs. For example, the set threshold can be set to 0.7. The embodiments of the present application do not limit this.

[0127] For example, "Telecom" is semantically similar to "Tianyi" in the preset thesaurus. The second keyword "Telecom Security" extracted from the website title "China Telecom Security Cloud Dyke Official Website" can be expanded to a new second keyword "Tianyi Security". Furthermore, when determining the correlation score between the second keyword and the first keyword subsequently, the second keyword includes the expanded second keyword.

[0128] S24. For each second keyword included in each web page data text, determine the correlation score between the second keyword and all the first keywords according to the edit distance between each first keyword and the second keyword.

[0129] Specifically, when implemented, the correlation score between the second keyword and all the first keywords represents the degree of association between the second keyword and the target organization name information.

[0130] For each second keyword included in each web page data text, calculate the correlation score between the second keyword and all the first keywords through the following formula:

[0131]

[0132] where S i represents the correlation score between the i-th second keyword in the web page data text and all the first keywords;

[0133] o i represents the i-th second keyword in the web page data text;

[0134] c j represents the j-th first keyword in the target organization name information, where j = 1, 2,..., m, and m represents the number of first keywords;

[0135] Levenshtein(c j ,o i ) represents the edit distance between the j-th first keyword c j in the target organization name information and the i-th second keyword o i in the web page data text;

[0136] L(o i ) represents the length of the i-th second keyword o i in the web page data text;

[0137] (L(c j ) represents the length of the j-th first keyword c j in the target organization name information.

[0138] The edit distance refers to the minimum number of single-character edit operations required to convert one string into another string. These edit operations include inserting, deleting, and replacing characters. The edit distance is a method to measure the similarity between two strings. The edit distance between the j-th first keyword c j in the target organization name information and the i-th second keyword o i in the web page data text represents the conversion of the first keyword c j into the second keyword o iThe minimum number of required editing operations. When the second keyword o i is in Chinese, the length of the second keyword o i refers to the number of Chinese characters of the second keyword o i . When the second keyword o i is a string such as English and / or numbers, the length of the second keyword o i refers to the string length of the second keyword o i . The length of the first keyword c j is similar and will not be elaborated.

[0139] S25. For each web page data text, extract the context sentences corresponding to the target second keywords in the web page data text whose relevance scores with all the first keywords are greater than the preset score threshold.

[0140] Specifically in implementation, for each web page data text, the server sorts the relevance scores of each second keyword in the web page data text with all the first keywords in descending order, takes the second keywords in the web page data text whose relevance scores with all the first keywords are greater than the preset score threshold as the target second keywords, extracts the context sentences corresponding to each target second keyword, and discards the other sentences in the web page data text. Thus, redundant texts with low relevance to the target organization name in the web page data text are filtered to improve the accuracy and efficiency of network asset recognition. Among them, the preset score threshold can be set according to requirements, such as 0.6 or 0.65, or other values, which are not limited in the embodiments of this application.

[0141] When extracting the context sentences corresponding to the target second keywords, it is possible to extract the sentence where the target second keyword belongs, the previous sentence and the next sentence of the sentence where the target second keyword belongs, or it is also possible to extract the sentence where the target second keyword belongs, multiple sentences before and after the sentence where the target second keyword belongs. This is not limited in the embodiments of this application.

[0142] In implementation, if it is determined that any keyword in the web page data text whose relevance scores with all the first keywords are greater than the preset score threshold is an expansion word, then the original word in the web page data text corresponding to the expansion word is determined as the target second keyword.

[0143] S26. Based on the target organization name information, the context sentences corresponding to the target second keywords, and the network asset recognition model, determine whether the web page data text belongs to the target organization.

[0144] In specific implementation, for each web page data text, the server inputs the target organization name information and the context statement corresponding to the target second keyword extracted from the web page data text into the network asset recognition model to obtain the probability that the web page data text belongs to the target organization, which can be denoted as the target prediction probability value. It is determined whether the web page data text belongs to the target organization according to the target prediction probability value. The network asset recognition model is used to identify whether the web page data text is the network asset of the target organization. The network asset recognition model includes a text splicing processing layer, a multi-granularity encoding and aggregation layer, a first feed-forward neural network layer, and an output layer. The granularity encoding and aggregation layer includes a multi-head attention network layer and a second feed-forward neural network layer.

[0145] In implementation, the server trains the network asset recognition model according to a preset network model and sample data. The structural schematic diagram of the preset network model is as Figure 4 shown. The preset network model includes a text splicing processing layer, a multi-granularity encoding and aggregation layer, a first feed-forward neural network (Feed-Forward Neural Network, FFN) layer, and an output layer. The multi-granularity encoding and aggregation layer includes a multi-head attention network layer and a second feed-forward neural network layer. Both the first feed-forward neural network layer and the second feed-forward neural network layer adopt the feed-forward neural network model structure.

[0146] The training process of the network asset recognition model is as Figure 5 shown and may include the following steps:

[0147] S41. Obtain training samples, where the training samples include sample web page data texts corresponding to the sample organization name and the domain name of the sample organization.

[0148] In specific implementation, the server can construct training samples in the following way: Use ICP (Internet Content Provider) filing data to obtain the organization names and main domain names of multiple different organizations. Among them, the ICP filing data includes information such as the organization name, the domain name of the organization, and the website name. Based on the obtained main domain names of different organizations, use a crawler tool to obtain the corresponding web page data texts. Take the organization name and the obtained web page data text of the organization as positive samples, and take the organization name and the web page data text of other organizations as negative samples to obtain training samples, that is: the training samples include sample web page data texts corresponding to the sample organization name and the domain name of the sample organization. The ratio of positive samples to negative samples also adopts a ratio of 1:1 to ensure the effectiveness of training. The ratio of positive samples to negative samples can also adopt any other ratio, and this application embodiment does not limit this.

[0149] S42. Extract the first keyword in each sample group name and extract the second keyword in each sample web page data text.

[0150] For the implementation of this step, refer to the implementation of step S23, which will not be elaborated here.

[0151] S43. For each second keyword included in each sample web page data text, determine the relevance score between the second keyword and all first keywords according to the edit distance between each first keyword of the sample group name and the second keyword.

[0152] For the implementation of this step, refer to the implementation of step S24, which will not be elaborated here.

[0153] S44. For each sample web page data text, extract the context statements corresponding to the target second keywords whose relevance scores between the sample web page data text and all first keywords are greater than the preset score threshold.

[0154] For the implementation of this step, refer to the implementation of step S25, which will not be elaborated here.

[0155] S45. For each sample web page data text, input the sample organization name and the context statements corresponding to each target second keyword into the text splicing processing layer in the preset network model.

[0156] Specifically in implementation, for each sample web page data text, the server inputs the sample organization name and the context statements of each target second keyword into the text splicing processing layer in the preset network model.

[0157] S46. Through the text splicing processing layer, splice the sample organization name with the context statements corresponding to all target keywords to obtain a sample splicing text, and splice the sample organization name with the context statements of each target second keyword respectively to obtain corresponding sample splicing statements. Send the sample splicing text and each sample splicing statement to the multi-granularity encoding and aggregation layer.

[0158] Specifically in implementation, through the text splicing processing layer, splice the sample organization name with the context statements corresponding to all target keywords to obtain a sample splicing text, and splice the sample organization name with the context statements of each target second keyword respectively to obtain corresponding sample splicing statements. The two statements can be connected by a delimiter [seq]. The spliced sample splicing text and each sample splicing statement are as Figure 4 shown. Furthermore, send the sample splicing text and each sample splicing statement to the multi-head attention network layer in the multi-granularity encoding and aggregation layer. The multi-head attention network layer is used to extract the semantic feature vectors of the text.

[0159] S47. Extract the semantic feature vectors of the sample concatenated text and the semantic feature vectors of each sample concatenated statement through the multi-granularity encoding and aggregation layer. Perform semantic aggregation based on the semantic feature vector of the sample concatenated text and the semantic feature vectors of each sample concatenated statement to obtain the target semantic vector of the sample concatenated statement, and send the comprehensive semantic feature vector obtained by concatenating the semantic feature vector of the sample concatenated text and the target semantic feature vector of the sample concatenated statement to the first feed-forward neural network layer.

[0160] Specifically, during implementation, it can be obtained according to the process as Figure 6 shown to obtain the target semantic vector of the sample concatenated statement, including the following steps:

[0161] S51. Extract the semantic feature vectors of the sample concatenated text and the semantic feature vectors of each sample concatenated statement through the multi-head attention network layer.

[0162] Specifically, during implementation, after the server inputs the sample concatenated text and each sample concatenated statement into the multi-head attention network layer, the multi-head attention network layer extracts the semantic feature vector of the sample concatenated text to complete the semantic modeling of the stable granularity, and the multi-head attention network layer extracts the semantic feature vectors of each sample concatenated statement to complete the semantic modeling of the sentence granularity. Furthermore, the semantic feature vector of the sample concatenated text is input into the second feed-forward neural network layer.

[0163] S52. Perform feature transformation on the semantic feature vector of the sample concatenated text through the second feed-forward neural network layer to obtain the target feature vector of the sample concatenated text.

[0164] Specifically, during implementation, the feed-forward neural network is used to perform feature transformation on the semantic feature vector of the text, convert it to a new feature space, and obtain a new feature vector representation of the text.

[0165] S53. Perform dot product of the target feature vector of the sample concatenated text with the semantic feature vector of each sample concatenated statement respectively to obtain the weight corresponding to the semantic feature vector of the corresponding sample concatenated statement.

[0166] Specifically, during implementation, the weight corresponding to the semantic feature vector of any sample concatenated statement can be calculated through the following formula:

[0167]

[0168] where, ω i represents the weight corresponding to the semantic feature vector of the i-th sample concatenated statement, i = 1, 2,..., k, and k represents the number of sample concatenated statements;

[0169] represents the semantic feature vector of the sample concatenated text, Represents the target feature vector of the sample splicing text;

[0170] Represents the semantic feature vector of the i-th sample splicing statement.

[0171] S54. Normalize the weights corresponding to the semantic feature vectors of each sample splicing statement, and perform weighted averaging on the normalized weights of the semantic feature vectors of each sample splicing statement and the semantic feature vectors of each splicing statement to obtain the target semantic vector of the sample splicing statement.

[0172] In specific implementation, the server performs a softmax operation on the weights corresponding to the semantic feature vectors of each sample splicing statement to normalize the weights corresponding to the semantic feature vectors of each sample splicing statement.

[0173] Specifically, the weights corresponding to the semantic feature vectors of any sample splicing statement can be normalized through the following formula to obtain the normalized weights of the semantic feature vectors of the sample splicing statement:

[0174] ω i ′ = softmax(ω i ), i = 1, 2, ……, k

[0175] where, ω i ′ Represents the normalized weight of the semantic feature vector of the i-th sample splicing statement.

[0176] Calculate the target semantic vector of the sample splicing statement through the following formula:

[0177]

[0178] where, Represents the target semantic vector of the sample splicing statement.

[0179] The target semantic vector of the sample splicing statement is a comprehensive sentence-level semantic vector.

[0180] The target semantic vector of the sample splicing statement can also be calculated through the following formula:

[0181]

[0182] where, Represents the target semantic vector of the sample splicing statement;

[0183] Represents the normalized weight of the semantic feature vector of the i-th sample splicing statement.

[0184] Furthermore, the semantic feature vector of the sample spliced text and the target semantic feature vector of the sample spliced statement are spliced to obtain a comprehensive semantic feature vector, and the comprehensive semantic feature vector is sent to the first feedforward neural network layer.

[0185] S48. Feature transformation is performed on the comprehensive semantic feature vector through the first feedforward neural network layer to obtain a target comprehensive semantic feature vector, and the target comprehensive semantic feature vector is output through the output layer to obtain the predicted probability value of the sample web page data text belonging to the sample organization.

[0186] In specific implementation, after the server splices the semantic feature vector of the sample spliced text and the target semantic feature vector of the sample spliced statement to obtain a comprehensive semantic feature vector and sends it to the first feedforward neural network layer, feature transformation is performed on the comprehensive semantic feature vector through the first feedforward neural network layer to obtain a target comprehensive semantic feature vector. Furthermore, the target comprehensive semantic feature vector is output through the output layer to obtain the predicted probability value of the sample web page data text belonging to the sample organization.

[0187] Specifically, the following activation function can be used to obtain the predicted probability value of the sample web page data text belonging to the sample organization:

[0188]

[0189] Among them, represents the predicted probability value of the sample web page data text belonging to the sample organization;

[0190] represents the semantic feature vector of the sample spliced text and the target semantic feature vector of the sample spliced statement after direct splicing, the comprehensive semantic feature vector;

[0191] represents the target comprehensive semantic feature vector.

[0192] S49. Train the preset network model according to the deviation between the predicted probability value of the sample web page data text belonging to the sample organization and the actual probability value until the model converges to obtain the trained network asset recognition model.

[0193] In specific implementation, for each sample network data text, the preset network model is trained according to the deviation between the predicted probability value of the sample web page data text belonging to the sample organization and the actual probability, and the parameters of the preset network model are continuously adjusted until the model converges to obtain the trained network asset recognition model.

[0194] Specifically, the cross-entropy loss function can be used to calculate the loss, and the model is updated based on the gradient descent method.

[0195] The loss function is specifically as follows:

[0196]

[0197] Among them, L represents the loss function;

[0198] y represents the actual probability value (i.e., the true probability value) that the sample web page data text belongs to the sample organization. When the sample web page data text belongs to the sample organization, y takes the value of 1; when the sample web page data text does not belong to the sample organization, y takes the value of 0;

[0199] represents the predicted probability value that the sample web page data text belongs to the sample organization;

[0200] N represents the number of samples.

[0201] In the embodiment of the present application, when training the network asset recognition model, by extracting feature information from two granularities of sentences and texts, and combining the interaction between local features and global features to train the network asset recognition model. Compared with the prior art that only depends on the overall content of web page text data and ignores the mutual relationship between different granularity data, the network asset recognition model trained by the embodiment of the present application improves the accuracy and comprehensiveness of network asset recognition when identifying the network assets of the target organization, and overcomes the deficiency of traditional methods in multi-granularity information fusion.

[0202] In this step, for each web page data text, after obtaining the context sentence corresponding to the target second keyword of the web page data text, it can be determined whether the web page data text belongs to the target organization according to the process as Figure 7 shown, including the following steps:

[0203] S61. Input the target organization name information and the context sentences corresponding to each target second keyword into the text splicing processing layer.

[0204] S62. Through the text splicing processing layer, splice the target organization name information with the context sentences corresponding to all target second keywords to obtain a spliced text, and splice the target organization name information with the context sentence of each target second keyword respectively to obtain corresponding spliced sentences, and send the spliced text and each spliced sentence to the multi-granularity encoding and aggregation layer.

[0205] S63. Through the multi-granularity encoding and aggregation layer, extract the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced sentence, perform semantic aggregation based on the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced sentence to obtain the target semantic vector of the spliced sentence, and send the comprehensive semantic feature vector after splicing the semantic feature vector of the spliced text and the target semantic feature vector of the spliced sentence to the first feedforward neural network layer.

[0206] In specific implementation, it can be obtained according to the process as Figure 8 shown to obtain the target semantic vector of the splicing statement, including the following steps:

[0207] S71. Extract the semantic feature vector of the splicing text and the semantic feature vectors of each splicing statement through the multi-head attention network layer.

[0208] S72. Perform feature transformation on the semantic feature vector of the splicing text through the second feed-forward neural network layer to obtain the target feature vector of the splicing text.

[0209] S73. Perform dot product of the target feature vector of the splicing text with the semantic feature vectors of each splicing statement respectively to obtain the weights corresponding to the semantic feature vectors of the corresponding splicing statements.

[0210] S74. Normalize the weights corresponding to the semantic feature vectors of each splicing statement, and perform weighted average of the normalized weights of the semantic feature vectors of each splicing statement and the semantic feature vectors of each splicing statement to obtain the target semantic vector of the splicing statement.

[0211] In implementation, the implementation of steps S71 - S74 can refer to the implementation of steps S51 - S54, which will not be elaborated here.

[0212] S64. Perform feature transformation on the comprehensive semantic feature vector through the first feed-forward neural network layer to obtain the target comprehensive semantic feature vector, and output the target prediction probability value through the output layer for the target comprehensive semantic feature vector.

[0213] Among them, the target prediction probability value represents the probability that the predicted web page data text belongs to the target organization.

[0214] In implementation, the implementation of steps S61 - S64 can refer to the implementation of steps S45 - S48, which will not be elaborated here.

[0215] S65. Determine whether the web page data text belongs to the target organization according to the target prediction probability value.

[0216] In specific implementation, for each web page data text, if the server determines that the target prediction probability value is greater than the preset probability threshold, it determines that the web page data text belongs to the target organization, that is, the web page data text and the URL corresponding to the web page data text are network assets of the target organization. If it determines that the target prediction probability value is less than or equal to the preset probability threshold, it determines that the web page data text does not belong to the target organization, that is, the web page data text and the URL corresponding to the web page data text are not network assets of the target organization.

[0217] In view of the problem of large interference from invalid information in the process of enterprise network asset identification, the redundant text filtering method provided by the embodiments of this application extracts keywords and expands keywords to extract keywords from parts such as titles and text in the content of enterprise websites. Combining with the edit distance sorting model, it filters out low-correlation keyword texts and only retains the context statements of high-correlation keywords for analysis, effectively reducing the interference of invalid information, significantly improving the data quality for network asset identification, and solving the problem of messy information and drowning of effective information in traditional methods.

[0218] In view of the problem that existing methods cannot accurately identify enterprise network assets at both the local and global levels, the network asset identification model proposed by the embodiments of this application uses the text after redundant information filtering, extracts features from two levels: sentence granularity and document granularity, and comprehensively analyzes the interaction information between local details and the overall structure, effectively enhancing the accuracy and comprehensiveness of network asset identification, making the extraction of enterprise network assets more accurate and comprehensive, and making up for the deficiencies of traditional methods in terms of accuracy and coverage.

[0219] When implemented, the network asset identification method provided by the embodiments of this application can be but is not limited to being applied to the following scenarios:

[0220] Through the network asset identification method, enterprises can comprehensively understand the status of their Internet assets. These assets are constantly changing with the expansion of business and the upgrade of technology, posing challenges to security management. The above network asset identification method enables enterprises to track these changes in real time, establish a complete and accurate network asset ledger, and ensure that all network assets exposed to the Internet are within the monitoring scope of the security team. This can timely discover and repair potential vulnerabilities, reduce the risk of being attacked, improve the overall network security protection ability, and effectively support the enterprise's rapid response ability in the face of security threats.

[0221] It can also help enterprises with compliance auditing and supervision:

[0222] Globally, network security regulations in many industries and regions have strict requirements for enterprise compliance, requiring enterprises to conduct regular audits and reports on their Internet assets. Through the network asset identification method provided by the embodiments of this application, enterprises can automate the asset inventory process and reduce errors and omissions in manual operations. These technologies can generate detailed audit reports, including the status of assets, historical change records, and related security risks. It can help enterprises meet regulatory requirements, reduce compliance risks, and ensure that the enterprise's network security practices comply with industry best standards. Through the identified website assets, timely discover non-compliance issues existing in the website, such as non-compliance of ICP filing, etc., ensure that the website operates within the scope of compliance, and reduce legal and financial risks that may be faced due to violations.

[0223] Based on the same inventive concept, an embodiment of the present application further provides a network asset identification device. Since the principle of solving problems by the above network asset identification device is similar to that of the above network asset identification method, the implementation of the above device can refer to the implementation of the method, and the repeated parts will not be described again.

[0224] As Figure 9 shown, it is a schematic structural diagram of a network asset identification device provided by an embodiment of the present application, and may include:

[0225] A receiving module 81, configured to receive a network asset identification request sent by a terminal, where the network asset identification request carries target organization name information;

[0226] An obtaining module 82, configured to construct a uniform resource locator (URL) based on domain names, network protocol IP addresses corresponding to the domain names, and preset port numbers in a preset domain name system (DNS) database, and obtain web data texts corresponding to each URL;

[0227] A keyword extraction module 83, configured to extract a first keyword from the target organization name information and a second keyword from each web data text;

[0228] A determination module 84, configured to, for each second keyword included in each web data text, determine a relevance score between the second keyword and all first keywords according to the edit distance between each first keyword and the second keyword, and the relevance score between the second keyword and all first keywords represents the degree of association between the second keyword and the target organization name information;

[0229] A context extraction module 85, configured to, for each web data text, extract context statements corresponding to target second keywords in the web data text whose relevance scores with all first keywords are greater than a preset score threshold;

[0230] An identification module 86, configured to determine whether the web data text belongs to the target organization based on the target organization name information, the context statements corresponding to the target second keywords, and a network asset identification model, where the network asset identification model is used to identify whether the web data text is a network asset of the target organization.

[0231] In an implementation manner, the device further includes:

[0232] An expansion module, configured to, after extracting the second keyword of each web data text, expand the second keyword of each web data text based on near-synonyms of the second keyword of each web data text.

[0233] In one embodiment, the obtaining module 82 is specifically configured to construct a first URL based on the domain names in the preset DNS database; construct a second URL based on the IP addresses corresponding to the domain names in the preset DNS database and a preset port number; and obtain the web data texts corresponding to each first URL and each second URL respectively.

[0234] In one embodiment, the keyword extraction module 83 is specifically configured to input the target organization name information into a large language model to obtain the first keywords in the target organization name information, and the large language model is used to extract the first keywords in the target organization name information.

[0235] In one embodiment, the keyword extraction module 83 is specifically configured to perform word segmentation on each web data text to obtain the words included in each web data text; determine the weights of the respective words included in each web data text, where the weight of a word represents the importance degree of the word in the web data text; and determine the words with weights greater than a preset weight threshold as the second keywords of the corresponding web data text.

[0236] In one embodiment, the keyword extraction module 83 is specifically configured to, for each word included in each web data text, determine the term frequency-inverse document frequency corresponding to the word according to the number of times the word appears in the web data text, the number of words included in the web data text, the number of web data texts, and the number of web data texts including the word; and determine the term frequency-inverse document frequency corresponding to the word as the weight of the word.

[0237] In one embodiment, the determining module 84 is specifically configured to calculate the correlation score between the second keyword and all the first keywords through the following formula:

[0238]

[0239] where S i represents the correlation score between the i-th second keyword in the web data text and all the first keywords;

[0240] o i represents the i-th second keyword in the web data text;

[0241] c j represents the j-th first keyword in the target organization name information, j = 1, 2,..., m, and m represents the number of first keywords;

[0242] Levenshtein(c j , o i ) represents the j-th first keyword c in the target organization name informationj the edit distance between the i-th second keyword o in the web page data text i ;

[0243] L(o i ) represents the length of the i-th second keyword o in the web page data text i ;

[0244] (L(c j ) represents the length of the j-th first keyword c in the target organization name information j ;

[0245] In one implementation, the network asset identification model includes a text splicing processing layer, a multi-granularity encoding and aggregation layer, a first feed-forward neural network layer, and an output layer;

[0246] The recognition module 86 is specifically configured to input the target organization name information and the context statements corresponding to each target second keyword into the text splicing processing layer; splice the target organization name information with the context statements corresponding to all target second keywords through the text splicing processing layer to obtain a spliced text, and splice the target organization name information with the context statements of each target second keyword respectively to obtain corresponding spliced statements, and send the spliced text and each spliced statement to the multi-granularity encoding and aggregation layer; extract the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement through the multi-granularity encoding and aggregation layer, perform semantic aggregation based on the semantic feature vectors of the spliced text and the semantic feature vectors of each spliced statement to obtain the target semantic vector of the spliced statement, and send the comprehensive semantic feature vector after splicing the semantic feature vector of the spliced text and the target semantic feature vector of the spliced statement to the first feed-forward neural network layer; perform feature transformation on the comprehensive semantic feature vector through the first feed-forward neural network layer to obtain a target comprehensive semantic feature vector, output a target prediction probability value through the output layer, where the target prediction probability value represents the probability that the web page data text belongs to the target organization; determine whether the web page data text belongs to the target organization according to the target prediction probability value.

[0247] In one implementation, the multi-granularity encoding and aggregation layer includes a multi-head attention network layer and a second feed-forward neural network layer;

[0248] The recognition module 86 is specifically configured to extract the semantic feature vector of the spliced text and the semantic feature vectors of the respective spliced statements through the multi-head attention network layer; perform feature transformation on the semantic feature vector of the spliced text through the second feed-forward neural network layer to obtain the target feature vector of the spliced text; perform dot multiplication on the target feature vector of the spliced text and the semantic feature vectors of each spliced statement respectively to obtain the weights corresponding to the semantic feature vectors of the corresponding spliced statements; normalize the weights corresponding to the semantic feature vectors of each spliced statement, and perform weighted averaging on the normalized weights of the semantic feature vectors of each spliced statement and the semantic feature vectors of each spliced statement to obtain the target semantic vector of the spliced statement.

[0249] Based on the same inventive concept, an embodiment of the present application further provides an electronic device 900. Referring to Figure 10 as shown, the electronic device 900 is used to implement the network asset recognition method described in the above method embodiment. The electronic device 900 in this embodiment may include: a memory 901, a processor 902, and a computer program stored in the memory and executable on the processor, such as a network asset recognition program. When the processor executes the computer program, the steps in the above various network asset recognition method embodiments are implemented.

[0250] In the embodiment of the present application, the specific connection medium between the above-mentioned memory 901 and the processor 902 is not limited. In the embodiment of the present application Figure 10 it is connected by a bus 903 between the memory 901 and the processor 902. The bus 903 is represented by a thick line in Figure 10 For the connection manners of other components, only a schematic illustration is given and is not taken as a limitation. The bus 903 may be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 10 only a thick line is used to represent it in

[0251] The memory 901 may be a volatile memory, such as a random-access memory (RAM); the memory 901 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or the memory 901 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 901 may be a combination of the above memories.

[0252] A processor 902 is configured to implement the network asset identification method provided by the embodiments of the present application.

[0253] The embodiments of the present application further provide a computer-readable storage medium storing computer-executable instructions required to be executed by the above-mentioned processor, which includes a program for executing the operations required to be executed by the above-mentioned processor.

[0254] In some possible implementation manners, various aspects of the network asset identification method provided by the present application may also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the network asset identification method according to various exemplary embodiments of the present application described above in this specification.

[0255] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, an apparatus, or a computer program product. Therefore, the present application may be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0256] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatuses), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0257] These computer program instructions may also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0258] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or boxes Figure 1 one process or a plurality of processes and / or boxes Figure 1 steps for implementing the functions specified in one box or a plurality of boxes.

[0259] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0260] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A network asset identification method, characterized in that: include: Receiving a network asset identification request sent by a terminal, wherein the network asset identification request carries target organization name information; Constructing a uniform resource locator URL based on the domain name in a preset domain name system DNS database, the network protocol IP address corresponding to the domain name and the preset port number, and obtaining the web page data text corresponding to each URL; Extracting a first keyword from the target organization name information and extracting a second keyword from each web page data text; For each second keyword contained in each webpage data text, determine the relevance score of the second keyword with all the first keywords according to the edit distance between each first keyword and the second keyword, wherein the relevance score of the second keyword with all the first keywords represents the degree of association between the second keyword and the target organization name information; For each webpage data text, extracting context sentences corresponding to target second keywords in the webpage data text whose relevance scores to all first keywords are greater than a preset score threshold; Based on the target organization name information, the context sentence corresponding to the target second keyword and the network asset identification model, it is determined whether the web page data text belongs to the target organization. The network asset identification model is used to identify whether the web page data text is a network asset of the target organization.

2. The method according to claim 1, characterized in that After extracting the second keyword of each web page data text, it also includes: The second keyword of each web page data text is expanded based on the synonyms of the second keyword of each web page data text.

3. The method according to claim 1, characterized in that Construct a URL based on the domain name in the preset DNS database, the IP address corresponding to the domain name, and the preset port number, including: Constructing a first URL based on the domain name in the preset DNS database; Constructing a second URL based on the IP address corresponding to the domain name in the preset DNS database and the preset port number; Get the web page data text corresponding to each URL, including: The web page data text corresponding to each first URL and each second URL is obtained.

4. The method according to claim 1, characterized in that Extracting the first keyword in the target organization name information specifically includes: The target organization name information is input into a large language model to obtain a first keyword in the target organization name information, and the large language model is used to extract the first keyword in the target organization name information.

5. The method according to claim 1, characterized in that Extracting the second keyword of each web page data text includes: Performing word segmentation processing on each web page data text to obtain words contained in each web page data text; Determine the weight of each word contained in each web page data text, wherein the weight of the word represents the importance of the word in the web page data text; A word with a weight greater than a preset weight threshold is determined as a second keyword of the corresponding web page data text.

6. The method according to claim 5, characterized in that Determine the weight of each word contained in each web page data text, including: For each word contained in each web page data text, determine the word frequency inverse document frequency corresponding to the word according to the number of times the word appears in the web page data text, the number of words contained in the web page data text, the number of web page data texts and the number of web page data texts containing the word; The word frequency inverse document frequency corresponding to the word is determined as the weight of the word.

7. The method according to claim 1, characterized in that For each second keyword included in each webpage data text, the relevance score of the second keyword to all first keywords is determined according to the edit distance between each first keyword and the second keyword, specifically including: The relevance score of the second keyword and all first keywords is calculated by the following formula: Among them, S i represents the relevance score between the ith second keyword and all first keywords in the webpage data text; o i represents the i-th second keyword in the web page data text; c j represents the jth first keyword in the target organization name information, j=1, 2, ..., m, where m represents the number of first keywords; Levenshtein (c j ,o i ) represents the jth first keyword c in the target organization name information j and the second keyword o in the web page data text i The edit distance between L(o i ) represents the i-th second keyword o in the web page data text i Length; (L(c j ) represents the jth first keyword c in the target organization name information j Length.

8. The method according to any one of claims 1 to 7, characterized in that: The network asset identification model includes a text splicing processing layer, a multi-granularity encoding and aggregation layer, a first feedforward neural network layer and an output layer; Determining whether the webpage data text belongs to the target organization based on the target organization name information, the context sentence corresponding to the target second keyword, and the network asset identification model specifically includes: Inputting the target organization name information and the context sentences corresponding to each target second keyword into the text splicing processing layer; The target organization name information is concatenated with the context sentences corresponding to all the target second keywords through the text concatenation processing layer to obtain a concatenated text, and the target organization name information is concatenated with the context sentences of each target second keyword to obtain corresponding concatenated sentences, and the concatenated text and each concatenated sentence are sent to the multi-granularity encoding and aggregation layer; Extracting the semantic feature vector of the concatenated text and the semantic feature vectors of each concatenated sentence through the multi-granularity encoding and aggregation layer, performing semantic aggregation based on the semantic feature vector of the concatenated text and the semantic feature vectors of each concatenated sentence to obtain a target semantic vector of the concatenated sentence, and sending the comprehensive semantic feature vector after concatenating the semantic feature vector of the concatenated text and the target semantic feature vector of the concatenated sentence to the first feedforward neural network layer; Performing feature transformation on the comprehensive semantic feature vector through the first feedforward neural network layer to obtain a target comprehensive semantic feature vector, and outputting a target prediction probability value of the target comprehensive semantic feature vector through the output layer, wherein the target prediction probability value represents the probability that the webpage data text belongs to the target organization; Determine whether the web page data text belongs to the target organization according to the target prediction probability value.

9. The method according to claim 8, characterized in that The multi-granularity encoding and aggregation layer includes a multi-head attention network layer and a second feed-forward neural network layer; Extracting the semantic feature vector of the concatenated text and the semantic feature vectors of each concatenated sentence through the multi-granularity encoding and aggregation layer, performing semantic aggregation based on the semantic feature vector of the concatenated text and the semantic feature vectors of each concatenated sentence, and obtaining the target semantic vector of the concatenated sentence, specifically includes: Extracting the semantic feature vector of the concatenated text and the semantic feature vectors of each concatenated sentence through the multi-head attention network layer; Performing feature transformation on the semantic feature vector of the concatenated text by the second feedforward neural network layer to obtain a target feature vector of the concatenated text; Performing a dot multiplication of the target feature vector of the concatenated text and the semantic feature vector of each concatenated sentence respectively, to obtain a weight corresponding to the semantic feature vector of the corresponding concatenated sentence; The weight corresponding to the semantic feature vector of each concatenated sentence is normalized, and the normalized weight of the semantic feature vector of each concatenated sentence and the semantic feature vector of each concatenated sentence are weighted averaged to obtain the target semantic vector of the concatenated sentence.

10. A network asset identification device, characterized in that: include: A receiving module, configured to receive a network asset identification request sent by a terminal, wherein the network asset identification request carries target organization name information; The acquisition module is used to construct a uniform resource locator URL based on the domain name in the preset domain name system DNS database, the network protocol IP address corresponding to the domain name and the preset port number, and obtain the web page data text corresponding to each URL; A keyword extraction module, used to extract the first keyword in the target organization name information and the second keyword in each web page data text; a determination module, for determining, for each second keyword contained in each webpage data text, a relevance score between the second keyword and all first keywords according to the edit distance between each first keyword and the second keyword, wherein the relevance score between the second keyword and all first keywords represents the degree of association between the second keyword and the target organization name information; A context extraction module, for extracting, for each web page data text, context sentences corresponding to target second keywords in the web page data text whose relevance scores to all first keywords are greater than a preset score threshold; An identification module is used to determine whether the web page data text belongs to the target organization based on the target organization name information, the context sentence corresponding to the target second keyword and the network asset identification model, and the network asset identification model is used to identify whether the web page data text is a network asset of the target organization.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the network asset identification method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the network asset identification method according to any one of claims 1 to 9 are implemented.