A network asset identification method and device, electronic equipment and storage medium

By dynamically adjusting feature weights through a multimodal feature fusion model, the problem of insufficient accuracy in traditional network asset ownership determination methods is solved, and a high-accuracy determination of network asset organization ownership is achieved.

CN122513291APending Publication Date: 2026-08-04CHINA ELECTRONICS CORP 6TH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA ELECTRONICS CORP 6TH RES INST
Filing Date
2026-04-28
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Traditional methods for determining the ownership of online assets rely on static matching from a single data source, resulting in insufficient accuracy and an inability to effectively handle situations with missing or contradictory data. Furthermore, they are ill-equipped to cope with the dynamic changes and complex relationships in the ownership of online assets.

Method used

By combining multi-source feature information and using a multimodal feature fusion model to dynamically determine the feature weights of each analysis dimension, and combining single-dimensional confidence and feature weights, the organizational affiliation of the target network assets is finally determined.

Benefits of technology

It improves the accuracy of determining the ownership of network assets, can handle data gaps and dynamic changes, and adapts to complex relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513291A_ABST
    Figure CN122513291A_ABST
Patent Text Reader

Abstract

The application provides a network asset identification method and device, electronic equipment and storage medium. Feature information of a target network asset about each analysis dimension is extracted, and a single-dimensional confidence that the target network asset belongs to a current judgment organization under the analysis dimension is determined. The feature information of each analysis dimension is encoded as a feature representation, and the single-dimensional confidence is input into a multi-modal feature fusion model to determine a first feature weight of the feature representation of each analysis dimension and a second feature weight of the single-dimensional confidence, and then determine a fusion confidence that the target network asset belongs to the current judgment organization to determine the organizational attribution of the target network asset and obtain a network asset attribution result. In this way, by combining multi-source feature information and dynamically determining the feature weight of each analysis dimension by the multi-modal feature fusion model, the organizational attribution of the target network asset is finally determined, which can improve the judgment accuracy of the organizational attribution of the network asset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cyberspace security technology, and in particular to a method, apparatus, electronic device, and storage medium for identifying network assets. Background Technology

[0002] With the rapid development of the internet and the deepening of digital transformation, the number of cyber assets has exploded. Accurate identification of their social attributes (such as the organizations to which they belong) has become a key challenge for cybersecurity, digital governance, and asset management. Traditional methods for determining cyber asset ownership mainly rely on static matching from a single data source, resulting in insufficient accuracy, an inability to effectively handle missing or contradictory data, and a greater difficulty in addressing the dynamic changes and complex relationships in cyber asset ownership. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method, device, electronic device and storage medium for identifying network assets, which can combine multi-source feature information and dynamically determine the feature weight of each analysis dimension by a multimodal feature fusion model, and finally combine the single-dimensional confidence and feature weight to obtain the fusion confidence, thereby determining the organizational affiliation of the target network asset, thereby improving the accuracy of determining the organizational affiliation of network assets.

[0004] This application provides a method for identifying network assets, the method comprising: Based on the basic data of the target network assets, feature information of the target network assets in each of the multiple analysis dimensions is extracted; wherein, the multiple analysis dimensions include at least two of the following: DNS historical change feature dimension, WHOIS feature dimension, website content feature dimension, certificate feature dimension, and IP association feature dimension; For each analysis dimension, the single-dimensional confidence level of the target network asset belonging to the currently judged organization is determined based on the feature information of that analysis dimension; The feature information of each analysis dimension is encoded into a corresponding feature representation, and the feature representation of each analysis dimension and the single-dimensional confidence are input into the multimodal feature fusion model to determine the first feature weight of the feature representation of each analysis dimension and the second feature weight of the single-dimensional confidence of each analysis dimension. Based on the feature information and corresponding first feature weight of each analysis dimension, as well as the single-dimensional confidence and corresponding second feature weight, the fusion confidence of the target network asset belonging to the currently determined organization is determined. Based on the fusion confidence level of the target network asset belonging to the currently determined organization, the organizational affiliation of the target network asset is determined, and the network asset affiliation result is obtained.

[0005] This application embodiment also provides a network asset identification device, the identification device comprising: The extraction module is used to extract feature information of the target network asset for each of multiple analysis dimensions based on the basic data of the target network asset; wherein, the multiple analysis dimensions include at least two of the following: DNS historical change feature dimension, WHOIS feature dimension, website content feature dimension, certificate feature dimension, and IP association feature dimension; The first confidence level determination module is used to determine the single-dimensional confidence level of the target network asset belonging to the currently judged organization under each analysis dimension based on the feature information of that analysis dimension. The weight determination module is used to encode the feature information of each analysis dimension into a corresponding feature representation, and input the feature representation of each analysis dimension and the single-dimensional confidence into the multimodal feature fusion model to determine the first feature weight of the feature representation of each analysis dimension and the second feature weight of the single-dimensional confidence of each analysis dimension. The second confidence determination module is used to determine the fusion confidence that the target network asset belongs to the currently determined organization based on the feature information of each analysis dimension and the corresponding first feature weight, as well as the single-dimensional confidence and the corresponding second feature weight. The determination module is used to determine the organizational affiliation of the target network asset based on the fusion confidence level of the target network asset belonging to the currently determined organization, and to obtain the network asset affiliation result.

[0006] This application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the network asset identification method described above are performed.

[0007] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the network asset identification method described above.

[0008] This application provides a method, apparatus, electronic device, and storage medium for identifying network assets. The method extracts feature information of the target network asset across multiple analytical dimensions and determines the single-dimensional confidence level of the target network asset belonging to the currently determined organization under each analytical dimension. A multimodal feature fusion model determines feature weights based on the feature information of each analytical dimension, and finally combines the single-dimensional confidence level and feature weights to obtain a fused confidence level, thereby determining the organizational affiliation of the target network asset.

[0009] In this way, multi-source feature information can be fused during the identification of network assets, and the multi-modal feature fusion model can dynamically adjust the feature weights based on the quality of the feature information. Therefore, compared with static matching from a single data source or multi-source fixed-weight methods, this approach can improve the accuracy of determining the organizational affiliation of network assets.

[0010] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a method for identifying network assets provided in an embodiment of this application is shown; Figure 2 A schematic diagram of the structure of a network asset identification device provided in an embodiment of this application is shown; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.

[0014] Research has revealed that with the rapid development of the internet and the deepening of digital transformation, the number of network assets has exploded. Accurate identification of their social attributes (such as the organization to which they belong) has become a key challenge for cybersecurity, digital governance, and asset management. Traditional methods for determining network asset ownership primarily rely on static matching from a single data source, leading to insufficient accuracy. For example, existing technologies often depend on single data sources such as WHOIS, IP geolocation, or SSL certificates. When WHOIS registrant information is outdated, SSL certificates are not updated in a timely manner, or IP addresses are shared by multiple organizations, the determination results will deviate significantly from reality, thus failing to effectively handle situations with missing or contradictory data. Furthermore, traditional methods for determining network asset ownership employ static feature analysis, making it even more difficult to address the dynamic changes and complex relationships within network asset ownership, further contributing to insufficient accuracy.

[0015] Based on this, embodiments of this application provide a method for identifying network assets to improve the accuracy of determining the organizational ownership of network assets.

[0016] Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for identifying network assets provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the identification method includes: S101. Based on the basic data of the target network asset, extract the feature information of the target network asset for each of the multiple analysis dimensions.

[0017] In this embodiment, network assets refer to information assets that have at least an IP address. Information assets include, but are not limited to, servers, PCs, network devices (such as switches and routers), and security devices (such as firewalls and IDS).

[0018] The multiple analysis dimensions include at least two of the following: DNS historical change feature dimension, WHOIS feature dimension, website content feature dimension, certificate feature dimension, and IP association feature dimension.

[0019] For the DNS historical change feature dimension, the historical resolution records (A record, MX record, NS record, etc.) of the domain name can be queried through the DNS historical database (such as DNSDB, CIRT.NET API) to obtain the domain name's ownership change trajectory. The feature information may include at least one of the following: (1) Domain name historical time series: records the resolution results of the domain name at different time points; (2) IP address change records pointed to by the domain name: records the history of IP address changes pointed to by the domain name; (3) Server geographical location change records resolved by the domain name: records the different geographical locations resolved by the domain name.

[0020] For WHOIS feature dimensions, the domain name registration information can be obtained through the WHOIS query interface. The feature information may include at least one of the following: (1) Registrant name: extract the name of the registered organization in the WHOIS record; (2) Registration email: extract the registration contact email in the WHOIS record; (3) Contact phone number: extract the contact information in the WHOIS record; (4) Registration address: extract the geographical location information in the WHOIS record.

[0021] For website content feature dimensions, website content can be obtained through non-intrusive web crawlers, and website content with organizational identification features can be extracted. The feature information can specifically include at least one of the following: (1) Text semantic features: extract the semantic vector of the web page text using the BERT model; (2) Keyword features: extract organizational identification keywords such as "About Us" and "Company Profile"; (3) Copyright statement: extract the copyright statement information at the bottom of the web page; (4) Contact information: extract contact address, telephone number and other information in the web page.

[0022] For the certificate feature dimension, the organization association feature can be extracted through TLS certificate parsing. The feature information may include at least one of the following: (1) Certificate subject name: extract the "O" (Organization) field in the certificate; (2) Certificate Authority: extract the certificate issuer information; (3) Certificate fingerprint: extract the SHA-256 fingerprint value of the certificate; (4) Certificate chain information: extract the complete certificate chain information.

[0023] For IP association feature dimension, network layer association features can be extracted through IP detection. The feature information may include at least one of the following: (1) IP geographical location: parsing the geographical location information of the IP address; it should be noted that this feature information is not applicable to cloud service asset tracing and cannot be used as the basis for cloud asset judgment. It is only used for enterprise distribution matching of non-cloud assets; (2) Automatic Numbering System (AS) number: parsing the AS number information to which the IP belongs; (3) Network topology: analyzing the network topology structure of the IP address; (4) Cloud service identifier: identifying whether the IP address belongs to the cloud service provider.

[0024] In practical implementation, each of the above analytical dimensions has its core characteristics (advantages) and different applicable scenarios. For the DNS historical change feature dimension, its core advantage lies in its ability to capture dynamic changes in ownership, making it particularly suitable for identifying domain migrations and organizational mergers. For the WHOIS feature dimension, as official registration information, it is especially suitable for identifying new domains and legitimately registered assets. For the website content feature dimension, containing semantic-level organizational identifiers, it is particularly suitable for identifying content-rich official websites. For the certificate feature dimension, as encrypted identity authentication, it is particularly suitable for identifying HTTPS sites and certificate reuse associations. For the IP association feature dimension, reflecting network layer ownership, it is particularly suitable for cloud assets and shared IP scenarios. In practical applications, multiple analytical dimensions can be selected as needed.

[0025] S102. For each analysis dimension, determine the single-dimensional confidence level of the target network asset belonging to the currently determined organization based on the feature information of that analysis dimension.

[0026] In this step, the current organization to which the target network asset belongs can be preliminarily determined based on the basic data of the target network asset. The current organization is determined based on the preliminary results of data from various analytical dimensions, such as domain names in DNS records, information in WHOIS registration information, organization names in website content, and IP addresses. These can all be used to preliminarily determine the organization and obtain the current organization. Then, based on the characteristic information of each analytical dimension, the single-dimensional confidence level of the target network asset belonging to the current organization under that analytical dimension is determined.

[0027] In a first possible implementation, DNS historical change records are a key clue for identifying dynamic changes in asset ownership, particularly regarding DNS historical change characteristics. However, existing technologies lack the ability to analyze ownership history and cannot identify dynamic changes in asset ownership. For example, the domain "example.com" pointed to IP_X (belonging to organization A0) in 2023, but changed to IP_Y (belonging to organization B0) in 2024. Analyzing only the current DNS records would not accurately determine the historical ownership trajectory of this domain. Therefore, this application provides a DNS historical ownership trajectory analysis technique that evaluates the reliability of DNS historical records from three aspects.

[0028] For the DNS historical change feature dimension, the feature information includes the domain name historical time series, the IP address change record pointed to by the domain name, and the server geographical location change record resolved by the domain name; then step S102 may include: S1021a. Based on the historical time series of the domain name, determine the frequency of changes in the domain name resolution record and obtain the time continuity evaluation index value.

[0029] For time continuity assessment, the frequency of changes in domain name resolution records can be analyzed; time continuity assessment index values The calculation formula is:

[0030] in, This represents the number of times the domain name resolution records have changed within the last n years (e.g., 3 years). When ≤1 time / year, The value is relatively high.

[0031] S1022 a. Based on the IP address change records pointed to by the domain name, determine the matching degree between the organization to which the IP pointed to by the domain name belongs and the currently determined organization, and obtain the IP attribution consistency evaluation index value.

[0032] For IP attribution consistency assessment, the matching degree between the historical IP address ownership and the currently identified organization can be analyzed; IP attribution consistency assessment index values. The calculation formula is:

[0033] in, This represents the text similarity calculated by the BERT model. The organization to which the current IP address belongs, i.e., the organization currently being determined. This refers to the organization to which the historical IP address belonged.

[0034] S1023a. Based on the server geographical location change records resolved by the domain name, determine the rationality of the domain name change and obtain the rationality evaluation index value of the change.

[0035] The value of the change rationality assessment index can be assigned based on the analysis of the business context. For example, if the business context includes "the domain name migration was caused by an organizational merger in 2023", the change rationality assessment index value will be... Relatively high.

[0036] S1024 a. Determine the single-dimensional confidence level corresponding to the DNS historical change feature dimension based on the time continuity assessment index value, the IP home consistency assessment index value, and the change rationality assessment index value.

[0037] Based on the above three aspects, this application's embodiments design a DNS historical confidence evaluation model:

[0038] For example, , , ; This represents the single-dimensional confidence level that the target network asset belongs to the currently identified organization under the DNS historical change feature dimension. The range is 0~1. Generally speaking, when When the value is greater than or equal to a preset threshold (e.g., 0.8), it can be determined that the DNS historical characteristics are highly correlated with the current organization being judged.

[0039] In a second possible implementation, regarding the website content feature dimension, website content is a crucial basis for identifying the organization to which an asset belongs. However, traditional methods rely solely on text matching, lacking in-depth analysis of content semantics and organizational relevance. This application provides a content-organization association analysis technique that evaluates the correlation between content and organization from three aspects.

[0040] For the website content feature dimension, the feature information includes website content; then step S102 may include: S1021b. Determine the text semantic matching index value based on the semantic similarity between the organization identifier in the website content and the organization information of the currently determined organization in the organization name database.

[0041] For the text semantic matching index, the BERT model can be used to calculate the semantic similarity between the webpage text and the organizational information of the currently identified organization in the organization name database. The formula for calculating the text semantic matching index is as follows:

[0042] in, This represents the cosine similarity calculated by the BERT model. Organize identifiers in web page text. The organization name is a name in the organization name database.

[0043] S1022b. Determine the content pattern matching index value based on the matching degree between the website structure in the website content and the organizational characteristics of the currently determined organization.

[0044] The content pattern matching metric can be used to analyze the degree of matching between the website structure (such as navigation bar, footer, contact information) and the organization's business characteristics. The formula for calculating the content pattern matching metric is:

[0045] in, The number of structural features that match the currently identified organization. This represents the total number of structural features analyzed.

[0046] S1023b. Determine the multi-source content cross-validation index value based on the consistency between the website content and the organizational information of the target network asset in at least one other information source. The organizational information of the target network asset in at least one other information source includes the registrant information in WHOIS and / or the SSL certificate organization name.

[0047] For multi-source content cross-validation metrics, the consistency between website content and SSL certificate organization name, WHOIS registrant information, etc., can be compared. The calculation formula is:

[0048] in, The number of data sources with consistent information. This represents the total number of data sources used in the comparison.

[0049] S1024b. Determine the single-dimensional confidence level corresponding to the website content feature dimension based on the text semantic matching index value, the content pattern matching index value, and the multi-source content cross-validation index value.

[0050] Based on the above three aspects, this application's embodiments design a content association confidence assessment model:

[0051] For example, , , ; This represents the single-dimensional confidence level that the target online asset belongs to the currently identified organization, based on the website content feature dimension. The range is 0~1. Generally speaking, when When the value is greater than or equal to a preset threshold (e.g., 0.8), it can be determined that the website's content features are highly correlated with the currently identified organization.

[0052] In the third possible implementation, for the WHOIS feature dimension, WHOIS registration information serves as a direct clue to the ownership of network assets. Its core value lies in quantifying the registration correlation between assets and the currently determined organization through direct matching of registration entity information (registering entity, registrant, contact information, etc.) with the organization currently being judged. During the judgment process, priority is given to matching the full / abbreviated name of the registering entity, followed by verifying the correlation between the contact email domain name, contact phone number location, and the organization currently being judged. Invalid and false registration information (such as anonymous registration or third-party agent registration) is eliminated, and finally, the independent confidence score of the WHOIS feature dimension is output. The value ranges from [0,1]. The closer the value is to 1, the higher the degree of association between the asset and the registration of the currently determined organization.

[0053] The calculation method is as follows:

[0054] In the formula, For the weighting coefficients, satisfying ,in (Registered units have the highest matching weight) (Email domain matching weight) (Telephone attribution matching weight) (Invalid information penalty weights) The initial values ​​of the weight coefficients can be determined through sample training and subsequently optimized through a closed-loop feedback mechanism.

[0055] : Matching degree of the registered entity, with values ​​of {0, 0.5, 1}; 1 indicates that the registered entity is completely consistent with the full name of the organization being judged, 0.5 indicates that the abbreviation is consistent or there is a relationship (such as a subsidiary and a parent company), and 0 indicates that there is no matching relationship.

[0056] : Contact email domain matching score, with a value of {0, 1}; 1 indicates that the email domain is the official domain of the organization being judged (such as xxx@target.org), and 0 indicates that it is a non-official domain.

[0057] : Phone number location matching score, with values ​​of {0, 0.5, 1}; 1 indicates that the phone number location is exactly the same as the registered address of the organization being judged, 0.5 indicates that the location is the city where the organization is being judged, and 0 indicates no relevance.

[0058] Invalid information penalty factor, with values ​​of {0, 0.5, 1}; 1 indicates that the registration information is anonymous or false, 0.5 indicates that the registration information is incomplete, and 0 indicates that the registration information is true and valid.

[0059] In a fourth possible implementation, for the certificate feature dimension, this application embodiment uses certificate chain association features for single-dimensional confidence determination. In specific implementation, certificate chain association features utilize information such as the SSL / TLS certificate holder, associated domain name, and issuing authority to mine the association between the asset and the currently determined organization, which is particularly suitable for determining the ownership of assets with hidden associations. During the determination process, priority is given to matching the consistency between the certificate holder (e.g., company name, organization code) and the currently determined organization; secondly, the association between the certificate-associated domain name and the known asset domain names of the currently determined organization is verified, eliminating interference from third-party issued generic certificates, and outputting the independent confidence level of the certificate feature dimension (…). The value ranges from [0,1]. The closer the value is to 1, the higher the correlation between the asset and the certificate of the currently judged organization.

[0060] The calculation method is as follows:

[0061] in, For the weighting coefficients, satisfying ,in (Certificate holders have the highest matching weight). (Associated domain name matching weight) (Certificate validity weights) The initial weight values ​​are determined through sample training and subsequently optimized through closed-loop feedback.

[0062] : Certificate holder matching degree, with values ​​of {0, 0.5, 1}; 1 indicates that the certificate holder is completely consistent with the full name / organization code of the currently judged organization, 0.5 indicates that there is an association relationship, and 0 indicates that there is no matching relationship.

[0063] : Certificate-associated domain name matching degree, with a value of [0,1]; calculated by the similarity between the associated domain name and the known asset domain name of the currently judged organization. The higher the similarity, the closer the value is to 1. If there is no associated domain name, the value is 0.

[0064] Certificate validity factor, with values ​​of {0, 1}; 1 indicates that the certificate is valid and the issuing authority is legitimate, while 0 indicates that the certificate is expired, revoked, or the issuing authority is illegitimate.

[0065] In the fifth possible implementation, for the IP association feature dimension, the IP association feature quantifies the degree of association between the asset and the network topology of the currently determined organization through information such as IP address, IP segment, and IP operator.

[0066] It should be noted that IP geolocation analysis cannot be used as the basis for determining the ownership of cloud assets. Therefore, in the determination process, for non-cloud assets, the focus is on matching the IP segment ownership, the correlation between the IP's carrier and the organization being determined; while for cloud assets, IP association features are only used for basic network topology analysis and do not participate in the ownership confidence calculation (weights are automatically set to 0), ultimately outputting an independent confidence score at the IP dimension. The value range is [0,1].

[0067] The calculation method is as follows:

[0068] In the formula, For non-cloud asset scenarios, the weighting coefficients must meet the following requirements. ,in (IP segment attribution matching weight) (Operator matching weights) The initial weight values ​​are determined through sample training and subsequently optimized through closed-loop feedback; in cloud asset scenarios, all weights are set to 0.

[0069] : IP segment attribution matching degree, with values ​​of {0, 0.5, 1}; 1 indicates that the IP segment is a unique IP segment of the currently identified organization, 0.5 indicates that the IP segment overlaps with the associated IP segments of the currently identified organization, and 0 indicates that there is no association.

[0070] : IP operator matching degree, with a value of {0, 1}; 1 indicates that the IP operator is consistent with the operator cooperating with the organization currently being judged, and 0 indicates that they are inconsistent.

[0071] Among them, IP geolocation analysis is only used for enterprise distribution matching of non-cloud assets and does not participate in confidence calculation in any scenario; the determination of ownership of cloud assets relies only on IP segment ownership, certificate chain association, domain name registration information, etc., and can also be combined with cloud service asset traceability technology, which will be detailed below.

[0072] Furthermore, the determination of ownership of cloud service assets has unique characteristics. Existing technologies mainly rely on external detection features (such as domain name naming patterns) to determine the ownership of cloud service assets, failing to fully utilize the metadata provided by cloud vendors (such as resource group names and tag information). This makes it difficult to directly associate the assets with the actual organization, resulting in low accuracy and difficulty in adapting to multi-cloud hybrid environments. Therefore, this application also provides a cloud service asset traceability technology that can directly identify the organization to which the asset belongs by analyzing the metadata information provided by the cloud vendor.

[0073] When an IP address is identified as belonging to a cloud service provider, the characteristic information includes metadata information provided by the cloud service provider; therefore, the identification method provided in this application embodiment further includes determining the organization to which the cloud service asset belongs based on the metadata information in the following ways: Step 1: Extract the first organization name from the cloud resource group name and the second organization name from the cloud resource tag using regular expressions or keyword extraction algorithms.

[0074] In this step, resource group naming analysis can be performed using regular expressions to parse the organization identifier in the cloud resource group name, such as "XX-Department-Prod" → XX Company R&D Department, and the first organization name can be extracted using regular expressions.

[0075] Tag association analysis can also be performed using keyword extraction algorithms to extract organization identifiers from cloud resource tags, such as "Owner=XX-Team" and "Project=XX-System", and then identify the second organization name using keyword extraction algorithms.

[0076] Step 2: Cross-validate the validity of the first organization name and the second organization name based on the single-dimensional confidence of WHOIS feature dimension and the single-dimensional confidence of certificate feature dimension, respectively. If the validation passes, determine the matching degree between the first organization name and the name of the currently determined organization, and determine the similarity between the second organization name and the name of the currently determined organization.

[0077] In this step, the attribution of the first organization name is traced through the resource group information of the cloud service assets. The validity of the tag information can be based on the feature information of the WHOIS feature dimension and certificate feature dimension extracted in the multimodal feature information in S101 above, while also referring to the single-dimensional confidence of the WHOIS feature dimension. ) and certificate chain association confidence ( Cross-validation is performed. For example, if the cloud resource group content matches the domain registration information, and... ≥Preset threshold (e.g., 0.7) or If the value is ≥ a preset threshold (e.g., 0.6), it indicates that the resource group information is valid. This helps to confirm the correlation between the resource group name and the currently identified organization, improving the accuracy of traceability.

[0078] For the attribution and tracing of the second organization name through the tag information (such as organization identifier, business tag) of cloud service assets, the validity of the tag information can be based on the feature information of WHOIS feature dimension and certificate feature dimension in the multimodal feature information extraction in the aforementioned S101, while also referring to the single-dimensional confidence of WHOIS feature dimension ( ) and certificate chain association confidence ( Cross-validation is performed. For example, if the label information matches the WHOIS registration entity and certificate holder, and ≥Preset threshold (e.g., 0.5) or If the value is ≥ a preset threshold (e.g., 0.5), the label information is considered valid. This helps confirm the correlation between the label information and the currently identified organization, improving the accuracy of traceability.

[0079] Next, the name of the first organization is matched with the name of the currently determined organization, and the matching degree is recorded as . (Range 0~1), the calculation formula is:

[0080] in, For regular expression matching similarity, For cloud resource group name, This is the name of the organization currently being identified.

[0081] Similarly, the similarity between the name of the second organization and the name of the currently identified organization is calculated and denoted as... (Range 0~1).

[0082] Step 3: If the matching degree is greater than the preset matching degree threshold and / or the similarity degree is greater than the preset similarity threshold, then the currently determined organization is identified as the target organization to which the cloud service asset belongs.

[0083] In this step, if the matching degree is greater than the preset matching degree threshold and / or the similarity degree is greater than the preset similarity threshold, the current identified organization can be directly identified as the target organization to which the cloud service asset belongs, thus obtaining the network asset ownership result.

[0084] S103. Encode the feature information of each analysis dimension into a corresponding feature representation, and input the feature representation of each analysis dimension and the single-dimensional confidence into the multimodal feature fusion model to determine the feature weight of each analysis dimension.

[0085] In this step, the feature information of each analysis dimension is first normalized and encoded into a unified feature representation (feature vector), with values ​​ranging from [0,1]. DNS historical change record feature vector: WHOIS registration information feature vector: Website content semantic feature vector: Certificate chain associated feature vector: IP-related feature vector: .

[0086] This application proposes an improved multimodal feature fusion model that combines the Grey Wolf Optimization (GWO) algorithm with an attention mechanism. As the core fusion module of the entire determination method, this model can receive all the valid data from the aforementioned steps. By combining the improved Grey Wolf Optimization (GWO) algorithm with the attention mechanism, it dynamically optimizes feature weights, resolves conflicts between multiple data sources, and achieves accurate determination of network asset ownership.

[0087] The model's input data includes two core categories. The first category is multimodal raw feature vectors, which, depending on the analysis dimension used, include at least two of the following: DNS historical change record feature vectors, WHOIS registration information feature vectors, website content semantic feature vectors, certificate chain association feature vectors, and IP association feature vectors. Among these, the geographic location feature in the IP association feature vector only participates in the calculation in non-cloud asset scenarios; in cloud asset scenarios, this feature component is automatically set to 0, strictly adhering to the IP geographic location usage rules. The second category is single-dimensional confidence scores for each analysis dimension, which, depending on the analysis dimension used, include confidence scores for DNS historical change records (…). ), website content semantic confidence ( ), WHOIS registration information confidence level ( ), Certificate chain association confidence ( IP association confidence ( ); among them, cloud asset scenarios =0, in non-cloud asset scenarios Calculate according to the aforementioned formula.

[0088] In this embodiment of the application, the model adopts a dual weight optimization strategy of initial GWO optimization + secondary adjustment by attention mechanism. Then step S103 may include: S1031. Configure the feature representation and initial feature weights corresponding to the confidence level of each analysis dimension.

[0089] In this step, initial feature weights can be configured for the feature representation and single-dimensional confidence of each analysis dimension based on the sample training of the multimodal feature fusion model. Among them, for the same analysis dimension, the initial feature weight of the single-dimensional confidence should be higher than the initial feature weight of the feature representation, so as to give priority to reflecting the value of the independent quantization results in S102.

[0090] Alternatively, the initial feature weights for each feature representation can be calculated using a self-attention mechanism, as shown in the formula:

[0091] in, Both are generated by linear transformation of the two types of data input to the model; multimodal feature vectors ( , , , , Using as the basic input, it is generated through a linear mapping. (Query vector) (Key vector); Single-dimensional confidence is used as an auxiliary input, participating in... Generation of (value vectors).

[0092] S1032. Using the improved Grey Wolf optimization algorithm, the initial feature weights are initially iteratively optimized. During the optimization process, the weight allocation of the conflict analysis dimensions is adjusted to obtain the feature representation of each analysis dimension and the first optimized weight of the single-dimensional confidence.

[0093] In this step, the optimization objectives are to maximize the accuracy of attribution determination and minimize the conflict of multi-source data. By improving the Grey Wolf optimization algorithm, the initial feature weights are initially iteratively optimized, the conflicting analysis dimensions (such as the large deviation between Content_conf and WHOIS_conf) are identified and the weight allocation is adjusted to ensure the reasonable use of the features and confidence of each analysis dimension, and the first optimized weight is obtained.

[0094] S1033. Based on the feature representation and single-dimensional confidence of each analysis dimension, the correlation score between each feature representation, each single-dimensional confidence and the currently determined organization is determined through an attention mechanism, and the weights are adjusted based on the correlation score to obtain the second optimized weights of the feature representation and single-dimensional confidence of each analysis dimension.

[0095] In this step, the correlation score between each feature representation and the current judgment organization can be determined based on the self-attention mechanism, as well as the correlation score between each single-dimensional confidence score and the current judgment organization, to ensure that the correlation score fits the aforementioned feature and confidence judgment results; the specific calculation formula can refer to existing technologies. Then, the weight ratio is increased for inputs with high correlation scores (e.g., WHOIS_conf≥0.8, Cert_conf≥0.8), and the weight ratio is decreased for inputs with low correlation scores (e.g., IP_conf=0, Content_conf<0.3), further improving the fusion accuracy.

[0096] S1034. Perform cross-validation on the feature representations of each analysis dimension and the first and second optimized weights of the single-dimensional confidence.

[0097] S1035. If the verification is successful, then based on the feature representation of each analysis dimension and the first and second optimized weights of the single-dimensional confidence, determine the first feature weight of the feature representation of each analysis dimension and the second feature weight of the single-dimensional confidence of each analysis dimension.

[0098] S1036. If the verification fails, the feature weights are iteratively optimized again using the improved Grey Wolf optimization algorithm to obtain the updated feature representation of each analysis dimension and the first optimized weight of the single-dimensional confidence, and cross-validation is performed again.

[0099] For steps S1034, S1035, and S1036, the second feature weight result after the second adjustment of the attention mechanism will be fed back to the GWO optimization stage and cross-validated with the first optimized weight output by the GWO algorithm.

[0100] If the weight deviation between the two for the same object is ≤0.1, the first optimized weight output by the Grey Wolf optimization algorithm, the second feature weight obtained after adjustment by the attention mechanism, or the average of the two can be used to obtain the final feature weight (i.e., for the feature representation of each analysis dimension, the corresponding first feature weight is obtained; for the single-dimensional confidence of each analysis dimension, the corresponding second feature weight is obtained). If the deviation is >0.1, the GWO algorithm is used for iterative optimization again, and the cross-validation in step S1034 is returned to implement the collaborative optimization logic of GWO preliminary optimization → secondary attention adjustment → GWO verification calibration, thereby improving the accuracy of weight allocation.

[0101] In one possible implementation, for step S1032 above, the initial feature weights are used as the initial solution of the improved gray wolf optimization algorithm. The initial population is generated by using the Tent chaotic mapping, and iterative optimization is performed to obtain the first optimization weights corresponding to each feature representation and each single-dimensional confidence under each analysis dimension. In each iteration, the dynamic parameters are determined according to the current iteration number, and the position is updated according to the dynamic parameters and the individual memory coefficient.

[0102] This application's embodiments improve upon the traditional GWO algorithm in three aspects, significantly enhancing the accuracy and efficiency of feature weight optimization; specifically including: Chaotic initialization strategy: An improved Tent chaotic map is used to generate the initial population, ensuring a uniform distribution of the population in the search space.

[0103] in, For chaotic parameters, This is the current iteration value.

[0104] Dynamic parameter a adjustment strategy: Design dynamic parameters that change over time. The global exploration and local development capabilities of the balancing algorithm:

[0105] in, This represents the current iteration number. This represents the maximum number of iterations. (Parameter) The algorithm decreases linearly from the initial value of 2 to 0, thus transforming the algorithm from global exploration to local development.

[0106] Individual memory coefficient κ: Introducing the individual memory coefficient To prevent the algorithm from converging to a local optimum too early:

[0107] in, For the individual's historical best position, This is the current optimal solution position. , , It is a random number. This is the individual memory coefficient, with a value range of [value range missing]. .

[0108] S104. Based on the feature representation and corresponding first feature weight of each analysis dimension, as well as the single-dimensional confidence and corresponding second feature weight, determine the fusion confidence of the target network asset belonging to the currently determined organization.

[0109] In this step, a multimodal confidence fusion algorithm is used to fuse and calculate the various input data corresponding to the optimized feature weights, and outputs the final confidence score (Final_conf) of the target network asset belonging to the current determining organization, with a value range of [0,1]. The specific calculation formula is as follows:

[0110] In the formula, The final confidence level of the target network asset's ownership by the current determining organization, with a value range of [0,1], serves as the core basis for ownership determination and subsequent automatic clue expansion and closed-loop feedback.

[0111] The feature representation of each analysis dimension corresponds to the first feature weight (i=1 to 5, corresponding to the aforementioned five analysis dimensions), satisfying... (C is a fixed coefficient, with an example value of 0.4, which forms a reasonable allocation with the confidence weight).

[0112] : The value of the feature representation obtained after normalization of the original feature vectors of each analysis dimension (value range [0,1]).

[0113] The second feature weights corresponding to the single-dimensional confidence scores of each analysis dimension (j=1 to 5, corresponding to DNS_conf, Content_conf, WHOIS_conf, Cert_conf, IP_conf) satisfy the following conditions: (A value of 0.6 is recommended).

[0114] : The confidence score of each analysis dimension (range [0,1]).

[0115] S105. Based on the fusion confidence level of the target network asset belonging to the currently determined organization, the organization affiliation of the target network asset is determined to obtain the network asset affiliation result.

[0116] In specific implementation, step S105 may include: When the fusion confidence level is greater than or equal to the first preset threshold, the network asset ownership result of the target network asset is determined to be the currently determined organization; when the fusion confidence level is less than the first preset threshold, it is submitted for manual review, and the manual review result is determined as the network asset ownership result of the target network asset.

[0117] In one example, when When the value is ≥0.8, it is considered an asset of the organization; when 0.6 ≤ When the value is less than 0.8, it is marked as "Pending Confirmation," triggering manual review. Furthermore, when... If the value is less than 0.6, it can still be judged as "cannot be determined", the original data can be retained, and manual review can still be triggered.

[0118] Furthermore, this application also proposes an automatic clue expansion technology based on the judgment result, which continuously discovers hidden related assets through a closed-loop feedback mechanism. The automatic clue expansion technology uses the final confidence score (Final_conf) output by the multimodal feature fusion model as the trigger condition, and relies on the multimodal feature information (representation) extracted from the front end and the single-dimensional confidence score judgment result to discover other network assets that have hidden connections with the current target network asset.

[0119] Here, when the fusion confidence level is greater than a preset confidence threshold, other related network assets associated with the target network asset are mined from different analysis dimensions based on the single-dimensional confidence level and feature information under each analysis dimension.

[0120] In practical implementation, regarding the triggering conditions, when When the reliability threshold is ≥ 0.7, the current asset is determined to be a definite related asset of the currently identified organization, and the lead expansion process is automatically triggered; when 0.3 < When the value is less than 0.7, it is considered a suspicious asset, and expansion will not be triggered temporarily. It can be included in the closed-loop feedback mechanism for subsequent verification; when... When the confidence level is ≤ 0.3, the asset is determined to be outside the currently identified organization and no expansion is triggered. The single-dimensional confidence level determines whether to perform related asset mining and its priority.

[0121] Regarding the specific basis and process for expanding clues, for certificate feature dimensions: based on the extracted certificate-related domain name, certificate holder information, and... Based on the judgment results, other domain names, servers, and other assets using the same certificate as the target network asset will be identified; among them, if If the value is ≥ 0.7, then priority will be given to expanding this type of related assets.

[0122] For the DNS historical change feature dimension: based on the extracted historical DNS resolution IPs, associated domain name information, and... Based on the judgment results, other assets that have DNS resolution associations with the target network assets are identified; among them, if If the value is ≥ 0.6, then priority should be given to expanding this type of asset.

[0123] Regarding IP association feature dimensions: In non-cloud asset scenarios, based on the extracted IP segments, IP-associated domain name information, and... Based on the judgment results, other related assets within the same IP segment are explored, without using IP geographic location information for expansion; among them, if If the value is ≥ 0.6, then this type of asset will be prioritized for expansion. In the cloud asset scenario, only the IP segment ownership and cloud service provider metadata are used to expand related assets within the same cloud environment; similarly, IP geographical location information is not involved in any expansion logic.

[0124] For WHOIS feature dimensions: based on the extracted WHOIS registration unit, contact information, and... Based on the judgment results, other domain names and assets registered by the same registrant will be investigated; among them, if If the value is ≥ 0.8, then priority should be given to expanding this type of asset.

[0125] After obtaining asset expansion results from different analytical dimensions, a list of related asset leads is generated, which clearly defines the analytical dimension, the basis for association, and the initial association confidence level for each expanded asset (which can be obtained from the target network assets). (or the corresponding single-dimensional confidence level), providing data support for the subsequent closed-loop feedback mechanism, while achieving deep coupling with the aforementioned modules.

[0126] Furthermore, this application embodiment also constructs a confidence-driven closed-loop feedback mechanism, covering the entire process of "multimodal feature extraction → independent confidence determination of each dimension → GWO-Attention fusion → automatic clue expansion," using the determination accuracy and expansion accuracy of each stage as core feedback data to reverse-optimize the parameters and weights of the entire process, continuously improving the accuracy of attribution determination. The identification method further includes: Step 1: Collect feedback data: The collected feedback data is comprehensively related to each step of the identification method in the embodiments of this application, specifically including: Accuracy of confidence level determination for each single dimension: Statistics , , , , Deviations from actual attribution results, with a focus on cloud asset scenarios. The rationale for =0, and in non-cloud asset scenarios The accuracy of the judgment.

[0127] Accuracy of GWO-Attention multimodal feature fusion model: Statistics The deviation from the actual attribution results was analyzed to determine the rationality of the weight allocation, with a focus on verifying the implementation of resetting IP-related rights to 0 in the cloud asset scenario.

[0128] Automatic lead expansion accuracy: Statistics on expanded assets The deviation between the judgment result and the actual attribution result is analyzed to determine the validity of the extended basis (features and confidence level).

[0129] Feature validity data: Statistically evaluate the utilization rate and validity of the original feature information of each analysis dimension, focusing on assessing the auxiliary role of IP geographic location features in non-cloud asset scenarios, and ensuring that they do not participate in any confidence calculations or cloud asset-related processes.

[0130] The second step, feedback optimization: Based on the collected feedback data, the parameters of each step of the identification method in this application embodiment are optimized in reverse (the parameters for determining single-dimensional confidence, the weights of the multimodal feature fusion model, the parameters for related asset mining, and the feature information of each analysis dimension), to achieve a closed-loop iteration throughout the entire process. Specific optimization content includes: Optimize the confidence level parameters for each single dimension: Adjust the weighting coefficients of the confidence level calculation formulas based on the decision bias of each confidence level (e.g., ...). ), correct matching degree (e.g. , The judgment criteria improve the accuracy of single-dimensional confidence and ensure , , The integration with front-end features and back-end fusion models is smoother.

[0131] Optimize the weights of the multimodal feature fusion model: Adjust the initial weights and optimization strategies of the model based on the fusion accuracy and data conflict situation. Focus on optimizing the logic of setting IP-related weights to 0 in the cloud asset scenario to ensure that the features and confidence of each analysis dimension are used reasonably and to further resolve multi-source data conflicts.

[0132] Optimize automatic lead expansion rules: Adjust the expansion trigger threshold based on expansion accuracy. The values ​​of the extensions and the priority of each extension basis are determined, the extension weight of certificate, WHOIS, and DNS associations is strengthened, the use of IP geolocation information in the extension is strictly avoided, and the extension results are ensured to be accurate.

[0133] Optimize multimodal feature extraction: Based on the feature validity data, supplement valid features and eliminate invalid features, focusing on improving the extraction logic of IP association features, clearly distinguishing the uses of IP segments, operators and IP geographical location features, and ensuring that the rule of "IP geographical location is not used as a basis for cloud asset judgment" continues to be effective; at the same time, the feature data of valid associated assets expanded by automatic clues will be added to the feature library of the multimodal feature extraction module to enrich the sample data and improve the accuracy of subsequent judgments.

[0134] Through the above methods, each optimization of the closed-loop feedback mechanism will synchronously update the parameters of each front-end stage (feature extraction, confidence judgment) and the weights of the back-end fusion model, ensuring the continuity and consistency of the entire process, realizing continuous iteration of "judgment → expansion → feedback → optimization", and ensuring that all features, confidence levels and technical rules mentioned above are fully utilized.

[0135] In practical implementation, the asset assessment results with high confidence ( ≥0.8) are directly added to the organization-asset mapping table; low confidence score results ( If the value is less than 0.8, manual review will be triggered, and the review results can update the feature weights.

[0136] For adaptive adjustment of feature weights, the weight coefficients of each feature module can be dynamically adjusted based on the accuracy of the judgment results. The formula for adaptive weight adjustment is expressed as:

[0137] in, For the first Weight coefficients for the next iteration For learning rate, For the first The accuracy of the determination of each feature module.

[0138] Furthermore, the audit results are also used to update the organization name database. The formula for calculating organization name similarity is as follows:

[0139] when If the value is less than 0.7, it is considered a new organization and added to the organization name database. Ultimately, an asset-organization knowledge graph can be constructed to visually represent asset ownership relationships; and a timeline view can be provided to show the historical evolution of asset ownership.

[0140] The method for identifying network assets provided in this application has the following beneficial technical effects: First, multi-source feature fusion improves judgment accuracy: By integrating five dimensions of features, including DNS history, WHOIS, content, certificate, and IP, and combining them with dynamic weight optimization of the multi-modal feature fusion model, the accuracy of organization affiliation determination is improved. Among them, dynamic affiliation analysis enhances timeliness: Through the analysis of DNS historical change records, domain name affiliation change events can be identified, effectively solving the affiliation drift problem.

[0141] Second, the ability to identify cloud assets has been significantly improved: by directly correlating and analyzing cloud service metadata, the accuracy of cloud asset identification has been improved.

[0142] Third, automated clue expansion improves coverage: Through a confidence-driven closed-loop feedback mechanism, hidden related assets are automatically discovered and identified, significantly improving asset mapping coverage.

[0143] Fourth, intelligent optimization of multimodal feature weights: The improved multimodal feature fusion model can dynamically adjust the weights according to the feature quality, which improves the confidence calculation accuracy compared with the fixed weight method.

[0144] Based on the same inventive concept, embodiments of this application also provide a device for identifying network assets. Please refer to... Figure 2 , Figure 2 This is a schematic diagram of the structure of a network asset identification device provided in an embodiment of this application. Figure 2 As shown, the identification device 200 includes: The extraction module 210 is used to extract feature information of the target network asset about each of multiple analysis dimensions based on the basic data of the target network asset; wherein, the multiple analysis dimensions include at least two of the following: DNS historical change feature dimension, WHOIS feature dimension, website content feature dimension, certificate feature dimension, and IP association feature dimension; The first confidence determination module 220 is used to determine the single-dimensional confidence that the target network asset belongs to the currently determined organization under each analysis dimension based on the feature information of that analysis dimension. The weight determination module 230 is used to encode the feature information of each analysis dimension into a corresponding feature representation, and input the feature representation of each analysis dimension and the single-dimensional confidence into the multimodal feature fusion model to determine the first feature weight of the feature representation of each analysis dimension and the second feature weight of the single-dimensional confidence of each analysis dimension. The second confidence determination module 240 is used to determine the fusion confidence of the target network asset belonging to the currently determined organization based on the feature information of each analysis dimension and the corresponding first feature weight, as well as the single-dimensional confidence and the corresponding second feature weight. The determination module 250 is used to determine the organizational affiliation of the target network asset based on the fusion confidence level of the target network asset belonging to the currently determined organization, and obtain the network asset affiliation result.

[0145] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device 300 includes a processor 310, a memory 320, and a bus 330.

[0146] The memory 320 stores machine-readable instructions that can be executed by the processor 310. When the electronic device 300 is running, the processor 310 and the memory 320 communicate via the bus 330. When the machine-readable instructions are executed by the processor 310, the steps of the network asset identification method as described in the above method embodiment can be performed. For specific implementation details, please refer to the method embodiment, which will not be repeated here.

[0147] This application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it can execute the steps of the network asset identification method as described in the above method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0148] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0150] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0151] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0152] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0153] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for identifying network assets, characterized in that, The identification method includes: Based on the basic data of the target network assets, feature information of the target network assets in each of the multiple analysis dimensions is extracted; wherein, the multiple analysis dimensions include at least two of the following: DNS historical change feature dimension, WHOIS feature dimension, website content feature dimension, certificate feature dimension, and IP association feature dimension; For each analysis dimension, the single-dimensional confidence level of the target network asset belonging to the currently judged organization is determined based on the feature information of that analysis dimension; The feature information of each analysis dimension is encoded into a corresponding feature representation, and the feature representation of each analysis dimension and the single-dimensional confidence are input into the multimodal feature fusion model to determine the first feature weight of the feature representation of each analysis dimension and the second feature weight of the single-dimensional confidence of each analysis dimension. Based on the feature information and corresponding first feature weight of each analysis dimension, as well as the single-dimensional confidence and corresponding second feature weight, the fusion confidence of the target network asset belonging to the currently determined organization is determined. Based on the fusion confidence level of the target network asset belonging to the currently determined organization, the organizational affiliation of the target network asset is determined, and the network asset affiliation result is obtained.

2. The identification method according to claim 1, characterized in that, The identification method further includes: When the fusion confidence level is greater than the preset confidence level threshold, other related network assets associated with the target network asset are mined from different analysis dimensions based on the single-dimensional confidence level and feature representation under each analysis dimension. A list of related asset clues is generated based on the mining results from different analytical dimensions.

3. The identification method according to claim 2, characterized in that, Based on the fusion confidence level of the target network asset belonging to the currently determined organization, the organizational affiliation of the target network asset is determined to obtain the network asset affiliation result, including: When the fusion confidence level is greater than or equal to the first preset threshold, the network asset ownership result of the target network asset is determined to be the current determined organization; When the fusion confidence level is less than the first preset threshold, it is submitted for manual review, and the manual review result is determined as the network asset ownership result of the target network asset. The identification method further includes: Collect feedback data; Based on the feedback data, at least one of the following is updated: the single-dimensional confidence judgment parameter, the weight of the multimodal feature fusion model, the parameters of associated asset mining, and the feature information of each analysis dimension.

4. The identification method according to claim 1, characterized in that, For the DNS historical change feature dimension, feature information This includes the domain name's historical time sequence, changes in the IP address the domain name points to, and changes in the geographical location of the server the domain name resolves to; then, determining the single-dimensional confidence level of the target network asset belonging to the currently judged organization based on the feature information of this analysis dimension includes: Based on the historical time series of the domain names, the frequency of changes in domain name resolution records is determined, and the time continuity evaluation index value is obtained; Based on the IP address change records pointed to by the domain name, the matching degree between the organization to which the IP pointed to by the domain name belongs and the currently determined organization is determined, and the IP attribution consistency evaluation index value is obtained; Based on the server geographical location change records resolved by the domain name, the rationality of the domain name change is determined, and the rationality evaluation index value of the change is obtained; Based on the time continuity assessment index value, the IP attribution consistency assessment index value, and the change rationality assessment index value, determine the single-dimensional confidence level corresponding to the DNS historical change feature dimension.

5. The identification method according to claim 1, characterized in that, For the website content feature dimension, the feature information includes website content; then, determining the single-dimensional confidence level of the target network asset belonging to the currently judged organization under this analysis dimension based on the feature information of this analysis dimension includes: Based on the semantic similarity between the organization identifier in the website content and the organization information of the currently identified organization in the organization name database, the text semantic matching index value is determined; The content pattern matching index value is determined based on the degree of matching between the website structure in the website content and the organizational characteristics of the currently determined organization. Based on the consistency between the website content and the information organized by the target network asset in at least one other information source, a multi-source content cross-validation index value is determined; wherein, the information organized by the target network asset in at least one other information source includes the registrant information in WHOIS and / or the SSL certificate organization name; Based on the text semantic matching index value, the content pattern matching index value, and the multi-source content cross-validation index value, determine the single-dimensional confidence level corresponding to the website content feature dimension.

6. The identification method according to claim 1, characterized in that, Regarding the IP association feature dimension, when identifying an IP address as belonging to a cloud service provider, the feature information includes metadata information provided by the cloud service provider; the identification method further includes determining the organization to which the cloud service asset belongs based on the metadata information in the following ways: Extract the first organization name from the cloud resource group name and the second organization name from the cloud resource tag using regular expressions or keyword extraction algorithms; The validity of the first organization name and the second organization name is cross-validated based on the single-dimensional confidence scores of WHOIS feature dimension and certificate feature dimension, respectively. If the verification passes, the matching degree between the first organization name and the name of the currently determined organization is determined, and the similarity between the second organization name and the name of the currently determined organization is determined. If the matching degree is greater than a preset matching degree threshold and / or the similarity degree is greater than a preset similarity threshold, then the currently determined organization is identified as the target organization to which the cloud service asset belongs.

7. The identification method according to claim 1, characterized in that, The multimodal feature fusion model combines an attention mechanism and an improved gray wolf optimization algorithm; the step of inputting the feature representation and single-dimensional confidence of each analysis dimension into the multimodal feature fusion model to determine the first feature weight of the feature representation of each analysis dimension and the second feature weight of the single-dimensional confidence of each analysis dimension includes: Configure the feature representation for each analysis dimension and the initial feature weights corresponding to the confidence scores of each dimension; By using the improved Grey Wolf optimization algorithm, the initial feature weights are initially iteratively optimized. During the optimization process, the weight allocation of the conflict analysis dimensions is adjusted to obtain the feature representation of each analysis dimension and the first optimized weight of the single-dimensional confidence. Based on the feature representation and single-dimensional confidence of each analysis dimension, the correlation score between each feature representation, each single-dimensional confidence and the currently judged organization is determined through an attention mechanism, and the weights are adjusted based on the correlation score to obtain the second optimized weights of the feature representation and single-dimensional confidence of each analysis dimension. Cross-validate the feature representations of each analysis dimension and the first and second optimized weights of the single-dimensional confidence; If the verification passes, the first feature weight of the feature representation of each analysis dimension and the second feature weight of the single-dimensional confidence are determined based on the feature representation of each analysis dimension and the first and second optimized weights of the single-dimensional confidence. If the validation fails, the feature weights are iteratively optimized again using the improved Grey Wolf optimization algorithm to obtain the updated feature representations for each analysis dimension and the first optimized weights for the single-dimensional confidence, and then cross-validation is performed again.

8. A device for identifying network assets, characterized in that, The identification device includes: The extraction module is used to extract feature information of the target network asset for each of multiple analysis dimensions based on the basic data of the target network asset; wherein, the multiple analysis dimensions include at least two of the following: DNS historical change feature dimension, WHOIS feature dimension, website content feature dimension, certificate feature dimension, and IP association feature dimension; The first confidence level determination module is used to determine the single-dimensional confidence level of the target network asset belonging to the currently judged organization under each analysis dimension based on the feature information of that analysis dimension. The weight determination module is used to encode the feature information of each analysis dimension into a corresponding feature representation, and input the feature representation of each analysis dimension and the single-dimensional confidence into the multimodal feature fusion model to determine the first feature weight of the feature representation of each analysis dimension and the second feature weight of the single-dimensional confidence of each analysis dimension. The second confidence determination module is used to determine the fusion confidence that the target network asset belongs to the currently determined organization based on the feature information of each analysis dimension and the corresponding first feature weight, as well as the single-dimensional confidence and the corresponding second feature weight. The determination module is used to determine the organizational affiliation of the target network asset based on the fusion confidence level of the target network asset belonging to the currently determined organization, and to obtain the network asset affiliation result.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the network asset identification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the network asset identification method as described in any one of claims 1 to 7.