Enterprise internet asset intelligent discovery and safety management and protection system

By optimizing the execution of directed graph parsing and the collaborative work of containerized toolsets, combined with real-time data fusion processing, and building a dynamic asset knowledge graph, we have solved the problems of rigid tool processes and inefficient data processing in enterprise Internet asset management, and achieved efficient and comprehensive asset management and accurate risk detection.

CN120707304AActive Publication Date: 2025-09-26HANGZHOU BILING SAFETY TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511142590.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-26
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing enterprise Internet asset management solutions have problems such as rigid discovery tool processes, inefficient multi-source heterogeneous data processing, and difficulty ensuring data status synchronization and consistency, which leads to blind spots in enterprise network security management and inefficient security measures.

Method used

An enterprise Internet asset intelligent discovery and security management system is adopted. By optimizing the execution of directed graph parsing and the collaborative work of containerized toolsets, the automation and efficient collaboration of asset discovery are realized. The multi-source heterogeneous data is converted into a structurally unified asset entity stream through the real-time data fusion processing layer, and a dynamic asset knowledge graph is constructed for risk detection and automated response.

Benefits of technology

It achieves efficient and comprehensive discovery and management of enterprise Internet assets, provides accurate risk detection and automated response capabilities, solves the problems of low discovery efficiency and complex data processing in traditional solutions, and improves the comprehensiveness and real-time nature of network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707304A_ABST
    Figure CN120707304A_ABST
Patent Text Reader

Abstract

The invention relates to the field of internet asset management, and particularly discloses an enterprise internet asset intelligent discovery and safety management and protection system which analyzes and converts a high-level asset discovery instruction into an optimized execution directed graph. The atlas can automatically drive a bottom-layer distributed and containerized discovery tool set to work efficiently and cooperatively, so that tedious and low-efficiency manual tool chain operation is replaced. Meanwhile, in order to solve the problem that multi-source heterogeneous data is difficult to fuse and utilize, the system introduces an original data stream generated by a discovery tool into a real-time data fusion processing layer, and the original data stream is converted into a unified asset entity stream with a unified structure and complete information through the steps of standardization, deduplication, aggregation and the like. Finally, the dynamic asset knowledge graph is constructed based on the high-quality data, accurate risk detection and automatic response processing are realized, and the problems of a traditional scheme in discovery efficiency and data processing are systematically solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet asset management, and more specifically, to an enterprise Internet asset intelligent discovery and security management system. Background Art

[0002] With the deepening of enterprise digital transformation and the widespread adoption of cloud computing, the number, variety, and distribution complexity of enterprise assets exposed to the internet, such as domain names, IP addresses, port services, web applications, and even cloud service configurations, have exploded. This dynamic, heterogeneous, and widely distributed asset landscape makes comprehensive, accurate, and timely understanding of one's digital asset landscape the cornerstone and primary challenge of enterprise network security. Without a clear and comprehensive understanding of the asset landscape, enterprises face significant security blind spots, making it impossible to effectively assess and manage potential attack surfaces. Any subsequent security measures will be significantly undermined by this weak foundation. Therefore, building a solution that intelligently discovers and effectively manages enterprise internet assets has become a pressing need for enterprise network security.

[0003] To address this challenge, the industry has proposed several attack surface management solutions. However, existing technologies often suffer from significant shortcomings. For one thing, most of these solutions rely on the combined use of multiple independent security tools (such as subdomain scanning, port scanning, and vulnerability scanning tools). The lack of a unified orchestration and scheduling framework leads to complex and inefficient data transfer between different tools during asset discovery, and difficulty ensuring state synchronization and consistency, making the entire discovery process time-consuming and error-prone. Furthermore, the raw asset data generated by these tools varies in format and quality, with significant amounts of redundant and conflicting information. Traditional processing methods, often in batch mode, not only lack timely information and are unable to adapt to the dynamic changes in modern enterprise assets, but also lack intelligent data fusion and correlation analysis capabilities, making it difficult to integrate scattered data points into a high-quality, structured asset view. This makes it difficult for security teams to form an accurate and comprehensive understanding of the asset landscape.

[0004] Therefore, an optimized enterprise Internet asset intelligent discovery and security management system is desired. Summary of the Invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide an enterprise Internet asset intelligent discovery and security management system.

[0006] According to one aspect of the present application, a system for intelligent discovery and security management of enterprise Internet assets is provided, comprising: An enterprise Internet asset discovery module, configured to respond to an asset discovery instruction and perform intelligent discovery of enterprise Internet assets based on an optimized execution directed graph to obtain an original asset data stream; An asset entity stream conversion module, configured to convert the original asset data stream into a unified asset entity stream; An asset knowledge graph construction and analysis module, configured to construct a dynamic asset knowledge graph and conduct association analysis on the unified asset entity flow to obtain an asset knowledge graph; A risk detection module, configured to perform risk detection on the asset knowledge graph based on a security policy to obtain a risk detection result; The response action generation module is used to generate an automated response action for the risk detection result based on the automation rules.

[0007] Compared with the existing technology, the present application provides an enterprise Internet asset intelligent discovery and security management system, which parses high-level asset discovery instructions and converts them into an optimized execution directed graph. This graph can automatically drive the underlying distributed, containerized discovery tool set to work efficiently and collaboratively, thereby replacing the cumbersome and inefficient manual tool chain operations. At the same time, in order to solve the problem of multi-source heterogeneous data being difficult to integrate and utilize, the system introduces the original data stream generated by the discovery tool into the real-time data fusion processing layer, and converts it into a unified asset entity stream with unified structure and complete information through steps such as standardization, deduplication and aggregation. Finally, a dynamic asset knowledge graph is constructed based on this high-quality data to achieve accurate risk detection and automated response and disposal, thereby systematically solving the problems of traditional solutions in discovery efficiency and data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0009] Figure 1 This is a system block diagram of an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application.

[0010] Figure 2 Schematic diagram of data flow of the enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application.

[0011] Figure 3 This is a block diagram of an enterprise Internet asset discovery module in the enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application.

[0012] Figure 4 This is a block diagram of the asset entity flow conversion module in the enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application.

[0013] Figure 5 This is a block diagram of a fingerprint enrichment unit in the enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application. DETAILED DESCRIPTION

[0014] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0015] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0016] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.

[0017] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0018] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0019] When conducting enterprise Internet asset management, the existing technology generally has the core technical problems of rigid discovery tool processes, fragmented multi-source heterogeneous data processing processes, and low efficiency. In order to systematically solve the above problems, the technical solution of this application proposes an enterprise Internet asset intelligent discovery and security management system. The system starts with a high-level asset discovery instruction. The system first intelligently parses it and converts it into an optimized execution directed graph. The graph can automatically orchestrate and schedule a series of independent containerized discovery tools at the bottom layer to achieve the optimal path collaborative execution of tasks, thereby completely breaking the deadlock of tool islands. When the tool set is executed, the multi-source heterogeneous raw data streams it generates are not processed in isolation, but are immediately sent to a real-time data processing pipeline. In this pipeline, the data stream is normalized, fingerprinted, efficiently deduplicated, and finally aggregated into stateful entities, transforming the scattered and chaotic raw data into a structured and information-complete asset entity stream. Ultimately, these high-quality asset entities are used to dynamically build a global asset knowledge graph, providing a solid and reliable data foundation for subsequent accurate risk detection and automated security management, thereby achieving full process automation and intelligence from command to response.

[0020] In the technical solution of this application, a system for intelligent discovery and security management of enterprise Internet assets is proposed. Figure 1 This is a system block diagram of an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of enterprise Internet asset intelligent discovery and security management system according to the embodiment of the present application. Figure 1 and Figure 2 As shown, the enterprise Internet asset intelligent discovery and security management system 100 according to the embodiment of the present application includes: an enterprise Internet asset discovery module 110, which is used to respond to asset discovery instructions and perform enterprise Internet asset intelligent discovery based on the optimized execution directed graph to obtain the original asset data stream; an asset entity stream conversion module 120, which is used to convert the original asset data stream into a unified asset entity stream; an asset knowledge graph construction and analysis module 130, which is used to perform dynamic asset knowledge graph construction and association analysis on the unified asset entity stream to obtain an asset knowledge graph; a risk detection module 140, which is used to perform risk detection on the asset knowledge graph based on security policies to obtain risk detection results; and a response action generation module 150, which is used to generate automated response actions for the risk detection results based on automation rules.

[0021] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the enterprise internet asset discovery module 110 is configured to respond to asset discovery instructions and perform intelligent discovery of enterprise internet assets based on an optimized execution directed graph to obtain the original asset data stream. It should be understood that existing technologies heavily rely on security personnel manually and sequentially running various single-function scanning tools. Specifically, when a security team needs to conduct a comprehensive asset inventory of a primary domain, the traditional approach is to manually run multiple tools, such as subdomain enumeration, port scanning, and service identification. This approach is not only inefficient and rigid, but also prone to omissions due to operational errors. Furthermore, the tools are isolated from each other, forming isolated tool islands that are difficult to coordinate, resulting in a time-consuming discovery process and prone to omissions. Therefore, in the practice of enterprise asset security management, in response to asset discovery instructions, intelligent discovery of enterprise internet assets is performed based on an optimized execution directed graph to obtain the original asset data stream. By parsing and converting the user's asset discovery requirements into a computable and optimizable execution directed graph, automated intelligent orchestration and scheduling of the underlying diverse discovery tools is achieved. In this way, the system can intelligently drive various tool containers to work in parallel or serially based on this graph. This greatly improves the overall efficiency and coverage of asset discovery, and automatically generates a continuous stream of raw asset data containing all original discovery results, laying an efficient and reliable foundation for subsequent unified processing and analysis.

[0022] Figure 3 FIG is a block diagram of an enterprise Internet asset discovery module in an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application. Figure 3 As shown, in an embodiment of the present application, the enterprise Internet asset discovery module 110 includes: an instruction parsing and graphing unit 111, which is used to perform task instruction parsing and execution path graphing on the asset discovery instruction to obtain the optimized execution directed graph; a container resource scheduling unit 112, which is used to perform graph instantiation and container resource scheduling based on the optimized execution directed graph to obtain a scheduled tool container instance set; a tool container instance set driving unit 113, which is used to drive the scheduled tool container instance set based on the topology of the optimized execution directed graph to obtain the original heterogeneous output; a standardization and streaming processing unit 114, which is used to perform standardization and streaming based on the Sidecar agent on the original heterogeneous output to obtain the original asset data stream.

[0023] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the instruction parsing and graphing unit 111 is used to perform task instruction parsing and execution path graphing on the asset discovery instructions to obtain the optimized execution directed graph. It should be understood that the traditional asset discovery process relies heavily on the personal experience of security engineers, combining and invoking various tools manually or through the writing of ad hoc scripts. This approach is not only inefficient and has poor reproducibility, but also fails to guarantee the global optimality of the execution path. Therefore, in the actual application scenario of enterprise asset discovery, performing task instruction parsing and execution path graphing on the asset discovery instructions to obtain the optimized execution directed graph systematically transforms vague human instructions into a precise, quantifiable, and machine-understandable and executable optimal task planning blueprint, thereby achieving automation, standardization, and intelligence in the asset discovery process. In this way, a structured optimized execution directed graph is generated that clearly defines the tools, execution order, and dependencies required to complete the specified discovery task, providing a deterministic and efficient execution basis for subsequent container resource scheduling and task-driven execution.

[0024] Specifically, the instruction parsing graph unit is used to: perform syntax parsing and semantic extraction on the asset discovery instruction to obtain a semantic task primitive; generate a candidate tool chain for the semantic task primitive based on a tool capability library; convert the candidate tool chain into a directed acyclic graph, and calculate the cost value of each graph node in the directed acyclic graph based on the execution constraints in the semantic task primitive to obtain a weighted unoptimized directed acyclic graph; perform shortest path solving on the weighted unoptimized directed acyclic graph to obtain the optimized execution directed graph.

[0025] In other words, the implementation of this process involves a multi-stage transformation process. More specifically, when the system receives an asset discovery instruction, such as "conduct a comprehensive web asset discovery of a target company," it first parses and extracts semantics from the instruction. Using natural language processing or a predefined domain-specific language parser, it decomposes the instruction into a set of structured, semantically defined task primitives, such as {Target: Target Company, Task Type: Web Asset Discovery, Scope: Comprehensive}. Next, based on a pre-defined tool capability library, the system searches for tools or tool combinations that meet these task primitives, generating multiple candidate tool chains. For example, one tool chain might be "subdomain enumeration tool -> port scanning tool -> web service identification tool," while another might be "passive domain collection tool -> web fingerprinting tool." The system then integrates these candidate toolchains and transforms them into a directed acyclic graph (DAG), where each tool is a graph node and data flows form directed edges. Simultaneously, based on the execution constraints in the semantic task primitives (such as time limits and resource consumption preferences), the system calculates the cost of each node in the graph. For example, tool nodes with long execution times or high resource consumption are assigned higher costs, resulting in a weighted, unoptimized DAG. Finally, the system applies a shortest path solving algorithm, such as Dijkstra, to this weighted graph to calculate the path with the lowest total cost from the starting node to the end node. This path is the final output, the optimized execution directed graph.

[0026] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the container resource scheduling unit 112 is configured to perform graph instantiation and container resource scheduling based on the optimized execution directed graph to obtain a set of scheduled tool container instances. It should be understood that in actual enterprise security operations and maintenance scenarios, the optimized execution directed graph generated in the previous step is merely a static, logical execution plan and cannot itself execute any operations. Therefore, graph instantiation and container resource scheduling are performed based on the optimized execution directed graph to obtain a set of scheduled tool container instances, transforming this abstract logical blueprint into executable software entities in the physical world. This allows the necessary computing resources to be allocated to each upcoming tool task in the computing cluster according to the graph plan, and the corresponding tool program to be launched. The system transitions from a purely planning phase to a ready-to-execute phase, producing a set of all launched, resource-allocated, and ready-to-run tool container instances. This ensures the efficiency and reliability of subsequent task-driven execution and avoids execution failures due to insufficient resources or environmental issues.

[0027] Specifically, in a specific example of the present application, the implementation of this process relies on a modern container orchestration platform, such as Kubernetes. More specifically, when the system's container resource scheduling unit receives the optimized execution directed graph, it will first traverse the nodes in the graph and parse out the tool represented by each node (such as the subdomain enumeration tool subfinder), its operating parameters, and preset resource requirements (such as CPU, memory). Then, for each tool node, the scheduling unit will dynamically generate a deployment description file that can be recognized by the container orchestration system, such as a YAML list of a Kubernetes Pod or Job. The list will clearly specify the tool container image to be pulled, the startup command, the parameters to be passed, and the resource requests and restrictions. Subsequently, the scheduling unit submits this description file to the control center of the container orchestration platform through the API. The scheduler of the orchestration platform will intelligently select the most suitable physical or virtual node to deploy the container instance of the tool based on the real-time load and available resources of each node in the cluster. Once a container instance is successfully started and enters the running state on the specified node, it will be registered and added to a dynamic set until all tools that need to be initially started in the graph have been instantiated, eventually forming the scheduled tool container instance set, fully prepared for the next step of driver execution.

[0028] In the above-mentioned enterprise Internet asset intelligent discovery and security management system 100, the tool container instance set driving unit 113 is used to drive the scheduled tool container instance set based on the topological structure of the optimized execution directed graph to obtain the original heterogeneous output. It should be understood that in the process of enterprise security asset discovery, it is not enough to simply instantiate the tool and allocate resources. These tool instances are in a standby state and will not execute automatically. Therefore, a coordinator is required to accurately command and trigger these independent tool units to work in sequence according to the preset battle plan. In the technical solution of the present application, based on the topological structure of the optimized execution directed graph, the scheduled tool container instance set is driven to obtain the original heterogeneous output. This realizes the transformation from static planning to dynamic execution, that is, through a central driving mechanism, strictly following the tool dependencies and data flow defined by the optimized execution directed graph to orchestrate the actual operation of the entire discovery task. This enables the entire asset discovery task to be completed efficiently and orderly, and produces a series of original discovery results generated by different tools with different formats and contents, namely original heterogeneous outputs, which provide the most basic original materials for the subsequent data standardization and fusion stages.

[0029] Specifically, in one specific example of this application, the optimized execution directed graph is first topologically sorted to identify all zero-indegree start nodes that do not depend on the output of any other tool. Next, the parameters in the initial asset discovery instructions (e.g., the primary domain name) are used as input to trigger the execution of the tool container instances corresponding to these start nodes (e.g., the subdomain enumeration tool). When a tool container instance completes its task, the driver captures the output data it generates. Simultaneously, based on the graph's topology, it identifies all successor nodes that directly depend on the node's output. The captured output data is then passed as input to the tool container instances corresponding to these successor nodes (e.g., passing a subdomain list to a port scanning tool), instructing them to start execution. This process iterates along the edges of the graph until all nodes have been executed. During this process, the outputs generated by all tool container instances are aggregated to form the original heterogeneous output.

[0030] In the above-mentioned enterprise Internet asset intelligent discovery and security management system 100, the standardization and streaming processing unit 114 is used to standardize and stream the original heterogeneous output based on the Sidecar agent to obtain the original asset data stream. It should be understood that in the complex scenario of enterprise asset discovery, the original heterogeneous outputs generated by different tools have different formats and structures. For example, the subdomain tool outputs a plain text list, while the vulnerability scanning tool may output a complex XML or JSON report. This inconsistency in the data source makes subsequent unified processing and analysis extremely difficult. Therefore, the original heterogeneous output is further standardized and streamed based on the Sidecar agent to solve this data heterogeneity problem in real time. In this way, the responsibility for data format conversion and streaming can be separated from the core discovery tool, and all scattered and batched original outputs can be converted into a unified and continuous data stream form in real time through an accompanying agent. The system integrates the originally chaotic, non-real-time multi-source data into a structured, easy-to-consume raw asset data stream. This not only greatly simplifies the design complexity of subsequent data processing modules, but also lays the foundation for the real-time response capability of the entire system.

[0031] Specifically, in one specific example of this application, this process is implemented by equipping each tool container instance with a sidecar proxy container. This sidecar proxy is deployed in the same execution unit (e.g., a Kubernetes pod) as the main tool container and is configured to capture the main tool container's standard output stream or monitor its output file. When the main tool container, such as a port scanning tool, generates a line of output indicating an open IP and port, the sidecar proxy immediately intercepts this data. The proxy has pre-built parsing logic for different tool output formats. Based on the current tool type, it parses this line of raw text and converts it into a predefined standardized data structure, such as a JSON object containing fields such as asset type, IP address, port, protocol, and timestamp. After completing the standardized conversion, the sidecar proxy does not cache the data but immediately publishes the JSON object as a message to a specific topic in a centralized message queue system (e.g., Apache Kafka). All tool sidecar proxies continuously push messages to this queue, thereby converging into a unified, real-time stream of raw asset data.

[0032] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the asset entity stream conversion module 120 is used to convert the raw asset data stream into a unified asset entity stream. It should be understood that while the raw asset data stream has a unified format, it is still essentially discrete, instantaneous data points, such as a certain IP address opening a certain port at a certain moment. These data points suffer from significant redundancy (e.g., multiple scans of the same port), incomplete information (e.g., only the IP address and port number, but no service details), and a lack of relevance. Therefore, converting the raw asset data stream into a unified asset entity stream aggregates these scattered, event-centric data points into an asset-centric, stable, and information-rich entity view. Through a series of in-depth data processing and fusion operations, the raw data stream is purified, enriched, and aggregated to construct a complete profile that comprehensively and accurately describes each of the enterprise's digital assets. This includes removing duplicate information, supplementing detailed asset fingerprint features (e.g., running services, used frameworks), and associating and aggregating different fragments describing the same asset (e.g., domain name, IP address, port number, application) to form a logically complete asset entity.

[0033] Figure 4 FIG is a block diagram of an asset entity flow conversion module in an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application. Figure 4As shown, in an embodiment of the present application, the asset entity stream conversion module 120 includes: an original asset data standardization unit 121, which is used to normalize and event the multi-source heterogeneous data stream of the original asset data stream to obtain a standardized discrete record stream; a fingerprint enrichment unit 122, which is used to perform multimodal fingerprint enrichment on the standardized discrete record stream to obtain a fingerprint enriched record stream; an efficient streaming deduplication unit 123, which is used to perform efficient streaming deduplication based on a hybrid strategy on the fingerprint enriched record stream to obtain a unique asset fragment stream; a stateful entity aggregation unit 124, which is used to perform stateful entity aggregation based on an aggregate primary key on each unique asset fragment in the unique asset fragment stream to obtain the unified asset entity stream.

[0034] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the raw asset data standardization unit 121 is used to normalize and eventize the raw asset data stream, a multi-source, heterogeneous data stream, to produce a standardized discrete record stream. It should be understood that in the actual enterprise asset discovery process, while the raw asset data stream generated by the sidecar agent tends to be consistent in transmission format, it remains heterogeneous at the semantic level of the data content. The output fields and meanings of different tools vary greatly, making unified analysis and processing impossible. Therefore, the raw asset data stream is further normalized and eventized to eliminate the semantic gap caused by tool diversity. The core purpose of this step is to establish a globally unified, standardized data model and transform each piece of raw discovery information into an independent, standardized event record with contextual metadata, thereby providing a homogenized data foundation for all subsequent processing steps. The system transforms a mixed, semantically diverse data stream into a clear, regular, standardized discrete record stream, where each record follows the same data structure, significantly reducing the complexity of subsequent data fusion and analysis.

[0035] More specifically, in one example of this application, a stream processing application (for example, built on Apache Flink or Kafka Streams) subscribes to and consumes an upstream stream of raw asset data. When the application receives a raw data record from a message queue, it first deserializes it. Next, the application loads the corresponding parsing and mapping rules from a preconfigured rule library based on the source identifier carried in the record (for example, the tool name injected by the sidecar). Based on these rules, the application then maps fields in the raw record (for example, the host field output by a tool) to corresponding fields in the standard data model (for example, asset.value), normalizing the data type and format. During this process, the application also generates a unique event ID for the record and appends metadata such as the processing timestamp to complete the eventing operation. Finally, this new, normalized and evented record, which fully conforms to the standard data model, is serialized and published to a new message queue topic, forming the standardized discrete record stream for consumption by downstream modules.

[0036] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the fingerprint enrichment unit 122 is configured to perform multimodal fingerprint enrichment on the standardized discrete record stream to obtain a fingerprint-enriched record stream. It should be understood that in the context of enterprise asset security management, while the standardized discrete record stream has a uniform structure, the information it contains is often basic and superficial, such as simply knowing that port 80 is open on an IP address. This level of information is insufficient to support accurate risk assessment and attack surface analysis. Therefore, multimodal fingerprint enrichment is performed on the standardized discrete record stream to deeply mine and detail the underlying asset information, revealing deeper technical attributes and characteristics. Through various methods, such as active detection and passive analysis, detailed technical fingerprint information is added to each underlying asset record. This includes identifying the specific type and version of the web server (e.g., Nginx / 1.21.6), the backend application framework and language (e.g., Spring Boot, PHP), the components used and their versions (e.g., jQuery / 3.5.1), and even the operating system type. This multimodal fingerprint recognition covers multiple dimensions from the network layer to the application layer, aiming to build a three-dimensional and detailed asset technology portrait.

[0037] Figure 5 FIG is a block diagram of a type-compliant response subunit in an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application. Figure 5As shown, in an embodiment of the present application, the fingerprint enrichment unit 122 includes: a type compliance response subunit 1221, which is used to determine whether each standardized discrete record in the standardized discrete record stream complies with a preset fingerprint type based on the type field of each standardized discrete record in the standardized discrete record stream, and if so, perform a dynamic append operation on the standardized discrete record to obtain a qualified record stream with a feature vector; a record stream filling subunit 1222, which is used to perform atomic feature extraction and feature vector filling on each qualified record with a feature vector in the qualified record stream with a feature vector to obtain a qualified record stream filled with feature vectors; a parallelization logic judgment subunit 1223, which is used to perform parallelization logic judgment based on feature vectors on the qualified record stream filled with feature vectors to obtain an intermediate inference result set; and a fusion arbitration subunit 1224, which is used to input the intermediate inference result set and the standardized discrete record stream into a fusion arbitrator module to obtain the fingerprint enriched record stream.

[0038] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the type conformance response subunit 1221 is configured to determine whether each standardized discrete record in the standardized discrete record stream conforms to a preset fingerprinting type based on the type field of each standardized discrete record. If so, a dynamic append operation is performed on the standardized discrete record to obtain a qualified record stream with a feature vector. It should be understood that in the practice of enterprise asset security management, while standardized discrete record streams solve the problem of inconsistent data structures, they contain a large number of different types of records, such as domain names, IP addresses, and open ports. Not all records are suitable for or require in-depth technical fingerprinting. Therefore, based on the type field of each standardized discrete record in the standardized discrete record stream, whether it conforms to the preset fingerprinting type is determined and a dynamic append operation is performed to avoid ineffective and resource-consuming detection operations on non-target asset records. In this way, classified processing and targeted enrichment of asset records are achieved. The system first screens the incoming records and identifies only those with detectable service features (such as port records that open Web services or database services). Then, for these qualified records, it calls the corresponding fingerprint recognition module to perform in-depth information mining and appends the mined structured feature information back to the original record.

[0039] Specifically, in one specific example of this application, a stream processing node continuously consumes a stream of standardized discrete records. When the node receives a standardized discrete record, it first checks the record's type field. For example, if a record's type is "Open Port" and the service field is "HTTP," the record meets the preset web fingerprinting type. The node then triggers a dynamic append operation: it passes the IP and port information in the record to a web fingerprinting submodule. This submodule initiates an HTTP request to the target, analyzes the response header, response body, and icon hash, and uses a pre-configured fingerprint library to identify the web server as Nginx and the backend framework as Django. The node then constructs this identified information into a structured feature vector, such as a JSON object {server: Nginx, framework: Django}, and appends this feature vector to the original record, forming a new record with the feature vector. This enriched record is output downstream, becoming part of the qualified record stream. Records that do not meet the preset fingerprinting type (such as domain name records) are bypassed or sent to other processing logic.

[0040] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the record stream filling subunit 1222 is configured to extract atomic features and fill feature vectors from each qualified record with a feature vector in the qualified record stream with feature vectors, thereby obtaining a qualified record stream filled with feature vectors. It should be understood that the feature vectors attached to a qualified record stream with feature vectors are often raw and unstructured, such as a complete HTTP response header string or a snippet of HTML code. While rich in information, this raw data cannot be directly used for accurate comparison and aggregation. Therefore, in the technical solution of the present application, atomic feature extraction and feature vector filling are further performed on each record in the qualified record stream with feature vectors to decompose this rough, semi-structured feature information into the smallest, indivisible, atomic features with clear business meaning, thereby achieving in-depth analysis and normalization of asset technical features. The system extracts mixed information from the raw feature vectors, such as the component name and version number from a Server response header, and fills these extracted atomic features into a predefined, standard feature vector template with a unified structure. This process ensures that no matter how the original features vary, the final representation is consistent and regular.

[0041] Specifically, in a specific example of the present application, a stream processing node consumes a stream of qualified records with feature vectors. When the node receives a record whose original feature vector contains an HTTP response header "Server:Apache / 2.4.54 (Ubuntu)," the node first applies a series of preset regular expressions or parsers to extract atomic features from the string. It successfully extracts three atomic features: the component name is Apache, the version number is 2.4.54, and the operating system is Ubuntu. Next, the node retrieves a standard feature vector template, which may be a JSON structure that defines multiple fields such as component_name, component_version, and os_type. The node then fills the extracted atomic features into the corresponding fields of the template, forming a filled feature vector {component_name: Apache, component_version: 2.4.54, os_type: Ubuntu, ...}. Finally, the record carrying the filled, structured feature vector is output and merged into the qualified record stream that has already been filled with feature vectors.

[0042] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the parallelized logical judgment subunit 1223 is configured to perform parallelized logical judgment based on the feature vectors on the qualified record stream populated with feature vectors to obtain an intermediate inference result set. It should be understood that within the overall framework of enterprise asset security management, a record stream containing standardized feature vectors has been generated. These records accurately describe the technical composition of the assets, but they themselves remain isolated factual statements and lack qualitative or quantitative security judgments. Therefore, parallelized logical judgment based on the feature vectors is performed on the qualified record stream populated with feature vectors to automatically collate and correlate these purely technical facts with a vast security knowledge base, thereby revealing their potential security implications. By efficiently parallelizing and logically reasoning each record's standardized feature vector with a pre-set rule base (such as a vulnerability library, compliance baseline, and threat intelligence), the system can automatically generate preliminary, atomic security conclusions for each asset record, such as determining whether a component version has a known vulnerability or whether a configuration violates a security policy.

[0043] Specifically, in an embodiment of the present application, the parallelized logic judgment sub-unit is used to: extract the original feature vector of the first qualified record from the qualified record stream with the feature vector filled in; input the original feature vector of the first qualified record into a rule-based deterministic inference engine to obtain a deterministic discovery result; input the original feature vector of the first qualified record into a feature engineering preprocessor to obtain a feature engineering vector of the first qualified record; input the feature engineering vector of the first qualified record into a probabilistic classification engine to obtain a probabilistic discovery result; and aggregate the probabilistic discovery result and the deterministic discovery result to obtain an intermediate inference result.

[0044] More specifically, the original feature vector of the first qualified record is extracted from the stream of qualified records populated with feature vectors, and this original feature vector is input into a rule-based deterministic inference engine to obtain a deterministic discovery result. It should be understood that while the stream of qualified records populated with feature vectors accurately describes the technical composition of the asset, this information itself is merely neutral technical facts and lacks direct security value judgments. Therefore, records are extracted from the stream of qualified records populated with feature vectors and input into a rule-based deterministic inference engine to associate these objective technical attributes with clear security knowledge to draw unambiguous security conclusions. By comparing the standardized feature vector of each asset record against a rigid rule base composed of expert knowledge, vulnerability information, and compliance baselines, the system aims to quickly identify undisputed security issues that meet deterministic criteria, such as a component version that clearly contains a critical, publicly disclosed vulnerability.

[0045] In a specific example of the present application, an inference processing unit continuously consumes a stream of qualified records that have been filled with feature vectors. When the unit receives a record whose filled feature vector is {component_name: Log4j, component_version: 2.14.1}, the unit inputs this feature vector as a fact into a rule-based deterministic inference engine. The engine is preloaded with a series of rules, one of which may be defined as: when the component name is Log4j and its version number is greater than or equal to 2.0 and less than 2.15.0, a high-risk vulnerability conclusion is triggered. When the engine is executed, it matches the input feature vector with the conditions of this rule and finds that it is fully met. Therefore, the engine generates a deterministic discovery result, such as a structured data containing information such as the vulnerability number CVE-2021-44228 and the risk level of serious. This discovery result is then attached to the record or output as an independent inference event, together forming part of the intermediate inference result set.

[0046] More specifically, the original feature vector of the first qualified record is input into a feature engineering preprocessor to obtain a feature engineering vector for the first qualified record. It should be understood that in the scenario of enterprise asset security management, although the rule-based deterministic inference engine can efficiently handle security issues with clear characteristics, it is powerless for potential risks with complex patterns, fuzzy boundaries, and reliance on multiple factors. Therefore, the original feature vector of the first qualified record is input into a feature engineering preprocessor, and the original feature vector is converted into a high-dimensional, standardized, purely numerical feature engineering vector by performing a series of mathematical transformations and encodings on the original feature vector. This process aims to maximize the retention of the original information while expressing it in a normalized form that is friendly to machine learning algorithms, so as to facilitate the model's subsequent pattern recognition and probabilistic inference.

[0047] In a specific example of this application, the feature engineering preprocessor receives a qualified record whose original feature vector is {component_name: Nginx, component_version: 1.21.6, server_header: nginx / 1.21.6}. The preprocessor first performs one-hot encoding on the categorical feature component_name. Assuming that the predefined component vocabulary includes Nginx, Apache, and IIS, Nginx is converted to the vector [1, 0, 0]. Next, for the numerical feature component_version, its string 1.21.6 is parsed into three independent numerical features [1, 21, 6] and may be normalized. For the text feature server_header, the preprocessor may apply the TF-IDF algorithm or a pre-trained word embedding model (such as BERT) to convert it into a fixed-dimensional numerical vector. Finally, the preprocessor concatenates these processed numerical vectors into a single, high-dimensional, purely numerical feature-engineered vector, which is the feature-engineered vector of the first qualified record and is ready to be input into the subsequent machine learning model.

[0048] More specifically, the feature-engineered vector of the first qualified record is input into a probabilistic classification engine to generate a probabilistic discovery result. It should be understood that in the complex scenario of enterprise asset security management, deterministic inference engines can effectively identify security issues with clear rules, but they are unable to detect potential risks that are composed of multiple subtle features, have complex patterns, and lack deterministic rules. Therefore, the feature-engineered vector of the first qualified record is input into a probabilistic classification engine, leveraging the power of machine learning to identify and assess subtle security risks hidden within deep data correlations. Using a pre-trained classification model, the feature-engineered vector of the asset is deeply analyzed to probabilistically predict whether the asset falls into a specific risk category (e.g., a tampered server or an improperly configured development environment). By learning the features of a large number of known positive and negative examples, this model can capture complex patterns that are difficult for human experts or hard-coded rules to describe and provide a quantified confidence score.

[0049] Specifically, in an embodiment of the present application, the feature engineering vector of the first qualified record is input into a probabilistic classification engine to obtain a probabilistic discovery result, including: performing feature de-redundancy on the feature engineering vector of the first qualified record to obtain a de-redundant feature engineering vector; and inputting the de-redundant feature engineering vector into the probabilistic classification engine to obtain the probabilistic discovery result.

[0050] Accordingly, it should be understood that although the feature engineering vector of the first qualified record is already in numerical form, it may still contain a large amount of redundant information and noise, and its initial representation may not be optimal for the downstream probabilistic inference model. Therefore, the feature engineering vector of the first qualified record is subjected to feature de-redundancy to transform this preliminary, possibly under-refined numerical representation into an enhanced feature vector with higher information density and more precise expression through a deeper, adaptive refinement process. Feature learning is transformed from a fixed transformation sequence to a dynamic polishing process driven by representation convergence. Its purpose is not simply to remove duplicate data, but to iteratively perform nonlinear remapping on the feature engineering vector of the first qualified record and use the redundancy of the feature distribution as a feedback signal to allow the model to dynamically find the optimal representation fixed point for each asset record. This process aims to automatically invest more computing resources in deep refinement of complex asset features in a self-supervised manner until their representation becomes stable, thereby eliminating redundancy at the representation level.

[0051] More specifically, in a specific example of the present application, feature de-redundancy is performed on the feature engineering vector of the first qualified record to obtain a de-redundant feature engineering vector, and the steps are as follows: The feature engineering vector of the first qualified record is subjected to feature refinement to obtain a first qualified record encoding refined feature vector, which is expressed by the formula:

[0052] in, is the feature engineering vector of the first qualified record, is the learnable weight matrix, is the learnable bias vector, is the random perturbation coefficient, Introducing random noise to improve robustness, for activation function, is the feature refinement function, A refined feature vector is encoded for the first qualified record.

[0053] It should be understood that although the feature engineering vector of the first qualified record has completed the conversion from raw data to numerical expression, its representation form is often not optimal for revealing deep and complex security patterns, and may still contain redundant dimensions that are invalid or interfere with downstream tasks. Therefore, the feature engineering vector of the first qualified record is subjected to feature refinement, that is, a deep polishing and nonlinear remapping of the current feature representation. The purpose is to actively enhance the discriminative power of features for downstream probabilistic classification tasks, eliminate noise, and reveal deeper internal patterns through a transformation in a high-dimensional space, thereby producing a first qualified record encoding refined feature vector with higher theoretical information density and better quality. It provides an evaluation object for subsequent feature distribution redundancy calculations and serves as a candidate input for the next round of iteration, thereby driving the entire system to converge one step towards the representation fixed point described in the corpus. It is a basic operation to achieve the relationship between dynamic computing resource allocation and intelligent balance efficiency performance.

[0054] The feature distribution redundancy of the first qualified record encoding refined feature vector relative to the first qualified record feature engineering vector is calculated, and is expressed as follows:

[0055] in, the mutual information between the feature engineering vector of the first qualified record and the encoded refined feature vector of the first qualified record, is the information entropy of the feature engineering vector of the first qualified record, For the first qualified record information assurance degree, is the first qualified record redundancy compression ratio, for Divergence, divide the first qualified record encoding refined feature vector into sub-vectors , Indicates calculation of each group The conditional distribution of With marginal distribution between Divergence, To control the weight of redundant items, Indicates the redundancy of feature distribution.

[0056] It should be understood that in the intelligent analysis process for enterprise asset security management, simply performing a single feature refinement operation is insufficient, as the system cannot determine whether the refinement was effective or whether it has reached its optimization limit. Therefore, the feature distribution redundancy of the first qualified record's encoded refined feature vector relative to its predecessor, the first qualified record's feature engineering vector, is calculated. This aims to accurately measure the information gain or representation change brought about by a single refinement operation through a quantitative mathematical metric. This calculation process converts the degree of change between the new and old feature vectors into a specific, verifiable redundancy value. This value directly reflects whether the feature representation has converged to a fixed point, or in other words, whether a representation equilibrium has been reached. The system obtains a key feedback signal, namely, the feature distribution redundancy. This value materializes the abstract concept of convergence into a concrete control variable. This variable is directly used in the next termination decision, determining whether to stop the iteration and use the current result as the optimal output, or to use the newly generated first qualified record's encoded refined feature vector as input for the next round of refinement for further refinement. This is the key to allocating computing resources on demand and achieving an intelligent balance between efficiency and performance.

[0057] In response to the distance between the feature distribution redundancy and the target redundancy meeting the preset tolerance, the first qualified record encoding refined feature vector is set as the de-redundancy feature engineering vector, which is expressed by the formula:

[0058] in, is the target redundancy, is the tolerance, is the absolute value; If the condition is met, output To remove redundant feature engineering vectors, otherwise As the new first qualified record coding feature, it is refined cyclically.

[0059] It's understandable that without a clear termination condition, the cyclical feature refinement process will either lead to infinite computation or a fixed, suboptimal number of iterations, defeating the purpose of allocating computing resources on demand. Therefore, executing a set operation in response to the distance between the feature distribution redundancy and the target redundancy meeting a preset tolerance provides an automated, data-driven exit for this dynamic refinement process. It uses the feature distribution redundancy calculated in the previous step as a feedback signal, comparing it with a target redundancy representing an ideal convergence state to determine whether the feature representation has converged to the representational fixed point described in the corpus. When the distance between the two is sufficiently small, that is, it meets the preset tolerance, it indicates that further refinement will no longer provide substantial improvement, and it is the optimal time to terminate the loop and lock in the results. This mechanism ensures that for simple asset features, the network converges quickly and saves computational overhead; for complex asset features, sufficient computation cycles are automatically invested until they stabilize. This not only ensures the highest quality and discriminative power of the feature vectors output to the downstream probabilistic classification engine, but also perfectly achieves the elegant and intelligent balance between efficiency and performance described in the corpus.

[0060] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the fusion arbitration sub-unit 1224 is configured to input the intermediate inference result set and the standardized discrete record stream into the fusion arbitrator module to generate the fingerprint-enriched record stream. It should be understood that within the overall enterprise asset security management process, the parallelized logical judgment steps generate two independent information streams: one is the original, standardized discrete record stream containing only objective facts, and the other is the intermediate inference result set containing deterministic and probabilistic security conclusions. These two data streams are logically interrelated but physically separate. Therefore, the intermediate inference result set and the standardized discrete record stream are input into the fusion arbitrator module to merge the newly generated security insights with the original asset facts and resolve any conflicts to form a single, complete, and consistent record. On the one hand, through the fusion operation, the inferred security attributes, such as vulnerabilities and risk classifications, are accurately attached back to their corresponding original asset records, completing the information loop. On the other hand, it uses an arbitration mechanism to select, merge, or discard conclusions from different inference engines (deterministic and probabilistic) according to preset strategies and confidence levels to ensure the accuracy and reliability of the final output results.

[0061] Specifically, in one specific example of this application, the fusion arbitrator module consumes a stream of standardized discrete records and a set of intermediate inference results simultaneously in a streaming manner. When the module receives a standardized record with the ID "Asset-001" and the content "{IP: 1.2.3.4, Port: 443, Service: HTTPS}," it waits for the inference result associated with that ID. Subsequently, it receives two associated inference results: a conclusion {ID: Asset-001, Vulnerability: Heartbleed} from the deterministic engine, and a conclusion {ID: Asset-001, Classification: Misconfigured_Server, Confidence: 0.85} from the probabilistic engine. The arbitrator first associates these three pieces of information by ID. Next, it applies pre-defined arbitration rules, such as: Rule 1: Deterministic results take precedence over probabilistic results; Rule 2: All results are merged when there is no conflict. In this scenario, since the two inference results do not conflict, the arbitrator adopts both and merges them with the original record, ultimately generating and outputting a fingerprint-enriched record: {ID:Asset-001, IP: 1.2.3.4,Port: 443, Service: HTTPS, Vulnerabilities:[Heartbleed], Classifications:[{Type:Misconfigured_Server, Confidence:0.85}]}. This record then enters the fingerprint-enriched record stream.

[0062] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the efficient streaming deduplication unit 123 is configured to perform efficient streaming deduplication on the fingerprint-enriched record stream using a hybrid strategy to obtain a unique asset fragment stream. It should be understood that while the fingerprint-enriched record stream contains complete information, due to the continuous nature of the asset discovery process, the parallel operation of multiple tools, and the dynamic nature of the network environment, this data stream contains a large number of duplicate reports on the same asset generated at different times or by different tools. Therefore, the fingerprint-enriched record stream is subjected to efficient streaming deduplication using a hybrid strategy to eliminate this data-level redundancy, prevent the downstream state aggregation and knowledge graph construction processes from generating erroneous and duplicate asset entities, and avoid unnecessary reprocessing of identical, unchanged asset information. This allows for the precise identification and filtering of redundant records describing the same asset fragment, ensuring that only asset information appearing for the first time or with significant state changes is passed through. By implementing a hybrid strategy that balances efficiency and accuracy, the system aims to determine the novelty of each incoming asset record in real time, thereby establishing a unique view of observed asset fragments.

[0063] Specifically, in a specific example of the present application, an efficient streaming deduplication unit consumes a stream of fingerprint-enriched records. The unit maintains two state structures internally: an in-memory Bloom filter for fast probabilistic judgment, and an external, persistent key-value store (such as Redis or RocksDB) for precise deterministic verification. When a fingerprint-enriched record flows in, the unit first generates a unique identifier based on the core identification fields in the record (for example, a combination of IP, port, and protocol). The identifier is then sent to the Bloom filter for query. If the Bloom filter clearly indicates that the identifier has never appeared before, the record is determined to be a new record, and its identifier is added to the Bloom filter and key-value store, while the record itself is output to the unique asset fragment stream. If the Bloom filter indicates that the identifier may already exist, the second stage of precise verification is initiated, that is, querying the key-value store for the identifier. If the query confirms that the identifier already exists, the record is considered a duplicate and discarded. If the query misses (a false positive in the Bloom filter), it is still treated as a new record, and both state structures are updated and the record is output. This two-stage hybrid strategy leverages the efficiency of the Bloom filter to filter out the vast majority of duplicate data, performing expensive precision checks on only a small amount of potentially duplicate data, thus achieving efficient streaming deduplication.

[0064] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the stateful entity aggregation unit 124 is configured to perform stateful entity aggregation based on an aggregate primary key on each unique asset fragment in the unique asset fragment stream to obtain the unified asset entity stream. It should be understood that, in the grand vision of enterprise asset security management, the unique asset fragment stream obtained through stream deduplication, while addressing data redundancy, is still essentially a collection of isolated facts. For example, one fragment may describe that a certain IP address has port 80 open, another fragment describes that the Apache service is running on the same IP address, and a third fragment indicates that the IP address is associated with a certain domain name. This information logically belongs to the same asset entity, but is discrete within the data stream. Therefore, stateful entity aggregation based on an aggregate primary key is further performed on each unique asset fragment in the unique asset fragment stream to associate and merge these fragmented, atomic asset fragments around a common identifier, thereby constructing a complete, coherent, and multi-dimensional view of the asset entity. By defining one or more aggregated primary keys (such as IP addresses or primary domain names) for each asset fragment, the system can group all fragments sharing the same primary key under the same logical entity. This process is stateful, meaning the system maintains a state for each entity that evolves over time. When new relevant fragments flow in, it updates rather than replaces existing entity information, dynamically accumulating and integrating asset information discovered at different points in time and across different dimensions. The resulting unified asset entity stream provides ideal, structured node data for the subsequent construction of the asset knowledge graph, successfully piecing together fragmented intelligence into clear asset portraits.

[0065] Specifically, in one specific example of this application, a stateful entity aggregation unit consumes a stream of unique asset fragments. The unit first extracts the aggregate primary key for each incoming asset fragment based on pre-defined rules, for example, using the IP address field in the record as the primary key. The unit then performs a partitioning operation based on the primary key, routing all fragments with the same IP address to the same processing instance. This processing instance maintains a state for each IP address, representing an asset entity object under construction. When the first fragment arrives for IP address 10.0.0.1 (e.g., "open port 22"), the processing instance creates a new asset entity object and records the port information. When a second fragment arrives for the same IP address 10.0.0.1 (e.g., "running the SSH service"), the processing instance reads the existing asset entity object state and appends the new service information to it, rather than creating a new object. When a third fragment arrives (e.g., "associated domain host1.example.com"), the unique entity object is also updated. After a trigger condition is met, such as at the end of a time window or when the entity state changes, the processing instance will serialize the current complete asset entity object that aggregates multiple fragment information and output it to the unified asset entity stream.

[0066] In the above-mentioned enterprise Internet asset intelligent discovery and security management system 100, the asset knowledge graph construction and analysis module 130 is used to perform dynamic asset knowledge graph construction and association analysis on the unified asset entity flow to obtain an asset knowledge graph. It should be understood that from the global perspective of enterprise asset security management, although the unified asset entity flow generated in the previous step provides a structured, multi-dimensional asset portrait, these portraits are essentially isolated from each other and lack an explicit expression of the complex relationships between them. Therefore, in the technical solution of the present application, the unified asset entity flow is further subjected to dynamic asset knowledge graph construction and association analysis to place these independent entities in a unified relationship network to reveal and solidify the complex and implicit connections between them, thereby forming a global, networked asset view. It aims to instantiate each asset object in the unified asset entity flow as a node in the graph, and based on the internal attributes of the entity and the common attributes between entities, automatically infer and generate edges representing various dependencies, subordinates or associations through association analysis, thereby dynamically weaving a real-time relationship network depicting the entire enterprise Internet exposure surface.

[0067] Specifically, in one example of this application, this process is implemented using a graph database (such as Neo4j or JanusGraph) as a storage backend, with a stream processing application continuously converting entity stream data into a graph structure. More specifically, the asset knowledge graph construction and analysis module continuously consumes a unified asset entity stream. When this module receives an asset entity object, such as {EntityType: Host, IP: 192.168.1.10, OpenPorts: [80, 443], RunningServices: [Apache / 2.4.54], AssociatedDomains: [test.example.com]}, it performs a series of graph operations. First, it sends a node creation or update instruction to the graph database, for example, using a MERGE statement to ensure the existence of a Host-type node representing the IP address 192.168.1.10 and then updates its properties. Next, the module performs association analysis. It parses the AssociatedDomains field and creates a Domain-type node for the domain name test.example.com (if it doesn't already exist). It then creates an edge between the Host node and the Domain node, representing the RESOLVES_TO relationship. Similarly, it creates a Service node for the service Apache / 2.4.54 and an edge between the Host node and the Service node, representing the HOSTS relationship. As new entity objects continue to flow in, the module continuously creates or updates nodes and edges in the graph database, dynamically and incrementally building and enriching the entire asset knowledge graph.

[0068] In the aforementioned enterprise internet asset intelligent discovery and security management system 100, the risk detection module 140 is used to perform risk detection on the asset knowledge graph based on security policies to obtain risk detection results. It should be understood that in the practice of enterprise asset security management, while a complete and dynamically updated asset knowledge graph provides a global view of asset relationships, it cannot directly reveal the security risks inherent therein. Therefore, risk detection is further performed on the asset knowledge graph based on security policies to obtain risk detection results. By identifying potential vulnerabilities and threats from massive amounts of asset data, a static asset list is transformed into a dynamic risk view. Specifically, by extracting target assets based on security policies, detection resources are focused on the most critical or most risky assets, avoiding blind and inefficient comprehensive scans. Secondly, by distinguishing between technical and non-technical assets and applying different detection methods, detection methods are specialized and precise, ensuring the use of the most appropriate analysis techniques for different risk sources, thereby maximizing the accuracy and depth of risk discovery.

[0069] Specifically, in an embodiment of the present application, the risk detection module is used to: extract a first target asset from the asset knowledge graph based on a security policy; in response to the first target asset being a technical dimension target asset, perform adaptive vulnerability scanning and cloud configuration risk detection on the first target asset to obtain the risk detection result; in response to the first target asset being a non-technical dimension target asset, perform cloud service permission exposure analysis and multi-platform information leakage monitoring on the first target asset to obtain the risk detection result.

[0070] Specifically, the risk detection module first loads a security policy that defines a security audit for all newly discovered server assets hosted on public clouds that expose database ports. The module converts this policy into a graph query, such as a Cypher query, to retrieve nodes from the asset knowledge graph that meet the criteria. The query returns a Host node with the IP address 3.4.5.6, which is marked as a cloud server and includes an open port 3306 as one of its attributes. This node is then designated as the first target asset. The module determines that this Host node is a technical target asset and triggers the appropriate detection process: it invokes the adaptive vulnerability scanning engine to scan the IP address 3.4.5.6 using a vulnerability template specific to MySQL. Simultaneously, it invokes the cloud configuration risk detection engine to query the security group rules associated with the server instance through the API to check for any configuration that opens port 3306 on 0.0.0.0 / 0. Ultimately, the scanning engine reports a known remote code execution vulnerability, and the configuration detection engine reports an overly permissive security group configuration. The module combines these two findings into a single risk detection result and outputs it.

[0071] In the above-mentioned enterprise Internet asset intelligent discovery and security management system 100, the response action generation module 150 is used to generate automated response actions for the risk detection results based on automation rules. It should be understood that the generation of risk detection results is only the identification of the problem. Without subsequent timely disposal, the risk will continue to exist and pose an actual threat. Therefore, automated response actions for the risk detection results are generated based on automation rules to convert security intelligence into executable, standardized disposal plans to bridge the gap between detection and response, thereby shortening the risk exposure window and reducing the manual burden on the security operations team. It aims to accurately and deterministically match and generate the most appropriate response action sequence based on multiple dimensions such as the type and severity of the input risk and the criticality of the associated assets, so as to achieve a transition from passive alerts to active intervention.

[0072] Specifically, the response action generation module begins its work after receiving the risk detection result output by the previous step. Suppose the received risk result is: {RiskID: R-123, Type: Critical_RCE_Vulnerability, Severity: Critical, TargetAsset: {Type: Host, IP: 3.4.5.6, CloudInstanceID: i-012345abcde}}. The rule engine within this module loads its rule base and matches the risk result. The engine discovers a high-priority rule with the condition: IF Risk.Severity is Critical AND Risk.Type contains RCE. This rule is associated with a response playbook that defines two parallel actions. The first action is Isolate, whose template is: {ActionType: Network_Isolate, Target: CloudInstance, Parameters: {InstanceID: {TargetAsset.CloudInstanceID}}}}. The second action is a notification, with the following template: {ActionType:Create_Ticket, System:Jira, Parameters: {Project:SEC_INCIDENT,Priority:Highest,Summary:Critical RCE detected on {TargetAsset.IP}, Assignee:SOC_Team}}. The rule engine then instantiates these two action templates using the specific values ​​in the risk result, generates two specific automated response action instructions, and outputs them. These two instructions can then be subscribed to and executed by the corresponding executors, completing the automated response loop to the risk.

[0073] In summary, the enterprise Internet asset intelligent discovery and security management system according to the embodiment of the present application is explained, which parses high-level asset discovery instructions and converts them into an optimized execution directed graph. This graph can automatically drive the underlying distributed, containerized discovery tool set to work efficiently and collaboratively, thereby replacing the cumbersome and inefficient manual tool chain operations. At the same time, in order to solve the problem of the difficulty in integrating and utilizing multi-source heterogeneous data, the system introduces the original data stream generated by the discovery tool into the real-time data fusion processing layer, and converts it into a unified asset entity stream with unified structure and complete information through steps such as standardization, deduplication and aggregation. Finally, a dynamic asset knowledge graph is constructed based on this high-quality data to achieve accurate risk detection and automated response and disposal, thereby systematically solving the problems of traditional solutions in discovery efficiency and data processing.

[0074] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An enterprise Internet asset intelligent discovery and security management system, characterized by: include: An enterprise Internet asset discovery module, configured to respond to an asset discovery instruction and perform intelligent discovery of enterprise Internet assets based on an optimized execution directed graph to obtain an original asset data stream; An asset entity stream conversion module, configured to convert the original asset data stream into a unified asset entity stream; An asset knowledge graph construction and analysis module, configured to construct a dynamic asset knowledge graph and conduct association analysis on the unified asset entity flow to obtain an asset knowledge graph; A risk detection module, configured to perform risk detection on the asset knowledge graph based on a security policy to obtain a risk detection result; The response action generation module is used to generate an automated response action for the risk detection result based on the automation rules.

2. The enterprise Internet asset intelligent discovery and security management system according to claim 1 is characterized in that: The enterprise Internet asset discovery module includes: An instruction parsing and graphing unit, configured to perform task instruction parsing and execution path graphing on the asset discovery instruction to obtain the optimized execution directed graph; a container resource scheduling unit, configured to execute graph instantiation and container resource scheduling based on the optimized execution directed graph to obtain a set of scheduled tool container instances; a tool container instance set driving unit, configured to drive the scheduled tool container instance set based on the topology of the optimized execution directed graph to obtain original heterogeneous output; A standardization and streaming processing unit is used to perform sidecar proxy-based standardization and streaming on the original heterogeneous output to obtain the original asset data stream.

3. The enterprise Internet asset intelligent discovery and security management system according to claim 2 is characterized in that: The instruction parsing graph unit is used to: Performing syntax parsing and semantic extraction on the asset discovery instruction to obtain semantic task primitives; Based on the tool capability library, generating candidate tool chains for the semantic task primitives; Converting the candidate toolchain into a directed acyclic graph, and calculating a cost value of each graph node in the directed acyclic graph based on the execution constraints in the semantic task primitive to obtain a weighted unoptimized directed acyclic graph; The shortest path solution is performed on the weighted unoptimized directed acyclic graph to obtain the optimized execution directed graph.

4. The enterprise Internet asset intelligent discovery and security management system according to claim 1 is characterized in that: The asset entity flow conversion module includes: The original asset data standardization unit is used to normalize and event the original asset data stream of multi-source heterogeneous data stream to obtain a standardized discrete record stream; a fingerprint enrichment unit, configured to perform multimodal fingerprint enrichment on the standardized discrete record stream to obtain a fingerprint enriched record stream; an efficient streaming deduplication unit, configured to perform efficient streaming deduplication based on a hybrid strategy on the fingerprint enriched record stream to obtain a unique asset segment stream; The stateful entity aggregation unit is configured to perform stateful entity aggregation based on an aggregation primary key on each unique asset fragment in the unique asset fragment stream to obtain the unified asset entity stream.

5. The enterprise Internet asset intelligent discovery and security management system according to claim 4 is characterized in that: The fingerprint enrichment unit comprises: a type conformance response subunit, configured to determine, based on the type field of each standardized discrete record in the standardized discrete record stream, whether each standardized discrete record conforms to a preset fingerprint type, and if so, perform a dynamic append operation on the standardized discrete record to obtain a qualified record stream with a feature vector; a record stream filling subunit, configured to extract atomic features and fill feature vectors from each qualified record with a feature vector in the qualified record stream with a feature vector to obtain a qualified record stream filled with feature vectors; A parallel logic judgment subunit, configured to perform parallel logic judgment based on the feature vector on the qualified record stream filled with the feature vector to obtain an intermediate inference result set; The fusion arbitration subunit is configured to input the intermediate inference result set and the standardized discrete record stream into a fusion arbitrator module to obtain the fingerprint enriched record stream.

6. The enterprise Internet asset intelligent discovery and security management system according to claim 5 is characterized in that: The parallelized logic judgment subunit is used to: extracting the original feature vector of the first qualified record from the qualified record stream populated with feature vectors; Inputting the original feature vector of the first qualified record into a rule-based deterministic inference engine to obtain a deterministic discovery result; Inputting the original feature vector of the first qualified record into a feature engineering preprocessor to obtain a feature engineering vector of the first qualified record; Inputting the feature engineering vector of the first qualified record into a probabilistic classification engine to obtain a probabilistic discovery result; The probabilistic discovery results and the deterministic discovery results are aggregated to obtain an intermediate inference result.

7. The enterprise Internet asset intelligent discovery and security management system according to claim 6 is characterized in that: Inputting the feature engineering vector of the first qualified record into a probabilistic classification engine to obtain a probabilistic discovery result, including: performing feature de-redundancy on the feature engineering vector of the first qualified record to obtain a de-redundant feature engineering vector; The de-redundant feature engineering vector is input into the probabilistic classification engine to obtain the probabilistic discovery result.

8. The enterprise Internet asset intelligent discovery and security management system according to claim 7 is characterized in that: Performing feature de-redundancy on the feature engineering vector of the first qualified record to obtain a de-redundant feature engineering vector, including: performing feature refinement on the feature engineering vector of the first qualified record to obtain a first qualified record encoding refined feature vector; Calculating feature distribution redundancy of the first qualified record encoding refined feature vector relative to the feature engineering vector of the first qualified record; In response to a distance between the feature distribution redundancy and the target redundancy satisfying a preset tolerance, the first qualified record encoding refined feature vector is set as the de-redundancy feature engineering vector.

9. The enterprise Internet asset intelligent discovery and security management system according to claim 1 is characterized in that: The risk detection module is used to: Extracting a first target asset from the asset knowledge graph based on a security policy; In response to the first target asset being a technical dimension target asset, performing adaptive vulnerability scanning and cloud configuration risk detection on the first target asset to obtain the risk detection result; In response to the first target asset being a non-technical dimension target asset, cloud service permission exposure analysis and multi-platform information leakage monitoring are performed on the first target asset to obtain the risk detection result.

Citation Information

Patent Citations

  • Internet situation assessment method based on knowledge graph

    CN117692198A

  • Specific network asset identification method based on network asset atlas construction

    CN117743479A

  • Asset risk tracing method and device

    CN119205351A

  • Asset data fusion and atlas analysis method and system in industrial control network environment

    CN120105332A

  • Network security protection method and system applied to regional digital and intelligent asset business

    CN120281586A

Cited By

  • Intelligent explanation method for inspection result of mart based on AI algorithm

    CN120600198A