Enterprise internet asset intelligent discovery and security management system

By optimizing the collaborative work of the directed graph-driven toolset and real-time data fusion, a dynamic asset knowledge graph is constructed, which solves the problems of tool silos and data inconsistency in enterprise Internet asset management, and achieves efficient and accurate asset discovery and secure management.

CN120707304BActive Publication Date: 2025-12-16HANGZHOU BILING SAFETY TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511142590.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-12-16
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing enterprise internet asset management solutions suffer from tool silos, inconsistent data formats, and low efficiency, making it difficult to achieve efficient and accurate asset discovery and secure management.

Method used

By adopting an enterprise internet asset intelligent discovery and security management system, the underlying toolset is driven by optimized execution of directed graphs to work collaboratively, and real-time data fusion and processing are used to build a dynamic asset knowledge graph, thereby achieving risk detection and automated response.

Benefits of technology

It improves the efficiency and coverage of asset discovery, generates a high-quality unified asset view, supports accurate risk detection and automated response, and solves the problems of inefficiency and complex data processing in traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707304B_ABST
    Figure CN120707304B_ABST
Patent Text Reader

Abstract

The application relates to the field of internet asset management, and specifically discloses an enterprise internet asset intelligent discovery and safety management and protection system, which analyzes and converts high-level asset discovery instructions into an optimized execution directed graph. The graph can automatically drive the bottom layer distributed and containerized discovery tool set to work efficiently in cooperation, thereby replacing the cumbersome and inefficient manual tool chain operation. Meanwhile, in order to solve the problem that multi-source heterogeneous data is difficult to fuse and utilize, the system introduces the original data stream generated by the discovery tool into a real-time data fusion processing layer, and through steps such as standardization, deduplication and aggregation, converts the original data stream into a unified asset entity stream with unified structure and complete information. Finally, based on the high-quality data, a dynamic asset knowledge graph is constructed, precise risk detection and automatic response disposal are realized, and the problems of the traditional scheme in the aspects of discovery efficiency and data processing are solved systematically.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of Internet asset management, and more specifically, to an enterprise Internet asset intelligent discovery and security management system. BACKGROUND

[0002] With the deepening of enterprise digital transformation and the wide application of cloud computing technology, the number, variety and distribution complexity of enterprise assets exposed on the Internet, such as domain names, IP addresses, port services, web applications and even cloud service configurations, have shown explosive growth. This dynamic, heterogeneous and widely distributed asset situation makes it a cornerstone and primary challenge for enterprise network security protection to comprehensively, accurately and timely grasp the digital asset landscape. If the overall situation of the assets cannot be clearly understood, the enterprise will face a huge security blind spot and be unable to effectively assess and manage potential attack surfaces, so that any subsequent security measures will be greatly discounted due to a weak foundation. Therefore, building an enterprise Internet asset solution that can intelligently discover and effectively manage security is an urgent need to protect enterprise network security.

[0003] To address this challenge, the industry has proposed some attack surface management solutions. However, the existing technical means generally have significant defects. On the one hand, most of these solutions rely on the combined use of multiple independent security tools (such as subdomain name scanning, port scanning, vulnerability scanning tools, etc.). Due to the lack of a unified orchestration and scheduling framework, this results in a complex and inefficient data transfer process between different tools when performing asset discovery tasks, and it is difficult to ensure the synchronization and consistency of the state, making the entire discovery process time-consuming and prone to errors. On the other hand, the original asset data generated by these tools varies in format and quality, with a lot of redundant and conflicting information. Traditional processing methods are usually in batch mode, not only is the information time-sensitive, making it difficult to respond to the dynamic changes of modern enterprise assets, but also lacks intelligent data fusion and correlation analysis capabilities, making it difficult to integrate scattered data points into high-quality, structured asset views, making it difficult for security teams to form a precise understanding of the overall situation of the assets.

[0004] Therefore, an optimized enterprise Internet asset intelligent discovery and security management system is desired. SUMMARY

[0005] To solve the above technical problems, the present application is proposed. The embodiments of the present application provide an enterprise Internet asset intelligent discovery and security management system.

[0006] According to one aspect of the present application, an enterprise Internet asset intelligent discovery and security management system is provided, comprising:

[0007] An enterprise Internet asset discovery module is configured to perform intelligent discovery of enterprise Internet assets based on an optimized execution directed graph to obtain an original asset data stream in response to an asset discovery instruction.

[0008] An asset entity stream conversion module is configured to convert the original asset data stream into a unified asset entity stream.

[0009] An asset knowledge graph construction analysis module is configured to perform dynamic asset knowledge graph construction and correlation analysis on the unified asset entity stream to obtain an asset knowledge graph.

[0010] A risk detection module is configured to perform risk detection on the asset knowledge graph based on a security policy to obtain a risk detection result.

[0011] A response action generation module is configured to generate an automated response action for the risk detection result based on an automated rule.

[0012] Compared with the prior art, the enterprise Internet asset intelligent discovery and security management system provided by the present application can analyze and convert high-level asset discovery instructions into an optimized execution directed graph. The graph can automatically drive the bottom-layer distributed and containerized discovery tool set to work efficiently and collaboratively, thereby replacing the cumbersome and inefficient manual tool chain operation. Meanwhile, to solve the problem of difficult fusion and utilization of multi-source heterogeneous data, the system introduces the original data stream generated by the discovery tool into a real-time data fusion processing layer, and converts it into a unified asset entity stream with unified structure and complete information through steps such as standardization, deduplication and aggregation. Finally, a dynamic asset knowledge graph is constructed based on the high-quality data, precise risk detection and automated response disposal are realized, and the problems of the traditional scheme in discovery efficiency and data processing are systematically solved. BRIEF DESCRIPTION OF DRAWINGS

[0013] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application, when taken in conjunction with the accompanying drawings. The drawings provided in the specification and the embodiments of the present application together serve to provide a further understanding that enables others skilled in the art to make or use the present application. The drawings provided are for illustrative purposes and are not intended to limit the present application thereto, and constitute part of the specification. In the drawings, the same reference numerals generally refer to the same components or steps throughout the drawings.

[0014] Figure 1 FIG. 1 is a system block diagram of an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application.

[0015] Figure 2 FIG. 2 is a data flow schematic diagram of an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application.

[0016] Figure 3A block diagram of an enterprise Internet asset discovery module in an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application.

[0017] Figure 4 A block diagram of an asset entity flow conversion module in an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application.

[0018] Figure 5 A block diagram of a fingerprint enrichment unit in an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application. DETAILED DESCRIPTION

[0019] Hereinafter, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. It should be understood that the exemplary embodiments described herein are merely some of the embodiments of the present application, and the present application is not limited to the exemplary embodiments described herein.

[0020] As shown in the present application and claims, unless the context clearly indicates otherwise, the words "one", "an", "a", and / or "the" do not mean "only one", but are used interchangeably with "at least one", "one or more" or "one or more of possibly a plurality of". Generally, the terms "comprise", "comprising", "include", "including" and the like are intended to mean including but not limited to, such that an enumerated step or element is not meant to be an exclusive list.

[0021] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or server. The modules are merely illustrative, and different aspects of the system and method can use different modules.

[0022] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in the exact order. Instead, various steps can be processed in reverse order or at the same time, as desired. Other operations can also be added to or removed from these processes, or one or more steps can be removed from these processes.

[0023] Hereinafter, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. It should be understood that the exemplary embodiments described herein are merely some of the embodiments of the present application, and the present application is not limited to the exemplary embodiments described herein.

[0024] The prior art has the core technical problems of rigid tool process in enterprise Internet asset management, and low efficiency of multi-source heterogeneous data processing process. In order to solve the above problems, an enterprise Internet asset intelligent discovery and security management system is proposed in the technical scheme of the present application. The system starts with a high-level asset discovery instruction. The system first intelligently analyzes and converts it into an optimized execution directed graph. The graph can automatically arrange and schedule a series of independent containerized discovery tools at the bottom layer, realize the optimal path collaborative execution of the task, and completely break the deadlock of tool island. When the tool set is executed, the multi-source heterogeneous raw data stream generated will not be processed in isolation, but will be immediately sent to a real-time data processing pipeline. In the pipeline, the data stream is standardized, fingerprinted, efficiently deduplicated, and finally aggregated into a state entity, converting the scattered and chaotic raw data into a unified structure and complete information asset entity stream. Finally, these high-quality asset entities are used to dynamically build a global asset knowledge graph, providing a solid and reliable data foundation for subsequent precise risk detection and automated security management, thereby realizing the full-process automation and intelligentization from instruction to response.

[0025] In the technical scheme of the present application, an enterprise Internet asset intelligent discovery and security management system is proposed. Figure 1 The system block diagram of the enterprise Internet asset intelligent discovery and security management system according to the embodiment of the present application. Figure 2 The data flow diagram of the enterprise Internet asset intelligent discovery and security management system according to the embodiment of the present application. As shown in Figure 1 and Figure 2 The enterprise Internet asset intelligent discovery and security management system 100 according to the embodiment of the present application, as shown in the figures, includes: an enterprise Internet asset discovery module 110, configured to respond to an asset discovery instruction, perform enterprise Internet asset intelligent discovery based on an optimized execution directed graph to obtain a raw asset data stream; an asset entity stream conversion module 120, configured to convert the raw asset data stream into a unified asset entity stream; an asset knowledge graph construction and analysis module 130, configured to perform dynamic asset knowledge graph construction and correlation analysis on the unified asset entity stream to obtain an asset knowledge graph; a risk detection module 140, configured to perform risk detection on the asset knowledge graph based on a security policy to obtain a risk detection result; and a response action generation module 150, configured to generate an automated response action for the risk detection result based on an automated rule.

[0026] In the enterprise Internet asset intelligent discovery and security management system 100, the enterprise Internet asset discovery module 110 is configured to perform intelligent discovery of enterprise Internet assets based on an optimized execution directed graph to obtain an original asset data stream in response to an asset discovery instruction. It should be understood that the prior art relies heavily on manual operation of various functional single scanning tools by security personnel, that is, when the security team needs to conduct a comprehensive asset inventory of a certain domain name, the traditional way is to manually run a plurality of tools such as subdomain enumeration, port scanning, and service identification one by one. This way is not only inefficient and rigid, but also prone to omissions due to operational errors. Moreover, the tools are isolated from each other, forming tool islands that are difficult to coordinate, resulting in a time-consuming discovery process and easy to miss. Therefore, in the practice of enterprise asset security management, intelligent discovery of enterprise Internet assets based on an optimized execution directed graph is performed in response to an asset discovery instruction to obtain an original asset data stream. By analyzing and converting the user's asset discovery requirements into a computable and optimized execution directed graph, the underlying diversified discovery tools can be automatically and intelligently arranged and scheduled. In this way, the system can intelligently drive each tool container to work in parallel or series according to the graph. In this way, the overall efficiency and coverage of asset discovery can be greatly improved, and a continuous original asset data stream containing all original discovery results can be automatically generated, laying a foundation for efficient and reliable subsequent unified processing and analysis.

[0027] Figure 3 A block diagram of an enterprise Internet asset discovery module in an enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, in an embodiment of the present application, the enterprise Internet asset discovery module 110 includes an instruction analysis and graphing unit 111, a container resource scheduling unit 112, a tool container instance set driving unit 113, and a standardization and streaming processing unit 114. Figure 3 The instruction analysis and graphing unit 111 is configured to perform task instruction analysis and execution path graphing on the asset discovery instruction to obtain the optimized execution directed graph. The container resource scheduling unit 112 is configured to perform graph instantiation and container resource scheduling based on the optimized execution directed graph to obtain a set of scheduled tool container instances. The tool container instance set driving unit 113 is configured to drive the set of scheduled tool container instances based on the topology of the optimized execution directed graph to obtain original heterogeneous output. The standardization and streaming processing unit 114 is configured to perform standardization and streaming based on a Sidecar agent on the original heterogeneous output to obtain the original asset data stream.

[0028] In the enterprise Internet asset intelligent discovery and security management system 100, the instruction analysis and mapping unit 111 is configured to perform task instruction analysis and execution path mapping on the asset discovery instruction to obtain the optimized execution directed graph. It should be understood that the traditional asset discovery process highly depends on the personal experience of security engineers, and various tools are combined and called through manual or temporary script writing. This method is not only inefficient and reproducible, but also cannot guarantee the global optimality of the execution path. Therefore, in the actual application scenario of enterprise asset discovery, the task instruction analysis and execution path mapping on the asset discovery instruction to obtain the optimized execution directed graph is to systematically convert the fuzzy human instruction system into an accurate, quantitative and machine understandable and executable optimal task planning blueprint, thereby realizing the automation, standardization and intelligentization of the asset discovery process. In this way, a structured optimized execution directed graph is generated, which clearly defines the required tools, execution order and dependency relationship for completing the specified discovery task, and provides a deterministic and efficient execution basis for subsequent container resource scheduling and task driving.

[0029] Specifically, the instruction analysis and mapping unit is configured to perform syntax analysis and semantic extraction on the asset discovery instruction to obtain semantic task primitives, generate a candidate tool chain for the semantic task primitives based on a tool capability library, convert the candidate tool chain into a directed acyclic graph, and calculate the cost value of each graph node in the directed acyclic graph based on the execution constraints in the semantic task primitives to obtain a weighted unoptimized directed acyclic graph. The shortest path of the weighted unoptimized directed acyclic graph is solved to obtain the optimized execution directed graph.

[0030] That is, the implementation of the process involves a multi-stage conversion process. More specifically, when the system receives an asset discovery instruction such as "perform a comprehensive Web asset probe on the target company", first, the system performs syntax analysis and semantic extraction on the instruction, using natural language processing or a pre-defined domain-specific language parser to decompose it into a set of structured semantic task primitives, such as {target: target company, task type: Web asset probe, scope: comprehensive}. Then, based on a pre-set tool capability library, the system queries tools or tool combinations that can meet these task primitives, generating multiple candidate tool chains, such as a tool chain that may be "subdomain enumeration tool -> port scanning tool -> Web service identification tool", and another that may be "passive domain name collection tool -> Web fingerprint identification tool". Subsequently, the system integrates and converts these candidate tool chains into a directed acyclic graph, where each tool is a graph node, and data flow constitutes a directed edge; at the same time, based on the execution constraints in the semantic task primitives (such as time limit, resource consumption preference, etc.), the system calculates the cost value of each node in the graph, for example, tool nodes with long execution time or high resource consumption are assigned higher costs, thus obtaining a weighted unoptimized directed acyclic graph. Finally, the system applies a shortest path solving algorithm such as Dijkstra on the weighted graph, calculates the path with the lowest total cost from the starting node to the end node, and this path is the final output of the optimized execution directed graph.

[0031] In the above enterprise Internet asset intelligent discovery and security management system 100, the container resource scheduling unit 112 is configured to perform graph instantiation and container resource scheduling based on the optimized execution directed graph to obtain a set of scheduled tool container instances. It should be understood that in actual enterprise security operation scenarios, the optimized execution directed graph generated in the previous step is only a static, logical-level execution plan, which cannot execute any operation itself. Therefore, based on the optimized execution directed graph, graph instantiation and container resource scheduling are performed to obtain a set of scheduled tool container instances, so as to convert this abstract logical blueprint into a physical world executable software entity. In this way, according to the graph planning, the necessary computing resources can be allocated to each tool task to be executed in the computing cluster, and the corresponding tool program is started. The system transitions from a purely planning stage to a ready-to-execute stage, and the output is a set of tool container instances that have been started and allocated resources, which are ready to be executed at any time, thereby ensuring the efficiency and reliability of subsequent task-driven execution and avoiding execution failures due to insufficient resources or environmental problems.

[0032] Specifically, in one specific example of the present application, the implementation of the process relies on a modern container orchestration platform such as Kubernetes. More specifically, when the container resource scheduling unit of the system receives the optimized execution directed graph, it will first traverse the nodes in the graph, parse out the tool represented by each node (for example, the subdomain enumeration tool subfinder), its running parameters, and the preset resource requirements (such as CPU, memory). Then, for each tool node, the scheduling unit will dynamically generate a deployment description file that can be recognized by the container orchestration system, such as a YAML manifest of a Kubernetes Pod or Job. The manifest will explicitly specify the required tool container image to be pulled, the start command, the parameters to be passed, and the resource request and limit. Subsequently, the scheduling unit will submit this description file to the control center of the container orchestration platform through the API. The scheduler of the orchestration platform will intelligently select the most suitable physical or virtual node to deploy the container instance of the tool according to the real-time load and available resources of each node in the cluster. Once the container instance is successfully started and enters the running state on the specified node, it will be registered and added to a dynamic collection until all the tools that need to be initially started in the graph have been instantiated, and finally form the scheduled tool container instance set, fully preparing for the next step of driving execution.

[0033] In the above enterprise Internet asset intelligent discovery and security management system 100, the tool container instance set driving unit 113 is configured to drive the scheduled tool container instance set based on the topology of the optimized execution directed graph to obtain the original heterogeneous output. It should be understood that in the process of enterprise security asset discovery, it is not enough to just instantiate the tool and allocate resources. These tool instances are in standby state and will not automatically execute. Therefore, a coordinator is needed to accurately command and trigger these independent tool units according to the preset operation plan. In the technical solution of the present application, the scheduled tool container instance set is driven based on the topology of the optimized execution directed graph to obtain the original heterogeneous output. This realizes the transformation from static planning to dynamic execution, that is, through a central driving mechanism, the tool dependency relationship and data flow defined by the optimized execution directed graph are strictly followed to arrange the actual operation of the entire discovery task. The entire asset discovery task is efficiently and orderly completed, and a series of original discovery results generated by different tools with different formats and contents, i.e., original heterogeneous output, are obtained, providing the most basic raw materials for the subsequent data standardization and fusion stage.

[0034] Specifically, in one specific example of the present application, the optimized directed graph is first topologically sorted, and all start nodes with in-degree zero are identified, which do not depend on the output of any other tool. Then, the parameters (such as the main domain name) in the initial asset discovery instruction are taken as input, triggering the start of execution of the tool container instances (such as the subdomain enumeration tool) corresponding to these start nodes. When a tool container instance completes its task, the driving unit captures the output data generated thereby. At the same time, it determines all successor nodes directly dependent on the output of the node according to the topological structure of the graph. Then, the captured output data is taken as input and passed to the tool container instances (such as passing the list of subdomain names to the port scanning tool) corresponding to these successor nodes, and they are instructed to start execution. This process is iterated along the edges of the graph until all nodes are executed. In this process, the outputs generated by all tool container instances are collected together to form the original heterogeneous output.

[0035] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the standardization and fluidization processing unit 114 is configured to perform standardization and fluidization on the original heterogeneous output based on a Sidecar agent to obtain the original asset data stream. It should be understood that in the complex scenario of enterprise asset discovery, the original heterogeneous output generated by different tools has different formats and structures, for example, the subdomain tool outputs a pure text list, while the vulnerability scanning tool may output a complex XML or JSON report. Such inconsistency in the data source makes subsequent unified processing and analysis extremely difficult. Therefore, the original heterogeneous output is further standardized and fluidized based on a Sidecar agent to solve the problem of data heterogeneity in real time. In this way, the responsibility for data format conversion and streaming can be separated from the core discovery tool, and an accompanying agent can be used to convert all scattered and batched original outputs into a unified and continuous data stream in real time. The system integrates the originally chaotic and non-real-time multi-source data into a consistent and easy-to-consume original asset data stream, which not only greatly simplifies the design complexity of the subsequent data processing module, but also lays the foundation for the real-time response capability of the entire system.

[0036] Specifically, in one specific example of the present application, the process is implemented by equipping each tool container instance with a Sidecar proxy container. The Sidecar proxy is deployed within the same execution unit (e.g. a Kubernetes Pod) as the main tool container and is configured to capture the standard output stream of the main tool container or monitor its output files. When the main tool container, e.g. a port scanning tool, produces a line of output indicating an open IP and port, the Sidecar proxy immediately intercepts this data. The proxy has built-in parsing logic for different tool output formats, which parses and converts the raw text line into a predefined standardized data structure, e.g. a JSON object containing fields such as asset type, IP address, port, protocol and timestamp, depending on the type of the current tool. Upon completing the standardization conversion, instead of caching the data, the Sidecar proxy immediately publishes the JSON object as a message to a specific topic of a centralized message queue system (e.g. Apache Kafka). All tool Sidecar proxies continuously push messages to this queue, thereby converging into a unified, real-time raw asset data stream.

[0037] In the above enterprise internet asset intelligent discovery and security management system 100, the asset entity stream conversion module 120 is configured to convert the raw asset data stream into a unified asset entity stream. It should be understood that although the raw asset data stream is uniform in format, it is essentially discrete and instantaneous data points, such as a certain IP opening a certain port at a certain time. These data points have problems such as a large amount of redundancy (e.g. multiple scans discovering the same port), incomplete information (e.g. only IP and port, lacking service details) and lack of correlation. Therefore, the raw asset data stream is converted into a unified asset entity stream to aggregate these scattered, event-centric data points into an asset-centric, stable-state and information-rich entity view. Through a series of deep data processing and fusion operations, the raw data stream is purified, enriched and aggregated, thereby constructing a complete portrait that can comprehensively and accurately describe each digital asset of the enterprise. This includes removing redundant information, supplementing detailed fingerprint features of the asset (e.g. running services, used frameworks), and associating and aggregating different fragments (e.g. domain name, IP, port, application) describing the same asset to form a logically complete asset entity.

[0038] Figure 4 A block diagram of the asset entity stream conversion module in the enterprise internet asset intelligent discovery and security management system according to an embodiment of the present application. As shown in FIG. 2, the asset entity stream conversion module 120 is configured to receive the raw asset data stream from the asset data stream collection module 110 and convert it into a unified asset entity stream. The asset entity stream conversion module 120 is configured to perform a series of data processing and fusion operations on the raw asset data stream, including but not limited to data cleaning, data enrichment and data aggregation, to construct a complete portrait of each digital asset of the enterprise. Figure 4As shown, in the embodiments of the present application, the asset entity stream conversion module 120 includes: an original asset data standardization unit 121, configured to normalize and eventize the original asset data stream to obtain a standardized discrete record stream; a fingerprint enrichment unit 122, configured to enrich the standardized discrete record stream in multiple modalities to obtain a fingerprint enriched record stream; an efficient stream deduplication unit 123, configured to perform efficient stream deduplication based on a hybrid strategy on the fingerprint enriched record stream to obtain a unique asset segment stream; and a stateful entity aggregation unit 124, configured to perform stateful entity aggregation based on an aggregation primary key on each unique asset segment in the unique asset segment stream to obtain the unified asset entity stream.

[0039] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the original asset data standardization unit 121 is configured to normalize and eventize the original asset data stream to obtain a standardized discrete record stream. It can be understood that, in the actual process of enterprise asset discovery, although the original asset data stream generated by the Sidecar agent tends to be consistent in transmission format, the semantic level of the data content is still heterogeneous, and the output fields and meanings of different tools are different, which cannot be directly analyzed and processed. Therefore, the original asset data stream is further normalized and eventized to eliminate the semantic gap caused by the diversity of tools. The core purpose of performing this step is to establish a globally unified and standardized data model, and to convert each piece of original discovery information into an independent and standardized event record with context metadata, thereby providing a homogeneous data basis for all subsequent processing links. The system converts a mixed and semantically inconsistent data stream into a clear and regular standardized discrete record stream, each record of which follows the same data structure, greatly reducing the complexity of subsequent data fusion and analysis.

[0040] More specifically, in one specific example of the present application, a stream processing application (e.g. built based on Apache Flink or Kafka Streams) subscribes and consumes the upstream raw asset data stream. When the application receives a raw data record from the message queue, it first deserializes it. Then, the application loads the corresponding parsing and mapping rules from a pre-configured rule base according to the source identification (e.g. tool name injected by Sidecar) carried in the record. Subsequently, the application maps the fields in the raw record (e.g. host field output by a certain tool) to the corresponding fields in the standard data model (e.g. asset.value) according to the rules, and normalizes the data types and formats. In this process, the application also generates a unique event ID for the record, and attaches metadata such as processing timestamp, completing the eventization operation. Finally, the new record that fully complies with the standard data model, is normalized and eventized, is serialized and published to a new message queue topic, forming the standardized discrete record stream for downstream modules to consume.

[0041] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the fingerprint enrichment unit 122 is configured to perform multi-modal fingerprint enrichment on the standardized discrete record stream to obtain a fingerprint enriched record stream. It should be understood that in the context of enterprise asset security management, although the standardized discrete record stream after normalization has a uniform structure, the information it contains is often basic and superficial, such as only knowing that an IP address has port 80 open. This information is not deep enough to support accurate risk assessment and attack surface analysis. Therefore, multi-modal fingerprint enrichment is performed on the standardized discrete record stream to deeply mine and supplement the basic asset information, so as to reveal its deeper technical attributes and characteristics. Through active detection or passive analysis, detailed technical fingerprint information is attached to each basic asset record. This includes identifying the specific type and version of the web server (e.g. Nginx / 1.21.6), the framework and language of the backend application (e.g. Spring Boot, PHP), the components and their versions used (e.g. jQuery / 3.5.1), and even the operating system type. This multi-modal fingerprint identification covers multiple dimensions from the network layer to the application layer, aiming to build a stereoscopic and detailed asset technical portrait.

[0042] Figure 5 A block diagram of a type compliance response subunit in the enterprise Internet asset intelligent discovery and security management system according to an embodiment of the present application is shown in FIG. 6. As shown in FIG. 6, the type compliance response subunit 1220 is configured to receive the standardized discrete record stream from the normalization unit 1210, and perform multi-modal fingerprint enrichment on the standardized discrete record stream to obtain a fingerprint enriched record stream. Figure 5As shown, in the embodiments of the present application, the fingerprint enrichment unit 122 comprises: a type compliance response subunit 1221 configured to determine whether each standardized discrete record in the standardized discrete record stream complies with a preset fingerprinting type based on a type field of each standardized discrete record, and if so, perform a dynamic additional operation on the standardized discrete record to obtain a qualified record stream with feature vectors; a record stream filling subunit 1222 configured to perform atomic feature extraction and feature vector filling on each qualified record with feature vectors in the qualified record stream with feature vectors to obtain a qualified record stream with filled feature vectors; a parallelization logical judgment subunit 1223 configured to perform feature vector-based parallelization logical judgment on the qualified record stream with filled feature vectors to obtain an intermediate inference result set; and a fusion arbitration subunit 1224 configured to input the intermediate inference result set and the standardized discrete record stream into a fusion arbitrator module to obtain the fingerprint enriched record stream.

[0043] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the type compliance response subunit 1221 is configured to determine whether each standardized discrete record in the standardized discrete record stream complies with a preset fingerprinting type based on a type field of each standardized discrete record, and if so, perform a dynamic additional operation on the standardized discrete record to obtain a qualified record stream with feature vectors. It should be understood that in the practice of enterprise asset security management, although the standardized discrete record stream solves the problem of inconsistent data structures, it contains a large number of records of different types, such as domain names, IP addresses, open ports, etc., and not all records are suitable or need to be subjected to deep technical fingerprint detection. Therefore, based on the type field of each standardized discrete record in the standardized discrete record stream, it is determined whether it complies with a preset fingerprinting type and a dynamic additional operation is performed, so as to avoid invalid and resource-consuming detection operations on asset records of non-target types. In this way, classified processing and targeted enrichment of asset records are achieved. The system first discriminates the incoming records, and only identifies those records with detectable service features (such as port records that have opened Web service or database service), then calls the corresponding fingerprint identification module to perform deep information mining on these qualified records, and attaches the mined structured feature information back to the original records.

[0044] In particular, in one specific example of the present application, one stream processing node continuously consumes a stream of normalized discrete records. When the node receives a normalized discrete record, it first checks the type field in the record. For example, if a record has a type of open port and a service field of HTTP, the record matches a pre-defined web fingerprinting type. Subsequently, the node triggers a dynamic enrichment operation: it passes the IP and port information in the record to a web fingerprinting submodule. The submodule initiates an HTTP request to the target, analyzes the response header, response body content, and icon hash, and identifies the web server as Nginx and the backend framework as Django using a pre-defined fingerprint library. Next, the node constructs a structured feature vector, such as a JSON object {server: Nginx, framework: Django}, from the identified information and appends the feature vector to the original record, forming a new record with a feature vector. The enriched record is output to the downstream as part of the qualified record stream; while records that do not match the pre-defined fingerprinting type (e.g., records with a type of domain name) are directly bypassed or sent to other processing logic.

[0045] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the record stream filling subunit 1222 is configured to perform atomic feature extraction and feature vector filling on each qualified record with a feature vector in the qualified record stream with a feature vector to obtain a qualified record stream with filled feature vectors. It should be understood that the appended feature vector of the qualified record stream with a feature vector is often raw and unstructured, such as a complete HTTP response header string or a piece of HTML code. Although such raw data is rich in information, it cannot be directly used for accurate comparison and aggregation. Therefore, in the technical solution of the present application, atomic feature extraction and feature vector filling are further performed on each record in the qualified record stream with a feature vector to decompose these rough, semi-structured feature information into the smallest, indivisible, and clearly business-meaning atomic features, thereby realizing deep analysis and normalization of asset technical features. The system extracts the mixed information in the original feature vector, such as the component name and version number from a Server response header, and fills these extracted atomic features into a pre-defined standard feature vector template with a uniform structure. This process ensures that the final representation is consistent and regular regardless of the changes in the original features.

[0046] In particular, in one specific example of the present application, one stream processing node consumes a stream of qualified records with feature vectors. When the node receives a record whose original feature vector contains a HTTP response header Server: Apache / 2.4.54 (Ubuntu), the node first applies a series of pre-set regular expressions or parsers to the string for atomic feature extraction. It successfully extracts three atomic features: component name Apache, version number 2.4.54, and operating system Ubuntu. Then, the node takes a standard feature vector template, which can be a JSON structure defining multiple fields such as component_name, component_version, os_type, etc. Subsequently, the node fills the extracted atomic features into the corresponding fields of the template to form a filled feature vector {component_name: Apache, component_version: 2.4.54, os_type: Ubuntu,...}. Finally, the record carrying the filled, structured feature vector is outputted into the stream of qualified records with filled feature vectors.

[0047] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the parallelization logic judgment subunit 1223 is configured to perform feature vector-based parallelization logic judgment on the stream of qualified records with filled feature vectors to obtain a set of intermediate inference results. It should be understood that in the overall framework of enterprise asset security management, the record stream containing standardized feature vectors has been generated, which accurately describes the technical composition of the assets, but they are still isolated factual statements and lack qualitative or quantitative judgments in the security aspect. Therefore, the feature vector-based parallelization logic judgment is performed on the stream of qualified records with filled feature vectors to automatically collide and associate these purely technical facts with a vast knowledge base of security, thereby revealing their potential security implications. By efficiently performing parallel matching and logical reasoning between the standardized feature vector of each record and the pre-set rule base (such as vulnerability library, compliance baseline, threat intelligence), the system can automatically generate preliminary, atomic security conclusions for each asset record, such as determining whether a certain component version has known vulnerabilities or whether a certain configuration violates security policies.

[0048] Specifically, in the embodiments of the present application, the parallelization logic judgment subunit is configured to: extract a first qualified record's original feature vector from the qualified record stream of the filled feature vector; input the first qualified record's original feature vector into a rule-based deterministic inference engine to obtain a deterministic discovery result; input the first qualified record's original feature vector into a feature engineering preprocessor to obtain a first qualified record's feature engineering vector; input the first qualified record's feature engineering vector into a probabilistic classification engine to obtain a probabilistic discovery result; and perform result aggregation on the probabilistic discovery result and the deterministic discovery result to obtain an intermediate inference result.

[0049] More specifically, the first qualified record's original feature vector is extracted from the qualified record stream of the filled feature vector, and the first qualified record's original feature vector is input into a rule-based deterministic inference engine to obtain a deterministic discovery result. It can be understood that the qualified record stream of the filled feature vector accurately describes the technical composition of the asset, but these information itself is only a neutral technical fact, lacking direct security value judgment. Therefore, the records are extracted from the qualified record stream of the filled feature vector and input into a rule-based deterministic inference engine to associate these objective technical attributes with explicit security knowledge to obtain unambiguous security conclusions. By comparing the standardized feature vector of each asset record with a rigid rule base composed of expert knowledge, vulnerability information and compliance baseline, the system aims to quickly identify those unambiguous security problems that meet the deterministic conditions, such as a certain component version that explicitly exists a publicly known serious vulnerability.

[0050] In a specific example of the present application, an inference processing unit continuously consumes the qualified record stream of the filled feature vector. When the unit receives a record whose filled feature vector is {component_name: Log4j, component_version: 2.14.1}, the unit inputs this feature vector as a fact into a rule-based deterministic inference engine. The engine is preloaded with a series of rules inside, one of which may be defined as: when the component name is Log4j and its version number is greater than or equal to 2.0 and less than 2.15.0, a high-risk vulnerability conclusion is triggered. When the engine is executed, it will match the input feature vector with the conditions of this rule and find that it fully meets the conditions. Therefore, the engine will generate a deterministic discovery result, such as a structured data containing vulnerability number CVE-2021-44228, risk level as serious, etc. This discovery result is then attached to the record or output as an independent inference event, which together constitutes part of the intermediate inference result set.

[0051] More specifically, the original feature vector of the first qualified record is input into a feature engineering preprocessor to obtain a feature engineering vector of the first qualified record. It should be understood that in the scenario of enterprise asset security management, although the rule-based deterministic inference engine can efficiently handle security problems with clear features, it is not competent for potential risks with complex patterns, fuzzy boundaries, and dependence on multiple factor correlations. Therefore, the original feature vector of the first qualified record is input into the feature engineering preprocessor, and through a series of mathematical transformations and encoding on the original feature vector, it is converted into a high-dimensional, standardized, and pure numerical feature engineering vector. This process aims to maximize the preservation of original information while expressing it in a paradigmatic form friendly to machine learning algorithms to facilitate subsequent pattern recognition and probability inference of the model.

[0052] In a specific example of the present application, the feature engineering preprocessor receives a qualified record, and the original feature vector of the qualified record is {component_name: Nginx, component_version: 1.21.6, server_header: nginx / 1.21.6}. The preprocessor first performs One-Hot Encoding on the categorical feature component_name. Assuming that the pre-defined component vocabulary has Nginx, Apache, and IIS, Nginx is converted to the vector [1, 0, 0]. Then, for the numerical feature component_version, the string 1.21.6 is parsed into three independent numerical features [1, 21, 6], and normalization processing can be performed. For the text feature server_header, the preprocessor can apply the TF-IDF algorithm or a pre-trained word embedding model (such as BERT) to convert it into a fixed-dimensional numerical vector. Finally, the preprocessor concatenates these processed numerical vectors to form a single, high-dimensional, and pure numerical feature engineering vector, which is the feature engineering vector of the first qualified record and is ready to be input into the subsequent machine learning model.

[0053] More specifically, the feature engineering vector of the first qualified record is input into a probabilistic classification engine to obtain a probabilistic discovery result. It should be understood that in the complex scenario of enterprise asset security management, the deterministic inference engine can effectively identify security problems that have clear rules, but it is difficult to identify and assess potential risks that are composed of multiple weak features, have complex patterns, and have no deterministic rules. Therefore, the feature engineering vector of the first qualified record is input into a probabilistic classification engine to utilize the ability of machine learning to identify and assess security risks hidden in deep data correlations. A pre-trained classification model is used to perform deep analysis on the feature engineering vector of the asset, thereby probabilistically predicting whether the asset belongs to a specific risk category (such as a tampered server or an improperly configured development environment). The model can capture complex patterns that are difficult for human experts or hard-coded rules to describe by learning the features of a large number of known positive and negative samples, and can give a quantitative confidence.

[0054] Specifically, in the embodiment of the present application, the feature engineering vector of the first qualified record is input into a probabilistic classification engine to obtain a probabilistic discovery result, including: performing feature redundancy reduction on the feature engineering vector of the first qualified record to obtain a redundancy-reduced feature engineering vector; and inputting the redundancy-reduced feature engineering vector into the probabilistic classification engine to obtain the probabilistic discovery result.

[0055] Accordingly, it should be understood that the feature engineering vector of the first qualified record, although in numerical form, may still have a large amount of redundant information and noise inside, and its initial representation may not be optimal for the downstream probabilistic inference model. Therefore, the feature engineering vector of the first qualified record is subjected to feature redundancy reduction, so that this preliminary, possibly insufficiently refined numerical representation is converted into an enhanced feature vector with higher information density and more accurate expression through a more in-depth, adaptive refinement process. Feature learning is innovated from a fixed transformation sequence to a dynamic polishing process driven by representation convergence. The purpose is not simply to remove duplicate data, but to nonlinearly remap the feature engineering vector of the first qualified record through iterative loops, and use feature distribution redundancy as a feedback signal to dynamically find the best representation fixed point for each asset record by the model. This process aims to automatically invest more computational resources in deep refinement of complex asset features through self-supervised means until their representation tends to be stable, thereby eliminating redundancy at the representation level.

[0056] More specifically, in one specific example of the present application, the feature engineering vector of the first qualified record is subjected to feature redundancy reduction to obtain a redundancy-reduced feature engineering vector, with the following steps:

[0057] characteristic engineering vector of the first qualified record is subjected to feature refinement to obtain a first qualified record coding refined feature vector, which is expressed by a formula:

[0058]

[0059] wherein, is a characteristic engineering vector of the first qualified record, is a learnable weight matrix, is a learnable bias vector, is a random disturbance coefficient, random noise is introduced to improve robustness, is an activation function, is a feature refinement function, is a first qualified record coding refined feature vector.

[0060] It can be understood that the characteristic engineering vector of the first qualified record has completed the conversion from the original data to the numerical expression, but its representation form is often not optimal for revealing deep and complex security patterns, and may still contain redundant dimensions that are invalid or interfere with downstream tasks. Therefore, the characteristic engineering vector of the first qualified record is subjected to feature refinement, that is, the current feature representation is once deeply polished and nonlinearly remapped. It aims to actively enhance the discriminability of the feature for the downstream probabilistic classification task through a transformation in a high-dimensional space, eliminate noise, and reveal deeper internal patterns, thereby outputting a first qualified record coding refined feature vector that is theoretically higher in information density and better in quality. It provides an evaluation object for subsequent feature distribution redundancy calculation and serves as a candidate input for the next round of iteration, thereby driving the entire system to converge one step closer to the representation fixed point described in the corpus, and is a basic operation for realizing the relationship between dynamic calculation resource allocation and intelligent balance efficiency performance.

[0061] The feature distribution redundancy of the first qualified record coding refined feature vector with respect to the characteristic engineering vector of the first qualified record is calculated, which is expressed by a formula:

[0062]

[0063] wherein, is mutual information between the characteristic engineering vector of the first qualified record and the first qualified record coding refined feature vector, is information entropy of the characteristic engineering vector of the first qualified record, is first qualified record information guarantee degree, is first qualified record redundancy compression rate, is divergence, the first qualified record coding refined feature vector is divided into Subvectors , This indicates that each group is calculated. Conditional distribution With marginal distribution Between divergence, To control the weight of redundant terms, This indicates the redundancy of the characteristic distribution.

[0064] It is understandable that in the intelligent analysis process of enterprise asset security management, performing only one feature refinement operation is insufficient, as the system cannot determine whether the refinement is effective or has reached its optimization limit. Therefore, further calculating the feature distribution redundancy of the first qualified record code refined feature vector relative to its predecessor, i.e., the feature engineering vector of the first qualified record, aims to accurately measure the information gain or representational change brought about by a single refinement operation through a quantitative mathematical index. This calculation process transforms the degree of change between the old and new feature vectors into a specific, quantifiable redundancy value. This value directly reflects whether the feature representation has converged to a fixed point, or in other words, whether it has reached a representational equilibrium state. The system obtains a crucial feedback signal: feature distribution redundancy. This value materializes an abstract concept of convergence into a concrete control variable. This variable will be directly used for the next termination judgment, determining whether to stop iteration and use the current result as the optimal output, or to use the newly generated first qualified record code refined feature vector as the input for the next round of refinement for deeper polishing. This is the key to achieving on-demand allocation of computing resources and striking an intelligent balance between efficiency and performance.

[0065] In response to the distance between the feature distribution redundancy and the target redundancy meeting a preset tolerance, the first qualified record encoding refined feature vector is set as the redundancy-removing feature engineering vector, expressed by the formula:

[0066]

[0067] in, For target redundancy, For tolerance, It is the absolute value;

[0068] If the condition is met, output This is to remove redundant feature engineering vectors; otherwise, it will be... As the new first qualified record coding feature, it is refined cyclically.

[0069] It can be appreciated that the feature refinement process of the loop will fall into an infinite calculation or use a fixed, suboptimal number of iterations without an explicit termination condition, which is contrary to the original intention of allocating computing resources on demand. Therefore, in response to the distance between the feature distribution redundancy and the target redundancy satisfying the preset tolerance and performing the set operation, an automated, data-driven exit is provided for this dynamic polishing process. It uses the feature distribution redundancy calculated in the previous step as a feedback signal to determine whether the feature representation has converged to the fixed point of representation described in the corpus by comparing it with a target redundancy representing the ideal convergence state. When the distance is small enough, i.e., the preset tolerance is satisfied, it means that subsequent refinement cannot bring substantial improvement. At this moment, it is the best time to terminate the loop and lock the results. This mechanism ensures that for simple asset features, the network can quickly converge and save computing overhead; while for complex asset features, it can automatically invest sufficient computing cycles until it is stable. This not only ensures that the feature vector output to the downstream probabilistic classification engine has the highest quality and discriminability, but also perfectly realizes the elegant and intelligent balance between efficiency and performance described in the corpus.

[0070] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the fusion arbitration subunit 1224 is configured to input the intermediate inference result set and the standardized discrete record stream into a fusion arbitrator module to obtain the fingerprint enriched record stream. It can be appreciated that in the overall process of enterprise asset security management, the parallel logical judgment step produces two independent information streams: one is the standardized discrete record stream containing only objective facts, and the other is the intermediate inference result set containing deterministic and probabilistic security conclusions. The two data streams are logically related but physically separated. Therefore, the intermediate inference result set and the standardized discrete record stream are input into the fusion arbitrator module to combine the newly generated security insights with the original asset facts and resolve possible conflicts to form a single, complete and consistent record. On the one hand, it accurately attaches the inferred security attributes such as vulnerabilities and risk classifications back to the original asset records through fusion operation, completing the information loop. On the other hand, it selects, combines or discards conclusions from different inference engines (deterministic and probabilistic) according to the preset strategy and confidence to ensure the accuracy and reliability of the final output results through the arbitration mechanism.

[0071] In particular, in one specific example of the present application, the fusion arbiter module consumes the stream of standardized discrete records and the set of intermediate inference results simultaneously in a streaming fashion. When the module receives a standardized record with ID Asset-001 and content {IP: 1.2.3.4, Port: 443, Service: HTTPS}, it waits for the inference results associated with this ID. Subsequently, it receives two associated inference results: one conclusion {ID: Asset-001, Vulnerability: Heartbleed} from the deterministic engine, and one conclusion {ID: Asset-001, Classification: Misconfigured_Server, Confidence: 0.85} from the probabilistic engine. The arbiter first correlates the three pieces of information by ID. Then, it applies pre-defined arbitration rules, for example: Rule 1, deterministic results have higher priority than probabilistic results; Rule 2, when there is no conflict, combine all results. In this scenario, since the two inference results do not conflict, the arbiter adopts them all and fuses them with the original record, resulting in one enriched record: {ID: Asset-001, IP: 1.2.3.4, Port: 443, Service: HTTPS, Vulnerabilities: [Heartbleed], Classifications: [{Type: Misconfigured_Server, Confidence: 0.85}]}. This record then enters the stream of enriched records.

[0072] In the enterprise internet asset intelligent discovery and security management system 100 described above, the high-efficiency stream deduplication unit 123 is configured to perform high-efficiency stream deduplication based on a hybrid strategy on the fingerprint enriched record stream to obtain a unique asset segment stream. It should be understood that, although the fingerprint enriched record stream is complete in information, due to the continuity of the asset discovery process, the parallelism of multiple tools, and the dynamic nature of the network environment, the data stream contains a large number of repetitive reports about the same asset at different time points or generated by different tools. Therefore, high-efficiency stream deduplication based on a hybrid strategy is performed on the fingerprint enriched record stream to eliminate such data-level redundancy, prevent errors and repetitive asset entities in the downstream state aggregation and knowledge graph construction process, and avoid unnecessary repeated processing of the same asset information without changes. In this way, redundant records describing the same asset segment can be accurately identified and filtered out, ensuring that only the first occurrence or asset information with significant changes can pass. By implementing a hybrid strategy that takes into account efficiency and accuracy, the system aims to determine the novelty of each incoming asset record in real time, thereby establishing a unique view of the observed asset segments.

[0073] Specifically, in one specific example of the present application, the high-efficiency stream deduplication unit consumes the fingerprint enriched record stream. The unit internally maintains two state structures: a Bloom filter in memory for fast probabilistic judgment, and an external, persistent key-value store (such as Redis or RocksDB) for accurate deterministic verification. When a fingerprint enriched record flows in, the unit first generates a unique identifier based on the core identification fields in the record (such as the combination of IP, port, and protocol). Then, the identifier is sent to the Bloom filter for query. If the Bloom filter explicitly indicates that the identifier has never appeared before, the record is determined to be a new record, and the identifier is added to the Bloom filter and the key-value store, while the record itself is output to the unique asset segment stream. If the Bloom filter indicates that the identifier may already exist, a second-stage accurate verification is initiated, i.e., the identifier is queried in the key-value store. If the query result confirms that the identifier already exists, the record is determined to be a duplicate record and is discarded directly; if the query misses (this is a false positive case of the Bloom filter), it is still treated as a new record, and the two state structures are updated and the record is output. This two-stage hybrid strategy filters out most of the duplicate data using the high efficiency of the Bloom filter, and only performs expensive accurate verification on a small amount of possible duplicate data, thereby achieving high-efficiency stream deduplication.

[0074] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the state entity aggregation unit 124 is configured to perform state entity aggregation on each unique asset fragment in the unique asset fragment stream based on an aggregation primary key to obtain the unified asset entity stream. It should be understood that in the macro view of enterprise asset security management, the unique asset fragment stream obtained through stream deduplication solves the problem of data redundancy, but it is essentially a collection of isolated facts. For example, one fragment may describe that an IP has opened port 80, another fragment describes that an Apache service is running on the IP, and a third fragment indicates that the IP is associated with a domain name. These information logically belong to the same asset entity, but are discrete in the data stream. Therefore, the state entity aggregation is further performed on each unique asset fragment in the unique asset fragment stream based on an aggregation primary key to associate and merge these scattered and atomized asset fragment information around a common identifier, thereby constructing a complete, coherent, and multi-dimensional asset entity view. By defining one or more aggregation primary keys (such as IP address, primary domain name) for each asset fragment, the system can attribute all fragments sharing the same primary key to the same logical entity. This process is stateful, meaning that the system maintains a state that evolves over time for each entity. When new related fragments flow in, it updates rather than replaces the existing entity information, thereby dynamically accumulating and fusing asset information discovered at different time points and in different dimensions. The generated unified asset entity stream provides ideal and structured node data for subsequent construction of asset knowledge graph, and successfully pieces together scattered intelligence into a clear asset portrait.

[0075] In particular, in one specific example of the present application, the stateful entity aggregation unit consumes the unique asset fragment stream. The unit first extracts the aggregation primary key for each asset fragment flowing in according to pre-set rules, for example taking the IP address field in the record as the primary key. Then, the unit performs a partitioning operation based on the primary key, routing all fragments with the same IP address to the same processing instance. The processing instance maintains a state for each IP address, which is an asset entity object being built. When the first fragment about IP address 10.0.0.1 (e.g. open port 22) arrives, the processing instance creates a new asset entity object and records the port information. When the second fragment about the same IP address 10.0.0.1 (e.g. running service SSH) arrives, the processing instance reads the existing asset entity object state and appends the new service information to the object instead of creating a new object. When the third fragment (e.g. associated domain name host1.example.com) arrives, it also updates this unique entity object. Upon meeting certain trigger conditions, for example at the end of a time window or when the entity state changes, the processing instance serializes the complete asset entity object that has aggregated multiple fragment information and outputs it to the unified asset entity stream.

[0076] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the asset knowledge graph construction and analysis module 130 is configured to perform dynamic asset knowledge graph construction and correlation analysis on the unified asset entity stream to obtain an asset knowledge graph. It should be understood that, in the overall perspective of enterprise asset security management, the unified asset entity stream generated in the previous steps provides structured and multi-dimensional asset portraits, but these portraits are essentially isolated from each other and lack explicit expression of complex mutual relationships among them. Therefore, in the technical solution of the present application, the unified asset entity stream is further subjected to dynamic asset knowledge graph construction and correlation analysis to place these independent entities in a unified relationship network to reveal and solidify the complex and implicit mutual relationships among them, thereby forming a global and networked asset view. Each asset object in the unified asset entity stream is instantiated as a node in the graph, and based on the properties within the entity and the common properties between entities, the edges representing various dependencies, subordinations or correlations are automatically inferred and generated through correlation analysis, thereby dynamically weaving a real-time relationship network depicting the entire enterprise Internet exposure surface.

[0077] In particular, in one specific example of the present application, the implementation of the process is to use a graph database (such as Neo4j, JanusGraph) as a storage backend, and a stream processing application to continuously transform entity stream data into graph structure. More specifically, the asset knowledge graph construction analysis module continuously consumes unified asset entity stream. When the module receives an asset entity object, for example {EntityType: Host, IP: 192.168.1.10, OpenPorts: [80, 443], RunningServices: [Apache / 2.4.54], AssociatedDomains: [test.example.com]}, it performs a series of graph operations. First, it sends a command to create or update a node to the graph database, such as using the MERGE statement to ensure that a Host type node representing IP address 192.168.1.10 exists, and update its properties. Then, the module performs association analysis, it parses the AssociatedDomains field, and creates a Domain type node for the domain name test.example.com (if it does not exist), and then creates an edge between the Host node and the Domain node representing the RESOLVES_TO relationship. Similarly, it creates a Service node for the service Apache / 2.4.54, and creates an edge between the Host node and the Service node representing the HOSTS relationship. As new entity objects continue to flow in, the module continuously creates or updates nodes and edges in the graph database, dynamically and incrementally constructing and enriching the entire asset knowledge graph.

[0078] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the risk detection module 140 is configured to perform risk detection on the asset knowledge graph based on a security policy to obtain a risk detection result. It should be understood that in the practice of enterprise asset security management, a complete and dynamically updated asset knowledge graph provides a global view of asset relationships, but it cannot directly reveal the security risks contained therein. Therefore, further risk detection on the asset knowledge graph based on a security policy to obtain a risk detection result can identify potential vulnerabilities and threats from massive asset data, thereby converting a static asset list into a dynamic risk view. Specifically, by extracting target assets based on a security policy, the detection resources are focused on the most critical or most likely to have risks assets, avoiding blind and inefficient comprehensive scanning. Secondly, by distinguishing technical and non-technical dimension assets and applying different detection means, the detection method is professionalized and refined, ensuring that the most appropriate analysis technology is used for different types of risk sources, thereby maximizing the accuracy and depth of risk discovery.

[0079] Specifically, in the embodiments of the present application, the risk detection module is configured to: extract a first target asset from the asset knowledge graph based on a security policy; in response to the first target asset being a technical-dimension target asset, perform adaptive vulnerability scanning and cloud configuration risk detection on the first target asset to obtain the risk detection result; and in response to the first target asset being a non-technical-dimension target asset, perform cloud service permission exposure analysis and multi-platform information leakage monitoring on the first target asset to obtain the risk detection result.

[0080] That is, the risk detection module first loads a security policy, which is defined as: performing security audit on all newly discovered server assets that are hosted on a public cloud and expose a database port externally. The module converts this policy into a graph query instruction, such as a Cypher query, to retrieve nodes that meet the conditions from the asset knowledge graph. The query returns a Host node with an IP of 3.4.5.6, which is marked as a cloud server and contains an open port 3306 in its attributes. This node is set as the first target asset. The module determines that the Host node belongs to a technical-dimension target asset, and thus triggers the corresponding detection process: it calls the adaptive vulnerability scanning engine to scan IP 3.4.5.6 using a vulnerability template for MySQL; at the same time, it calls the cloud configuration risk detection engine to query the security group rules associated with the server instance through an API, and checks whether there is a configuration that opens port 3306 to 0.0.0.0 / 0. Finally, the scanning engine reports a known remote code execution vulnerability, and the configuration detection engine reports that the security group configuration is too loose. The module integrates these two findings into a risk detection result and outputs it.

[0081] In the enterprise Internet asset intelligent discovery and security management system 100 described above, the response action generation module 150 is configured to generate an automated response action for the risk detection result based on an automated rule. It should be understood that the generation of the risk detection result only identifies the problem, and without subsequent timely disposal, the risk will continue to exist, constituting an actual threat. Therefore, an automated response action for the risk detection result is generated based on an automated rule to convert security intelligence into an executable and standardized disposal plan, to bridge the gap between detection and response, thereby shortening the risk exposure window and reducing the manual burden of the security operation team. The purpose is to accurately and deterministically match and generate the most suitable response action sequence according to the type, severity, key degree of associated assets, and other dimensions of the input risk, to realize the transition from passive alarm to active intervention.

[0082] Specifically, the response action generation module starts working upon receiving the risk detection result outputted by the preceding step. Suppose the received risk result is: {RiskID: R-123, Type: Critical_RCE_Vulnerability, Severity: Critical, TargetAsset: {Type: Host, IP: 3.4.5.6, CloudInstanceID: i-012345abcde}}. The rule engine within this module loads its rule base and matches the risk result against it. The engine finds a high-priority rule whose condition is: IF Risk.Severity is Critical AND Risk.Type contains RCE. This rule is associated with a response playbook, which defines two parallel actions. The first action is isolation, whose template is: {ActionType: Network_Isolate, Target: CloudInstance, Parameters: {InstanceID: {TargetAsset.CloudInstanceID}}}. The second action is notification, whose template is: {ActionType: Create_Ticket, System: Jira, Parameters: {Project: SEC_INCIDENT, Priority: Highest, Summary: Critical RCE detected on {TargetAsset.IP}, Assignee: SOC_Team}}. The rule engine then instantiates these two action templates with the concrete values from the risk result, generates two concrete automated response action instructions, and outputs them. These two instructions can then be subscribed and executed by the corresponding executors, thus completing the automated response loop for the risk.

[0083] In summary, the enterprise Internet asset intelligent discovery and security management system according to the embodiments of the present application is illustrated, which parses high-level asset discovery instructions and converts them into an optimized execution directed graph. This graph can automatically drive the bottom layer distributed, containerized discovery tool set to work efficiently in collaboration, thus replacing the tedious and inefficient manual tool chain operation. At the same time, to solve the problem of difficult fusion and utilization of multi-source heterogeneous data, the system introduces the original data stream generated by the discovery tool into the real-time data fusion processing layer, and through steps such as standardization, deduplication and aggregation, converts it into a unified asset entity stream with unified structure and complete information. Finally, based on this high-quality data, a dynamic asset knowledge graph is constructed, precise risk detection and automated response disposal are realized, thus systematically solving the problems of discovery efficiency and data processing of traditional solutions.

[0084] Having described various embodiments of the disclosure above, the descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical application, or improvement over the technology in the art, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A smart discovery and security management system for enterprise internet assets, characterized in that, include: The Enterprise Internet Asset Discovery Module is used to respond to asset discovery commands and perform intelligent discovery of enterprise internet assets based on optimized execution directed graphs to obtain the original asset data stream; The asset entity flow conversion module is used to convert the original asset data flow into a unified asset entity flow; The asset knowledge graph construction and analysis module is used to perform dynamic asset knowledge graph construction and association analysis on the unified asset entity flow to obtain the asset knowledge graph. The risk detection module is used to perform risk detection on the asset knowledge graph based on security policies to obtain risk detection results; The response action generation module is used to generate automated response actions for the risk detection results based on automated rules; The enterprise internet asset discovery module includes: The instruction parsing and graphing unit is used to parse the asset discovery instruction into a task instruction and graph the execution path to obtain the optimized execution directed graph. The container resource scheduling unit is used to perform graph instantiation and container resource scheduling based on the optimized directed graph to obtain a set of scheduled tool container instances; The tool container instance set driving unit is used to drive the scheduled tool container instance set to obtain the original heterogeneous output based on the topology of the optimized execution directed graph. The standardization and streaming processing unit is used to standardize and stream the original heterogeneous output based on the Sidecar agent to obtain the original asset data stream.

2. The enterprise internet asset intelligent discovery and security management system according to claim 1, characterized in that, The instruction parsing and graphing unit is used for: The asset discovery instructions are parsed and semantically extracted to obtain semantic task primitives; Based on the tool capability library, a candidate toolchain for the semantic task primitives is generated; The candidate toolchain is transformed into a directed acyclic graph, and the cost value of each graph node in the directed acyclic graph is calculated based on the execution constraints in the semantic task primitives to obtain a weighted unoptimized directed acyclic graph. The shortest path is solved on the weighted unoptimized directed acyclic graph to obtain the optimized directed graph.

3. The enterprise internet asset intelligent discovery and security management system according to claim 1, characterized in that, The asset entity flow conversion module includes: The original asset data standardization unit is used to standardize and eventify the original asset data stream from multiple heterogeneous data streams to obtain a standardized discrete record stream. A fingerprint enrichment unit is used to perform multimodal fingerprint enrichment on the standardized discrete recording stream to obtain a fingerprint-enriched recording stream; A high-efficiency streaming deduplication unit is used to perform high-efficiency streaming deduplication on the fingerprint enriched record stream based on a hybrid strategy to obtain a unique asset fragment stream. The stateful entity aggregation unit is used to perform stateful entity aggregation based on the aggregation primary key on each unique asset fragment in the unique asset fragment stream to obtain the unified asset entity stream.

4. The enterprise internet asset intelligent discovery and security management system according to claim 3, characterized in that, The fingerprint enrichment unit includes: The type conformance response subunit is used to determine whether each standardized discrete record conforms to a preset fingerprint type based on the type field of each standardized discrete record in the standardized discrete record stream. If so, a dynamic append operation is performed on the standardized discrete record to obtain a qualified record stream with feature vectors. The record stream filling subunit is used to perform atomic feature extraction and feature vector filling on each qualified record with feature vector in the qualified record stream with feature vector to obtain a qualified record stream with filled feature vector; A parallel logic judgment subunit is used to perform parallel logic judgment based on feature vectors on the qualified record stream with filled feature vectors to obtain an intermediate inference result set; The fusion arbitration subunit is used to input the intermediate inference result set and the standardized discrete record stream into the fusion arbitrator module to obtain the fingerprint enriched record stream.

5. The enterprise internet asset intelligent discovery and security management system according to claim 4, characterized in that, The parallelized logic judgment subunit is used for: Extract the original feature vector of the first qualified record from the qualified record stream that has been filled with feature vectors; The original feature vector of the first qualified record is input into the rule-based deterministic inference engine to obtain the deterministic discovery result; The original feature vector of the first qualified record is input into the feature engineering preprocessor to obtain the feature engineering vector of the first qualified record; The feature-engineered vector of the first qualified record is input into the probabilistic classification engine to obtain the probabilistic discovery result; The probabilistic and deterministic findings are aggregated to obtain intermediate inference results.

6. The enterprise internet asset intelligent discovery and security management system according to claim 5, characterized in that, The feature-engineered vector of the first qualified record is input into the probabilistic classification engine to obtain probabilistic discovery results, including: The feature engineering vector of the first qualified record is deredundant to obtain a deredundant feature engineering vector. The deredundant feature engineering vector is input into the probabilistic classification engine to obtain the probabilistic discovery result.

7. The enterprise internet asset intelligent discovery and security management system according to claim 6, characterized in that, The feature engineering vector of the first qualified record is subjected to feature redundancy removal to obtain a redundancy-free feature engineering vector, including: The feature engineering vector of the first qualified record is refined to obtain the coded refined feature vector of the first qualified record. Calculate the feature distribution redundancy of the first qualified record's encoded refined feature vector relative to the feature engineering vector of the first qualified record; In response to the distance between the feature distribution redundancy and the target redundancy meeting a preset tolerance, the first qualified record encoding refined feature vector is set as the redundancy removal feature engineering vector.

8. The enterprise internet asset intelligent discovery and security management system according to claim 1, characterized in that, The risk detection module is used for: Based on the security strategy, the first target asset is extracted from the asset knowledge graph; In response to the fact that the first target asset is a technical target asset, adaptive vulnerability scanning and cloud configuration risk detection are performed on the first target asset to obtain the risk detection results; Since the first target asset is a non-technical target asset, cloud service permission exposure analysis and multi-platform information leakage monitoring are performed on the first target asset to obtain the risk detection results.

Citation Information

Patent Citations

  • Network security protection method and system applied to regional digital and intelligent asset business

    CN120281586A