Business relationship analysis method and system, storage medium and electronic equipment
By automatically identifying business-related communities between network assets through graph databases and community discovery algorithms, the problems of manual configuration prone to errors and insufficient dynamic capture in existing technologies are solved, and efficient, accurate and intuitive business relationship analysis is achieved for network management.
Patent Information
- Application Number
- CN202510942735.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-03
AI Technical Summary
Existing network management and analysis technologies cannot automatically and accurately identify business-related communities in the network. They rely on manual configuration, which is prone to errors. They lack deep structured analysis and dynamic capture capabilities, and cannot effectively associate interaction patterns with business semantics.
Graph databases are used to abstract network assets as nodes, traffic interaction data is collected as directed relationships, time windows are defined for filtering and aggregation, UNION operators are combined with prior knowledge to generate virtual relationships, directed graphs are constructed and visualized, and community discovery algorithms are applied to divide communities.
It achieves automatic and efficient dynamic mining and visual analysis of business relationships, reduces the burden of manual configuration, adapts to network changes, improves the accuracy of community identification and the comprehensiveness of results, and outputs results that are intuitive and easy to understand.
Smart Images

Figure CN120750587A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer network technology and data processing technology, and in particular to a technology for modeling, analyzing and visualizing network traffic data using a graph database. Background Art
[0002] In modern network environments, various assets within the network, such as servers, controllers, and terminal devices, interact with each other in significant quantities. These interactions are not random or chaotic, but often naturally form relatively tight interaction clusters centered around specific business functions or workflows, known as "business-related communities." Accurately identifying these implicit business-related communities is crucial for improving network management.
[0003] Existing network management and analysis technologies have significant shortcomings in automatically and accurately completing this task. For one thing, traditional asset grouping or business system segmentation methods still rely heavily on network administrators' prior knowledge and manual configuration, such as entering information into a configuration management database (CMDB) or manually drawing network topology diagrams. This approach is not only labor-intensive and prone to human error, but also incapable of automatically discovering unknown, newly formed, or time-evolving business relationships. Furthermore, while basic network traffic monitoring tools, such as NetFlow / sFlow analyzers, can count communication information such as IP address pairs, service ports, and protocol types, they typically provide flat data lists or simple aggregated statistical results. These tools generally lack the ability to conduct in-depth, structured analysis of interaction data. Furthermore, existing methods generally lack the ability to effectively capture the dynamic nature of network interactions. Business relationships may evolve over time, but static analysis or configuration struggles to reflect these changes. Furthermore, they also fail to automatically associate identified interaction patterns with underlying "business" semantics.
[0004] Therefore, the current technical field is in urgent need of an innovative technical solution. Summary of the Invention
[0005] Purpose of the invention: To solve the problems of large workload and simple analysis results in the prior art, the present invention proposes a business relationship analysis method, system, storage medium and electronic device.
[0006] Technical solution: The present invention proposes a business relationship analysis method, including:
[0007] Collect traffic interaction data to build a graph database. Network assets are abstracted as nodes in the graph database and assigned unique identifiers as attributes. Each piece of traffic interaction data is abstracted as a directed relationship connecting two nodes.
[0008] Define a time window, capture the business relationships of the graph database within the time window, filter the business relationships according to predefined rules, and aggregate the filtered business relationships to obtain an aggregated relationship view;
[0009] The UNION operator is used to merge the aggregate relationship view with the supplementary connection based on prior knowledge to generate aggregate virtual relations based on the aggregate relationship graph and special virtual relations generated based on prior knowledge.
[0010] The aggregated virtual relationship and the special virtual relationship are both regarded as valid connection edges, and a directed graph representing the core connection relationship is constructed for visual display. The same method is used to construct an undirected graph, and the community discovery algorithm is used to divide the undirected graph into several communities, which are organized into a list for output. The list includes the community, the network assets that constitute the community, and the business relationships between the network assets.
[0011] Furthermore, the traffic interaction data comes from network traffic mirroring, NetFlow / sFlow records or traffic records of terminal probes, and the recorded information includes at least one of the source IP address, target IP address, target service port, traffic size, number of connections, and timestamp.
[0012] Furthermore, the filtering according to predefined rules includes at least one of the following:
[0013] Exclude business relationships corresponding to preset non-business ports;
[0014] Exclude business relationships corresponding to the IP addresses of preset scanners, management tools, or other non-business hosts;
[0015] The scope of the network asset analysis is limited to focus on interactions between hosts in the server area. Furthermore, defining the time window includes defining a start timestamp and an end timestamp. Aggregating includes aggregating service relationships occurring on the same source node, destination node, and service port into a class, calculating the total traffic volume and total number of connections for the class, and removing classes whose total traffic volume and total number of connections do not meet preset thresholds.
[0016] Furthermore, the directed graph is a graph data view, which is visualized using the graph database's own functions. The lines between nodes represent aggregated virtual relationships or special virtual relationships, that is, there are business interactions that meet the conditions between the nodes.
[0017] The present invention also proposes a business relationship analysis system, comprising:
[0018] The collection module collects traffic interaction data to build a graph database. It abstracts network assets into nodes in the graph database and assigns them unique identifiers as attributes. It also abstracts each piece of traffic interaction data into a directed relationship connecting two asset nodes.
[0019] An aggregation module is used to define a time window, capture business relationships in the graph database within the time window, filter the business relationships according to predefined rules, and aggregate the filtered business relationships to obtain an aggregated relationship view;
[0020] A query module is used to combine the aggregate relationship view with the supplementary connection based on prior knowledge using the UNION operator to generate aggregate virtual relationships based on the aggregate relationship graph and special virtual relationships generated based on prior knowledge;
[0021] The output module is used to treat the aggregated virtual relationship and the special virtual relationship as valid connection edges, construct a directed graph representing the core connection relationship, and perform visual display; use the same method to construct an undirected graph, use the community discovery algorithm to divide the undirected graph into several communities, and organize it into a list for output. The list includes the community, the network assets that constitute the community, and the business relationships between the network assets.
[0022] Furthermore, the traffic interaction data comes from network traffic mirroring, NetFlow / sFlow records or traffic records of terminal probes, and the recorded information includes at least one of the source IP address, target IP address, target service port, traffic size, number of connections, and timestamp.
[0023] Furthermore, the filtering according to predefined rules includes at least one of the following:
[0024] Exclude business relationships corresponding to preset non-business ports;
[0025] Exclude business relationships corresponding to the IP addresses of preset scanners, management tools, or other non-business hosts;
[0026] The scope of the network asset analysis is limited to focus on interactions between hosts in the server area. Furthermore, defining the time window includes defining a start timestamp and an end timestamp. Aggregating includes aggregating service relationships occurring on the same source node, destination node, and service port into a class, calculating the total traffic volume and total number of connections for the class, and removing classes whose total traffic volume and total number of connections do not meet preset thresholds.
[0027] Furthermore, the query operation utilizes a UNION operator to merge the aggregate relationship view with the supplementary connection based on prior knowledge, and the community discovery algorithm regards both the aggregate virtual relationship and the special virtual relationship as valid connection edges.
[0028] Furthermore, the directed graph is a graph data view, which is visualized using the graph database's own functions. The lines between nodes represent aggregated virtual relationships or special virtual relationships, that is, there are business interactions that meet the conditions between the nodes.
[0029] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the aforementioned methods when executed by a processor.
[0030] The present invention also provides an electronic device, comprising a computer program, characterized in that the computer program implements the steps of any of the aforementioned methods when executed by a processor.
[0031] Beneficial effects: This invention proposes a method for dynamic mining and visualization analysis of business relationships, which has significant advantages over existing technologies:
[0032] Automatic and efficient operation: This invention uses graph models to deeply explore the structured associations between network assets, combines rule filtering and threshold screening to focus on core business interactions, and significantly reduces the burden of manual configuration.
[0033] Dynamic adaptation to changes: This invention captures the dynamic evolution of relationships through time window analysis, can automatically adapt to changes in the network and business, and improves the accuracy of community identification and business relevance;
[0034] Comprehensive analysis foundation: This invention uses the UNION operator to combine aggregate relationship views with supplementary connections based on prior knowledge. This prior knowledge fusion mechanism can make up for the shortcomings of pure data-driven analysis and improve the comprehensiveness of the analysis.
[0035] The output results are intuitive: This paper uses a community discovery algorithm to divide the connection relationship into several business-related communities, ensuring the stability and good interpretability of the results in small and medium-sized networks. Combined with visualization methods, the results are intuitive and easy to understand. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flow chart of the present invention;
[0037] Figure 2 A flowchart of the business community discovery and result presentation of the present invention;
[0038] Figure 3 Discover visualizations for the business community;
[0039] Figure 4 This is a visualization effect diagram of the business community discovery after adding prior knowledge in the present invention. DETAILED DESCRIPTION
[0040] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0041] Example 1
[0042] like Figure 1 As shown in the figure, the core of the proposed method and system for dynamic mining and visualization analysis of business relationships based on graph databases is to utilize the powerful relationship processing capabilities of graph databases, combine time windows and business rules to conduct in-depth analysis of network traffic data, and automatically discover business-related communities between network assets through graph algorithms. The specific implementation steps are as follows:
[0043] (1) Network traffic data collection and graph database modeling
[0044] (1.1) Data Collection: Continuously collect traffic interaction data in the network. The present invention preferably collects traffic records containing key information such as source IP address, destination IP address, destination service port (specifically TCP open port), traffic size, number of connections, timestamp, etc. These data can be derived from network traffic mirroring, NetFlow / sFlow records, or terminal probes. The focus is on service-oriented TCP connection data because it can better reflect deterministic business interaction behaviors, such as application layer communication in power monitoring systems.
[0045] (1.2) Graph Database Modeling and Data Injection: The Neo4j graph database was selected as the data storage and analysis engine. Network assets (such as servers and terminals) were abstracted as nodes in the graph database and assigned unique identifiers (such as IP addresses) as attributes. Each collected traffic record that met the criteria was abstracted as a directed relationship connecting two asset nodes in the graph database.
[0046] Specifically, use a Cypher statement similar to the following to dynamically inject the collected traffic data records into the graph database:
[0047] MERGE(s:Host{ip:$source}) / / Create if the source IP node does not exist
[0048] MERGE(t:Host{ip:$target}) / / If the target IP node does not exist, create it
[0049] MERGE(s)-[r:CONNECT{port:$port,traffic:$traffic,conncount:$conncount,timestamp:$timestamp}]->(t) / / Create a CONNECT relationship from source to target, including properties such as port, traffic, number of connections, and timestamp
[0050] $source, $target, $port, $traffic, $conncount, and $timestamp are parameters extracted from collected traffic records. Host is the node label, and CONNECT is the relationship type. This approach enables continuous and incremental mapping of network interaction behaviors into a graph model.
[0051] (2) Aggregation and screening of interactive relationships based on time windows and business rules
[0052] (2.1) Define the analysis time window: To implement dynamic analysis, users must specify an analysis time window, setting the start timestamp t0 and the end timestamp t1. The system will only analyze the CONNECT relationships within this time window. This allows the system to capture business interaction patterns within a specific time period and supports sliding window analysis to observe the evolution of business relationships. This corresponds to the WHERE r.timestamp>=$t0 AND r.timestamp<=$t1 clause in a Cypher query.
[0053] (2.2) Filter non-business and interference traffic: Filter out interactions that are not strongly related to core business or are known interference based on predefined rules.
[0054] Port filtering: excludes known non-business ports, such as the SSH management port (22) and specific remote management ports (such as 10022). This corresponds to the AND r.port<>22 AND r.port<>10022 clauses. This list can be customized according to the actual network environment.
[0055] Specific source / destination filtering: Exclude IP addresses of known scanners, management tools, or other non-business hosts. For example, exclude traffic with the source IP address server5 (AND s.ip<>'server5').
[0056] Interaction Scope Limitation: Limits the scope of assets analyzed to interactions between hosts in the server area, focusing on relationships within core business systems. This corresponds to the AND s.ip STARTS WITH 'server' AND t.ip STARTS WITH 'server' clauses.
[0057] (2.3) Interaction Relationship Aggregation: Within the selected time window, aggregate multiple CONNECT relationships that pass the filtering in step (2.2) and occur on the same source node, destination node, and service port. Calculate the total traffic volume (sum(r.traffic)) and the total number of connections (sum(r.conncount)). This helps to consolidate scattered connection records into a statistically meaningful interaction intensity metric. This corresponds to the WITHs,t,r.port AS servicePort,sum(r.traffic) AS totalTraffic,sum(r.conncount) AS totalConnCount clause.
[0058] (2.4) Relationship screening based on aggregation thresholds: Aggregated interaction relationships are further screened based on their strength metrics (total traffic, total number of connections). These connections, even after aggregation, appear weak or sporadic, and may not represent stable business relationships. For example, the total number of connections after aggregation may be greater than a certain threshold (e.g., 2, WHERE totalConnCount>2) and the total traffic volume may be greater than a certain threshold (e.g., 200kb, AND totalTraffic>200). These thresholds can be adjusted based on actual business traffic characteristics.
[0059] (2.5) Generate an Aggregate Relationship View: After the above filtering and aggregation steps, the result is a set of aggregate relationships between assets that meet business rules and have a significant level of interaction intensity within the specified time window. These relationships represent the main interaction paths of the core business during that time period. This corresponds to the RETURN s,apoc.create.vRelationship(s,'AGGCONNECT',{port:servicePort,totalTraffic:totalTraffic,totalConnCount:totalConnCount},t) AS aggRel,t clause.
[0060] (3) Connection enhancement based on prior knowledge
[0061] To make up for the deficiency of relying solely on data-driven analysis, which may miss important business associations (for example, associations caused by non-TCP protocol communications, interactions below a set threshold, or known but sparsely populated business logic), this step allows the use of prior knowledge to supplement or forcibly define the connection relationships between network assets to improve the comprehensiveness and accuracy of community discovery. The prior knowledge includes known business logic, system design documents, or operation and maintenance experience.
[0062] (3.1) Joint query and virtual relationship generation
[0063] This enhancement mechanism is implemented by extending the Cypher query statement. The core is to use the UNION operator to merge the analysis results based on traffic data with the supplementary connections based on prior knowledge.
[0064] In terms of the union query structure, the query body first includes the aggregate virtual relationship (AGGCONNECT) generated by filtering and aggregating actual traffic data. Subsequently, for each pair of assets that requires a mandatory connection based on prior knowledge, this is achieved by adding one or more independent UNION clauses.
[0065] To implement knowledge-driven connections, in each UNION clause used for knowledge supplementation, explicitly match the target asset node pair using a MATCH statement. For example, a UNION clause could use MATCH(s6:Host{ip:'server6'}),(t7:Host{ip:'server7'}) to match server6 and server7. To further supplement other known relationships, such as the important business logic connection between server6 and server9, a second UNION clause could be added, using a statement similar to MATCH(s6:Host{ip:'server6'}),(t9:Host{ip:'server9'}) to match this pair of nodes.
[0066] In terms of virtual relationship generation, in the RETURN part of the UNION clause of each such knowledge supplement, a special virtual relationship representing a connection based on prior knowledge is created for the matched node pairs. The present invention preferably names it FKCONNECT (Foreign Knowledge Connect) to clearly identify its non-data-driven source. This is usually implemented by calling the virtual relationship creation function provided by the graph database, such as apoc.create.vRelationship(node1,'FKCONNECT',{port:-1,totalTraffic:0,totalConnCount:1},node2)AS aggRelationship. Among them, the assigned attribute values (such as the port is set to -1, the traffic is set to 0, and the number of connections is set to 1) are mainly used to indicate the existence of the connection, rather than the actual traffic measurement, and to ensure that it meets the data structure requirements of the subsequent processing steps.
[0067] (3.2) Integrate the results to construct a directed graph
[0068] After executing a complete Cypher query containing one or more UNION clauses, the returned result set will uniformly contain: a) AGGCONNECT relationships based on actual traffic data that meet the filtering and aggregation conditions; and b) all FKCONNECT virtual relationships generated based on prior knowledge. This integrated set of relationships (containing two or more types of relationships, but all representing connections between nodes) together constitutes the input graph structure required by the community discovery algorithm (such as connected component analysis) in step (3), and a directed graph is constructed based on this. When processing, the algorithm will treat both AGGCONNECT and FKCONNECT types of relationships as valid connection edges. This ensures that even if the actual communication between some assets is not significant, not captured by traffic data, or is below the threshold, as long as there are FKCONNECT relationships added based on prior knowledge, they can still be correctly identified as connected when performing community division, so that the business association community finally discovered is closer to the actual business structure and dependency relationship.
[0069] (4) Discovery of business-related communities based on aggregation relationships and prior knowledge
[0070] This step aims to combine the significant business interaction relationships obtained by screening and aggregation in step (2) with the business interaction relationships obtained based on prior knowledge in step (3) to discover business association communities in the network and present the results in the form of visualization and lists. Figure 2 The flowchart of step (4) is shown, which specifically includes:
[0071] (4.1) Generation and visualization of business association diagrams
[0072] (4.1.1) Graph Data Preparation: Execute a Cypher query statement. The query returns a set of nodes (s, t representing network assets) and virtual aggregation relationships connecting these nodes (aggRel, of type AGGCONNECT, containing aggregated attributes such as ports, traffic, and number of connections). Together, they form a graph data view that reflects significant business interactions within a specified time window. This view can then be supplemented by combining the prior knowledge from step (3).
[0073] (4.1.2) Visualization using the graph database's own functions: The present invention preferably utilizes the built-in visualization function of the Neo4j graph database to directly display the business association graph. The specific steps are as follows:
[0074] Step 1: Configure the graph visualization parameters in the settings of a graph database client tool (such as Neo4j Browser). In this example, the automatic connection options such as "Connect result nodes" are canceled to ensure that only AGGCONNECT relationships or FKCONNECT relationships returned by the query are displayed.
[0075] Step 2: Execute the Cypher query statement generated in step (3.1) and pass in the time window parameters $t0 and $t1 set by the user.
[0076] Step 3: After the query is successfully executed, switch to the graph view mode in the client tool. Figure 3 The following is a general business community discovery visualization diagram, Figure 4 Shown is a visualization effect diagram of business community discovery after adding prior knowledge in the present invention. Figure 3 and Figure 4 The nodes displayed on the left side of the figure represent network assets (including sever1-sever4, sever6-sever10) that participate in significant interactions or have prior knowledge support, while the connections between the nodes (including AGGCONNECT relationships and FKCONNECT relationships) represent the existence of business interactions that meet the conditions. Figure 3 and Figure 4 The right side shows the attribute information of nodes and edges, including IP address, port, traffic, and number of connections.
[0077] By observing the visualization graph, users can directly identify interconnected node clusters, which are potential business-related communities (represented by connected components of the graph). Users can use interactive operations (such as clicking, dragging, and expanding / collapsed neighbor nodes) to view detailed properties of nodes and relationships, intuitively understanding community structure and interaction details.
[0078] (4.2) Programmatic extraction of business-related community members
[0079] To obtain an accurate and automated list of business-related community members, the present invention further provides a programmatic extraction method to supplement the visual identification in step (4.1). Considering that the data direction and network connection direction are not completely consistent, this method uses a programming interface to obtain the asset pair data returned by the Cypher query in step (3), including the aforementioned AGGCONNECT relationship and FKCONNECT relationship, removes the direction attribute of the connection relationship, and uses this data to construct an undirected graph model representing the core connection relationship in memory. The construction method is consistent with the graph database view construction method in (4.1).
[0080] Subsequently, a graph theory community discovery algorithm is applied based on this undirected graph model. In this embodiment, a connected component algorithm is preferably used, which is implemented using the connected_components function of the NetworkX library. This algorithm can effectively divide the graph into several business-related communities.
[0081] In this embodiment, after steps such as aggregation and screening of interaction relationships based on time windows and business rules, and calculation of connected components, the business community discovery results without prior knowledge are divided into two independent business communities: business community 1 includes sever10, sever3, sever4, sever7, sever8, and sever9, and business community 2 includes sever1, sever2, and sever6. In contrast, the business community discovery results with prior knowledge added in the present invention are divided into one independent business community, including sever10, sever1, sever2, sever3, sever4, sever6, sever7, sever8, and sever9.
[0082] Finally, the system organizes the asset identifiers of each identified community into a structured list and outputs it, providing a precise, machine-readable list of the community's components. This list includes the community, the network assets that make up the community, and the business relationships between network assets. Community discovery results can be used to formulate security policies or detect anomalies, improving network management efficiency.
[0083] Example 2
[0084] The present invention also proposes a business relationship analysis system, comprising:
[0085] The collection module collects traffic interaction data to build a graph database. It abstracts network assets into nodes in the graph database and assigns them unique identifiers as attributes. It also abstracts each piece of traffic interaction data into a directed relationship connecting two asset nodes.
[0086] An aggregation module is used to define a time window, capture business relationships in the graph database within the time window, filter the business relationships according to predefined rules, and aggregate the filtered business relationships to obtain an aggregated relationship view;
[0087] A query module is used to combine the aggregate relationship view with the supplementary connection based on prior knowledge using the UNION operator to generate aggregate virtual relationships based on the aggregate relationship graph and special virtual relationships generated based on prior knowledge;
[0088] The output module is used to treat the aggregated virtual relationship and the special virtual relationship as valid connection edges, construct a directed graph representing the core connection relationship, and perform visual display; use the same method to construct an undirected graph, use the community discovery algorithm to divide the undirected graph into several communities, and organize it into a list for output. The list includes the community, the network assets that constitute the community, and the business relationships between the network assets.
[0089] Furthermore, the traffic interaction data comes from network traffic mirroring, NetFlow / sFlow records or traffic records of terminal probes, and the recorded information includes at least one of the source IP address, target IP address, target service port, traffic size, number of connections, and timestamp.
[0090] Furthermore, the filtering according to predefined rules includes at least one of the following:
[0091] Exclude business relationships corresponding to preset non-business ports;
[0092] Exclude business relationships corresponding to the IP addresses of preset scanners, management tools, or other non-business hosts;
[0093] The scope of the network asset analysis is limited to focus on interactions between hosts in the server area. Furthermore, defining the time window includes defining a start timestamp and an end timestamp. Aggregating includes aggregating service relationships occurring on the same source node, destination node, and service port into a class, calculating the total traffic volume and total number of connections for the class, and removing classes whose total traffic volume and total number of connections do not meet preset thresholds.
[0094] Furthermore, the query operation utilizes a UNION operator to merge the aggregate relationship view with the supplementary connection based on prior knowledge, and the community discovery algorithm regards both the aggregate virtual relationship and the special virtual relationship as valid connection edges.
[0095] Furthermore, the directed graph is a graph data view, which is visualized using the graph database's own functions. The lines between nodes represent aggregated virtual relationships or special virtual relationships, that is, there are business interactions that meet the conditions between the nodes.
[0096] Example 3
[0097] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the aforementioned methods when executed by a processor.
[0098] It will be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented in various computer languages, for example, the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0099] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0100] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0102] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0103] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
[0104] Example 4
[0105] The present invention also provides an electronic device, comprising a computer program, characterized in that the computer program implements the steps of any of the aforementioned methods when executed by a processor.
Claims
1. A business relationship analysis method, characterized in that: include: Collect traffic interaction data to build a graph database, abstract network assets into nodes in the graph database, assign unique identifiers to nodes as attributes, and abstract each piece of traffic interaction data into a directed relationship connecting two nodes; Define a time window, capture the business relationships of the graph database within the time window, filter the business relationships according to predefined rules, and aggregate the filtered business relationships to obtain an aggregated relationship view; The UNION operator is used to merge the aggregate relationship view with the supplementary connection based on prior knowledge to generate aggregate virtual relations based on the aggregate relationship graph and special virtual relations generated based on prior knowledge. The aggregate virtual relationship and the special virtual relationship are both regarded as connecting edges, and a directed graph is constructed by combining the nodes for visual display; Call a programming interface to query the aggregated virtual relationship and the special virtual relationship, regard the network asset pairs obtained from the query as connection edges, combine the nodes to construct an undirected graph representing the core connection relationship, use a community discovery algorithm to divide the undirected graph into several communities, and organize them into a list for output. The list includes the communities, the network assets that constitute the communities, and the business relationships between the network assets.
2. The business relationship analysis method according to claim 1, characterized in that: The traffic interaction data comes from network traffic mirroring, NetFlow / sFlow records or traffic records of terminal probes, and the recorded information includes at least one of the source IP address, target IP address, target service port, traffic size, number of connections, and timestamp.
3. The business relationship analysis method according to claim 1, characterized in that: The filtering according to predefined rules includes at least one of the following: Exclude business relationships corresponding to preset non-business ports; Exclude business relationships corresponding to the IP addresses of preset scanners, management tools, or other non-business hosts; Limit the scope of the analysis of network assets to focus on interactions between hosts in the server area.
4. The business relationship analysis method according to claim 1, characterized in that: Defining the time window includes defining a start timestamp and an end timestamp.
5. The business relationship analysis method according to claim 1, characterized in that: The aggregation includes aggregating business relationships occurring on the same source node, target node and service port into a class, calculating the total traffic size and total connection times of the class, and removing classes whose total traffic size and total connection times do not meet preset thresholds.
6. The business relationship analysis method according to claim 1, characterized in that: The directed graph is a graph data view, which is visualized using the graph database's own functions, and the lines between nodes represent aggregated virtual relationships or special virtual relationships.
7. A business relationship analysis system, characterized in that: include: The collection module collects traffic interaction data to build a graph database. It abstracts network assets into nodes in the graph database and assigns them unique identifiers as attributes. It also abstracts each piece of traffic interaction data into a directed relationship connecting two asset nodes. An aggregation module is used to define a time window, capture business relationships in the graph database within the time window, filter the business relationships according to predefined rules, and aggregate the filtered business relationships to obtain an aggregated relationship view; A query module is used to combine the aggregate relationship view with the supplementary connection based on prior knowledge using the UNION operator to generate aggregate virtual relationships based on the aggregate relationship graph and special virtual relationships generated based on prior knowledge; an output module, configured to regard the aggregated virtual relationship and the special virtual relationship as connecting edges, and construct a directed graph in combination with the nodes for visual display; Call a programming interface to query the aggregated virtual relationship and the special virtual relationship, regard the network asset pairs obtained from the query as connection edges, combine the nodes to construct an undirected graph representing the core connection relationship, use a community discovery algorithm to divide the undirected graph into several communities, and organize them into a list for output. The list includes the communities, the network assets that constitute the communities, and the business relationships between the network assets.
8. The business relationship analysis system according to claim 7, characterized in that: The traffic interaction data comes from network traffic mirroring, NetFlow / sFlow records or traffic records of terminal probes, and the recorded information includes at least one of the source IP address, target IP address, target service port, traffic size, number of connections, and timestamp.
9. The business relationship analysis system according to claim 7, characterized in that: The filtering according to predefined rules includes at least one of the following: Exclude business relationships corresponding to preset non-business ports; Exclude business relationships corresponding to the IP addresses of preset scanners, management tools, or other non-business hosts; Limit the scope of the analysis of network assets to focus on interactions between hosts in the server area.
10. The business relationship analysis system according to claim 7, characterized in that: Defining the time window includes defining a start timestamp and an end timestamp.
11. The business relationship analysis method according to claim 7, characterized in that: The aggregation includes aggregating business relationships occurring on the same source node, target node and service port into a class, calculating the total traffic size and total connection times of the class, and removing classes whose total traffic size and total connection times do not meet preset thresholds.
12. The business relationship analysis system according to claim 7, characterized in that: The directed graph is a graph data view, which is visualized using the graph database's own functions, and the lines between nodes represent aggregated virtual relationships or special virtual relationships.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
14. An electronic device comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.